Key Requirements
· Hands-on experience as a Data Engineer, designing and building scalable data pipelines and cloud-based data processing solutions.
· Strong hands-on expertise in Python and SQL, including data transformation, performance optimization, data modelling, workflow automation and large-scale data processing.
· Experience working with Google Cloud Platform (GCP) services such as Big Query, Cloud Storage and Cloud Functions, or equivalent cloud-native data platforms, is mandatory.
· Hands-on experience with Databricks, including notebooks, clusters, Spark/PySpark development, data transformation, and pipeline orchestration.
· Strong understanding of distributed data processing frameworks, particularly PySpark/Spark Data Frames, for large-scale data transformation and analytics workloads.
· Experience working with relational databases such as PostgreSQL, SQL Server, Oracle or similar platforms, including complex SQL development, query optimization and data modelling.
· Proven experience developing and maintaining ETL/ELT pipelines, including ingestion from APIs, databases, cloud storage platforms and third-party data sources.
· Understanding of modern data architecture patterns such as Medallion Architecture (Bronze/Silver/Gold), Data Lake, Data Warehouse, and Lakehouse architectures.
· Experience working with version control systems and CI/CD pipelines using tools such as GitHub Actions, GitLab CI/CD, Azure DevOps or equivalent platforms.
· Experience working with modern data processing and analytics platforms such as Databricks, Snowflake, Redshift, Synapse Analytics or similar technologies is desirable.
· Experience designing and implementing reusable Python libraries, utility packages, frameworks, and automation solutions for data engineering workflows.
· Familiarity with cloud-based orchestration and workflow management tools such as Apache Airflow, Data form, Azure Data Lake, Azure SQL DW, Azure Synapse, etc. or equivalent platforms is good to have.
· Experience integrating and processing data from diverse sources, including APIs, cloud storage, structured databases and file formats such as JSON, CSV and Parquet.
· Exposure to containerization and cloud-native deployment technologies such as Docker and Kubernetes would be an advantage.
· Strong problem-solving skills with the ability to independently analyze business requirements, propose scalable data solutions, and collaborate effectively with analysts and stakeholders.
· Any understanding of Web Analytics data such as Adobe Analytics and Google Analytics alongside any Ad tech like Google Ad Manager (working with programmatic and affiliate partners) is a plus, but not essential.
Hyderabad, Telangana, India
Hyderabad, Telangana, India
Data Engineer
(Annual Package)
Lorem ipsum dolor sit amet, consectetur adipiscing elit.