Role Overview
We are looking for a Data Engineer to design, develop, and deploy data ingestion and transformation pipelines. You will work extensively within the Google Cloud Platform (GCP) ecosystem to ensure high-quality data delivery and pipeline reliability.
Responsibilities
- Design, develop, and deploy data ingestion and transformation pipelines using Apache Spark (Scala).
- Work within Google Cloud Platform (GCP) components including Dataproc, BigQuery & Cloud Storage.
- Write complex SQL queries for data analysis, validation, and transformation in BigQuery.
- Develop Unix shell scripts for automation, data movement, and pipeline orchestration.
- Optimize and troubleshoot data pipelines for performance, scalability, and reliability.
- Collaborate with data engineers, analysts, and business stakeholders to ensure data quality.
- Contribute to code reviews, documentation, and CI/CD integration of data workflows.
Requirements
- 4–6 years of hands-on experience in data engineering or related roles.
- Proven experience developing Spark applications in Scala.
- Strong experience working in Google Cloud Platform (GCP) ecosystem — including Dataproc, BigQuery, Cloud Storage.
- Proficient in SQL (especially BigQuery SQL dialects).
- Strong experience with Unix/Linux scripting for data automation.
- Familiarity with version control (Git) and CI/CD processes.
- Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
Nice to Have
- Experience with Python for data processing or automation.
- Knowledge of data governance, data quality, or metadata management best practices.
Skills
- Apache Spark
- Scala
- Google Cloud Platform
- SQL
- Unix