Role Overview
We are seeking an experienced ETL Developer with 4 to 8 years of experience to assist in solution design, delivery, and the build of new ETLs. The ideal candidate will have strong expertise in Big Data and the ability to implement end-to-end serverless data architectures using AWS services.
Responsibilities
- Build and maintain high volume ETL/ELT pipelines across Hadoop and AWS.
- Develop distributed data processing solutions using PySpark and Spark SQL.
- Implement reusable data ingestion frameworks and orchestration processes.
- Optimize data workflows using partitioning, bucketing, and file formats like Parquet or ORC.
- Define technical specifications and make architecture decisions for long-term scalability.
- Implement best practices including unit tests, automation, and code reviews.
- Troubleshoot Spark performance issues, job failures, and cluster bottlenecks.
- Collaborate with business stakeholders to translate data requirements into technical solutions.
Requirements
- 4 to 8 years of overall ETL experience.
- Strong expertise in Big Data (Spark, Cloudera).
- Experience orchestrating workflows using AWS Step Functions.
- Proficiency in Python, Scala, or Shell scripting.
- Strong experience with SQL, HiveQL, and Impala.
- Experience with CI/CD, GitHub, and Git.
- Understanding of data modelling (star/snowflake) and schema evolution.
- Familiarity with serverless patterns and containerization (Docker, ECS/EKS).
Skills
- PySpark
- AWS
- Apache Spark
- Hadoop
- Python
Nice to Have
- AWS certifications (Data Engineer or Developer).
Benefits
- Competitive total rewards package.
- Continuing education and training.
- Tremendous potential with a growing worldwide organization.