Role Overview
We are looking for an experienced AWS Data Engineer with expertise in building scalable data pipelines and modern data lake solutions on AWS. The ideal candidate should have hands-on experience with Amazon EMR, Apache Iceberg, and AWS data services to design, develop, and optimize batch and streaming data processing solutions.
Responsibilities
- Design, develop, and maintain scalable data pipelines on AWS.
- Build and optimize ETL/ELT workflows using Python, PySpark, and Apache Spark.
- Develop and manage data lake solutions using Apache Iceberg on Amazon S3.
- Configure and manage Amazon EMR clusters for large-scale data processing.
- Process streaming data using Amazon Kinesis (preferred).
- Work with AWS services such as Glue, Athena, Redshift, Lambda, and S3.
- Optimize data processing performance, partitioning, and query execution.
- Ensure data quality, reliability, security, and governance.
- Collaborate with data analysts, data scientists, and application teams to deliver business requirements.
- Monitor, troubleshoot, and enhance existing data pipelines.
Requirements
- 6+ years of experience as a Data Engineer.
- Strong programming skills in Python.
- Hands-on experience with Apache Spark / PySpark.
- Experience with Amazon EMR.
- Strong knowledge of Apache Iceberg.
- Experience with Amazon S3 and AWS data services.
- Good understanding of SQL and relational databases.
- Experience with ETL/ELT pipeline development.
- Familiarity with Git and CI/CD practices.