Role Overview
We are seeking a Data Engineer specializing in Agentic AI to design, build, and maintain ETL/ELT pipelines within an AWS environment. This role focuses on enabling real-time and batch data processing to support advanced AI-driven workflows and model training.
Responsibilities
- Design, build, and maintain ETL/ELT pipelines for structured and unstructured data in AWS.
- Work with AWS services such as S3, Glue, EMR, Lambda, Kinesis, Step Functions, Redshift, Athena, DynamoDB, and RDS.
- Ensure data quality, governance, and security across the full lifecycle.
- Enable real-time and batch data processing to support AI-driven workflows.
- Collaborate with AI/ML teams to prepare datasets for model training, inference, and fine-tuning.
- Optimize data infrastructure for scalability, cost-efficiency, and performance.
- Implement CI/CD pipelines and best practices for data engineering in a cloud-native environment.
- Monitor and troubleshoot data pipelines to ensure high availability and reliability.
Requirements
- 3–7 years of experience in data engineering.
- Strong programming skills in Python, SQL, and/or Scala/Java.
- Expertise in the AWS cloud ecosystem.
- Experience with data pipeline orchestration tools like Airflow, Step Functions, or Dagster.
- Proficiency with big data frameworks such as Spark, Hadoop, or Flink.
- Familiarity with data modeling, warehousing, and schema design.
- Solid understanding of data governance, lineage, and security (IAM, Lake Formation, encryption).
- Experience with real-time streaming data using Kafka or Kinesis.
- Knowledge of DevOps practices including Terraform, CloudFormation, CI/CD, Git, and Docker.
Skills
- AWS
- Python
- Apache Spark
- SQL
- Airflow