Role Overview
The Data Engineer is responsible for designing, developing, and maintaining scalable data pipelines and robust data architecture within the AWS ecosystem. This role focuses on building high-performance data processing solutions, executing complex data migrations, and ensuring data integrity across diverse cloud-based environments. The engineer will bridge the gap between raw data and actionable insights through advanced processing and architectural design.
Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using PySpark for large-scale data processing.
- Build and optimize complex data workflows and transformation logic to handle diverse datasets.
- Execute end-to-end data migration strategies, moving data from legacy systems (such as SAS environments) to modern cloud architectures.
- Design and implement robust data models and schemas to support advanced analytics and business intelligence.
- Manage data ingestion, cleansing, validation, and orchestration processes across various formats (Parquet, Avro, JSON, CSV, etc.).
- Leverage AI-assisted coding tools (such as GitHub Copilot) to enhance productivity and ensure high code quality.
- Architect and implement data solutions utilizing AWS components (e.g., AWS Glue, EMR, S3, Redshift, Lambda, Kinesis).
- Ensure strict adherence to cloud security best practices, including identity management, encryption, and data protection.
- Optimize the performance, reliability, and cost-efficiency of cloud data workloads.
- Troubleshoot and proactively resolve data pipeline failures and production issues to ensure system stability.
- Collaborate with cross-functional teams to design and deploy scalable, cloud-native data architectures.
Requirements
- Strong hands-on experience with PySpark and distributed computing frameworks.
- Proven experience working with SAS for data manipulation and statistical analysis.
- Deep expertise in AWS services (S3, Glue, EMR, Redshift, Athena, etc.).
- Expert knowledge of data migration patterns, schema mapping, and validation techniques.
- Strong understanding of Data Modeling, Star/Snowflake schemas, and Data Lakehouse architecture.
- Proficiency in SQL, Python, and advanced data structures.
- Proficiency in DevOps practices and CI/CD pipelines using AWS Code Pipeline or GitHub Actions.
- Experience with workflow orchestration tools like Apache Airflow or AWS Step Functions.
- Knowledge of real-time data processing using AWS Kinesis or Kafka.
- B. tech / BE in Computer Science or Information Technology.