Role Overview
We are looking for a skilled Big Data Engineer with strong experience in Scala, PySpark, AWS Glue, Databricks, Delta Lake, and Amazon EMR to design, develop, and optimize scalable data engineering solutions. The ideal candidate should have hands-on experience in big data pipeline development, data validation and profiling, data quality frameworks, performance optimization, and Snowflake migration.
Responsibilities
- Design, develop, and maintain scalable Big Data pipelines using Scala, PySpark, AWS Glue, and Databricks.
- Develop robust ETL/ELT pipelines for structured, semi-structured, and large-volume datasets.
- Work with Apache Spark/PySpark for distributed data processing and transformation.
- Develop and optimize data processing solutions using Delta Lake and Parquet.
- Build and manage data pipelines using AWS Glue and Amazon EMR.
- Perform data validation, profiling, reconciliation, and quality checks across data pipelines.
- Implement and maintain Data Quality Frameworks to ensure data accuracy, completeness, consistency, and reliability.
- Participate in Snowflake migration projects, including data extraction, transformation, loading, validation, and reconciliation.
- Optimize Spark jobs, SQL queries, data transformations, partitioning, caching, and resource utilization for improved performance.
- Troubleshoot production data pipelines and resolve data quality, performance, and processing issues.
- Implement appropriate error handling, logging, monitoring, and recovery mechanisms within data pipelines.
- Work with Delta Lake features such as ACID transactions, schema evolution, partitioning, and optimization.
- Utilize AI Coding Assistants to improve development productivity, code quality, debugging, documentation, and testing.
- Collaborate with data architects, analysts, application teams, and cloud engineers to deliver end-to-end data solutions.
- Follow best practices for coding, version control, CI/CD, security, and data governance.
Requirements
- Strong hands-on experience with Scala and PySpark.
- Strong understanding of Apache Spark architecture and distributed data processing.
- Experience developing enterprise-grade Big Data pipelines.
- Strong SQL skills for data transformation and validation.
- Hands-on experience with AWS Glue.
- Experience with Amazon EMR and Spark-based workloads.
- Strong experience with Databricks.
- Hands-on experience with Delta Lake, Delta tables, partitioning, schema evolution, and optimization.
- Experience with Parquet and other distributed data storage formats.
- Experience in ETL/ELT pipeline development.
- Experience implementing Data Quality Frameworks.
- Hands-on experience or strong understanding of Snowflake migration projects.
- Experience optimizing Spark jobs and Big Data pipelines.
- Experience using AI-assisted coding tools such as GitHub Copilot or similar.
Skills
- Scala
- PySpark
- AWS Glue
- Databricks
- Delta Lake
Nice to Have
- Experience with AWS S3, Lambda, Step Functions, CloudWatch, or related AWS services.
- Knowledge of Apache Airflow or other workflow orchestration tools.
- Experience with CI/CD tools such as Git, Jenkins, GitHub Actions, or Azure DevOps.
- Knowledge of data governance, metadata management, and data security.
- Experience working in Agile/Scrum environments.
- Exposure to Terraform/IaC and cloud automation.