Role Overview
We are seeking an experienced Big Data Engineer with strong expertise in Scala, PySpark, AWS Glue, Big Data Pipeline Development, and Data Validation & Profiling. The ideal candidate will be responsible for designing, developing, and optimizing scalable big data pipelines and data processing solutions. You will work with modern cloud data technologies and contribute to data migration, data quality, validation, and performance optimization initiatives.
Responsibilities
- Design, develop, and maintain scalable big data pipelines using Scala, PySpark, and AWS Glue.
- Develop robust ETL/ELT workflows for large-scale data processing.
- Build and optimize distributed data processing solutions using Apache Spark.
- Perform comprehensive data validation, profiling, and quality checks across data pipelines.
- Develop reusable frameworks for data ingestion, transformation, validation, and processing.
- Work with Databricks, Delta Lake, AWS EMR, Snowflake, and other modern data platforms.
- Support Snowflake migration and cloud data modernization initiatives.
- Process and optimize data stored in formats such as Parquet.
- Implement data quality frameworks and automated validation processes.
- Identify and resolve data inconsistencies, pipeline failures, and performance bottlenecks.
- Optimize Spark jobs, data pipelines, queries, and resource utilization for performance and scalability.
- Collaborate with Data Architects, Data Engineers, Analysts, QA teams, and business stakeholders.
- Follow software engineering best practices including code reviews, version control, testing, and documentation.
- Leverage AI coding assistants to improve development productivity, code quality, and engineering efficiency.
- Monitor production pipelines and troubleshoot data processing issues.
Requirements
- 4+ years of experience in Big Data Engineering or Data Engineering.
- Strong hands-on experience with Scala.
- Strong experience with PySpark and Apache Spark.
- Hands-on experience with AWS Glue.
- Strong experience in big data pipeline development.
- Experience with data validation and data profiling.
- Strong understanding of distributed data processing and ETL/ELT concepts.
- Experience working with large-scale datasets and cloud-based data platforms.
- Good understanding of data quality, data consistency, and validation techniques.
- Strong analytical and problem-solving skills.
Skills
- Scala
- PySpark
- AWS Glue
- Big Data Pipeline Development
- Data Validation & Profiling