Role Overview
We are seeking a skilled and motivated Big Data Engineer with 2-4 years of experience in designing, developing, and maintaining large-scale data processing systems. The ideal candidate should have strong expertise in AWS cloud services, Python, and PySpark, along with hands-on experience in building scalable data pipelines, data lakes, and ETL solutions. You will work closely with data architects, analysts, and business stakeholders to deliver high-quality data solutions that support analytics and business intelligence initiatives.
Responsibilities
- Design, develop, and maintain scalable and reliable data pipelines for processing large volumes of structured and unstructured data.
- Build and optimize ETL/ELT workflows using PySpark and Python.
- Develop and manage data lake and data warehouse solutions on AWS Cloud.
- Implement data ingestion frameworks from multiple data sources, including databases, APIs, files, and streaming platforms.
- Leverage AWS services such as S3, Glue, EMR, Lambda, Redshift, Athena, CloudWatch, and IAM.
- Perform data transformation, cleansing, validation, and enrichment to ensure high data quality.
- Optimize Spark applications for performance, scalability, and cost efficiency.
- Collaborate with data scientists, analysts, and business teams to understand data requirements.
- Monitor data pipelines and resolve production issues to ensure data availability and reliability.
- Implement security best practices, governance, and compliance standards.
- Participate in code reviews, unit testing, and deployment activities.
- Create technical documentation and maintain operational runbooks.
Requirements
- 2-4 years of experience in Big Data Engineering.
- Strong experience in Python programming.
- Hands-on expertise in PySpark and distributed data processing.
- Experience working with AWS Cloud Platform.
- Knowledge of AWS services including S3, Glue, EMR, Lambda, Redshift, Athena, IAM, and CloudWatch.
- Good understanding of ETL/ELT concepts and data pipeline development.
- Experience with SQL and relational databases.
- Knowledge of Data Warehousing concepts and dimensional modeling.
- Familiarity with version control tools such as Git.