Role Overview
We are looking for a Jr. data engineer to design and develop scalable data pipelines and manage cloud-based data architectures. You will work with large-scale data formats and ensure the quality, performance, and reliability of our data ecosystems.
Responsibilities
- Design and develop scalable data pipelines using Azure Data Factory (ADF) and equivalent AWS services.
- Work with Data Lakes, Lakehouse architectures, Delta Lake, and Data Warehouses across cloud platforms.
- Develop efficient data processing solutions using Apache Spark / PySpark.
- Build and maintain Semantic Models for analytics and reporting.
- Develop and optimize Advanced SQL queries, transformations, and data models.
- Work with large-scale formats such as Parquet and implement efficient storage and processing strategies.
- Apply Spark optimization techniques including partitioning, caching, broadcast joins, shuffle optimization, and query tuning.
- Work with cloud data services such as Azure Data Lake, Synapse, Fabric, S3, Glue, Athena, Redshift, and EMR.
- Ensure data quality, performance, scalability, security, and reliability of data pipelines.
Requirements
- 2–3 years of hands-on Data Engineering experience.
- Strong expertise in Azure Data Factory (ADF).
- Strong Advanced SQL skills.
- Good hands-on experience with Apache Spark / PySpark.
- Strong understanding of Data Lake, Lakehouse, Delta Lake, and Data Warehouse concepts.
- Good understanding of Parquet and columnar data formats.
- Experience with Semantic Models / Power BI datasets.
- Understanding of ETL/ELT, data modeling, partitioning, and Spark performance optimization.
- Good understanding of cloud-based data engineering architectures.
- 3 years of experience with Azure.
Skills
- Azure Data Factory
- Apache Spark
- PySpark
- SQL
- Delta Lake
Benefits
- Flexible schedule
- Provident Fund