Role Overview
We are looking for a Data Engineer to develop and maintain robust data pipelines, ELT processes, and workflow orchestration to ensure efficient and reliable data delivery.
Responsibilities
- Develop and maintain data pipelines, ELT processes, and workflow orchestration using Apache Airflow, Python, PySpark, Hive(Trino) and Snowflake.
- Design and implement custom connectors to facilitate the ingestion of diverse data sources, including structured and unstructured data.
- Collaborate with cross-functional teams to translate data needs into technical solutions.
- Design and implement data CI/CD pipelines for automated integration and deployment.
- Monitor and troubleshoot pipelines to resolve ingestion, transformation, and loading issues.
- Conduct data validation and testing to ensure accuracy, consistency, and compliance.
- Document data workflows and technical specifications for data governance.
Requirements
- Proven experience in designing and implementing scalable data architectures.
- Strong understanding of data orchestration and ETL/ELT methodologies.
- Experience with CI/CD processes in a data environment.
Skills
- Apache Airflow
- Python
- PySpark
- Hive(Trino)
- Snowflake