Role Overview
Design, implementation, and management of data platform, data pipelines to help fulfill data access and aggregation functionality for users in cloud environment.
Responsibilities
- Design, build, and manage scalable data pipelines for data extraction, transformation, and loading (ETL).
- Integrate data from various sources, ensuring data quality and consistency.
- Identify and implement process improvements to enhance data reliability, efficiency, and quality.
- Work with data scientists, analysts, and other stakeholders to support their data infrastructure needs.
- Monitor data pipelines and systems, troubleshoot issues, and ensure data integrity.
- Maintain comprehensive documentation of data processes and systems.
Requirements
- B.E / B. Tech in Computer Science/IT/ or Masters in Computer Application(MCA) with minimum 60% marks.
- Minimum 3-5 years of experience in data engineering or a related field.
- Proficiency in Py-Spark, SQL, Python, and data platforms like Databricks.
- Experience in storage mechanisms in AWS S3, Layers of Medallion Architecture, OTF like Delta, Iceberg etc.
- Hands-on in Data schema like Star, Snowflake etc.
- Good knowledge of next generation ETL/ELT tools and data consolidation frameworks.
- Ability to work independently or within a team in a fast-paced AGILE-SCRUM environment.
Skills
- Py-Spark
- SQL
- Python
- Databricks
- AWS S3
Nice to Have
- Certification in Big Data/Data Science/Cloud Computing or equivalent.
- Experience in IoT time series data domain.
- Exposure to automotive data.