Role Overview
The Data Engineer, Advanced Analytics platforms will work with our core platform development team as well as domain experts, application developers, controls engineers and data scientists. Their primary responsibility will be to develop reliable and instrumented data ingestion pipelines that land inbound data from multiple process and operational data stores throughout the company to on-premises and cloud-based data lakes. These pipelines will require data validation and data profiling automation along with version control and CI/CD to ensure ongoing resiliency and maintainability of the inbound data flows supporting our advanced analytics projects.
Responsibilities
- Design, test, deploy and maintain production big-data ingestion pipelines using established frameworks, patterns of practice, agile software development and CI/CD practices.
- Work with cross-organizational data source teams to define data ingestion requirements for structured, unstructured and semi-structured data.
- Define and implement automated validation and profiling capabilities needed to ensure reliable data delivery.
- Work with data source teams, domain experts and data scientists to define data cleansing and data enrichment requirements for landed data.
- Implement data cleansing and enrichment code using established patterns of practice.
- Actively participate in code reviews and technical information sharing.
- Provide support in a DevOps environment to monitor tokens, jobs and overall system performance.
Requirements
- Bachelor's degree in computer science, engineering, mathematics, or a related technical discipline.
- Understand concepts of big data engineering, developing and maintaining ETL and ELT pipelines for data warehousing, on-premise and cloud data lake environments.
- Demonstrated production programming proficiency in at least one modern JVM language such as Java, Scala or Kotlin, as well as an interpreted declarative programming language such as Python.
- Entry level experience with AWS platform services, including AWS S3 & EC2, Data Migration Services (DMS), RDS, EMR, RedShift, Lambda, DynamoDB, CloudWatch, CloudTrail.
- Hands-on technical familiarity with Apache Spark architecture, S3, parquet and Delta Lake architecture, technologies and tools.
Skills
- Python
- Java
- AWS
- Apache Spark
- ETL