Role Overview
Amgen harnesses the best of biology and technology to fight the world’s toughest diseases. As an Associate Data Engineer, you will develop, test, and maintain data pipelines to support analytics, reporting, and machine learning use cases, ensuring reliable datasets for a variety of cross-functional teams.
Responsibilities
- Develop, test, and maintain data pipelines using Databricks, PySpark, and Python.
- Ingest, transform, and process structured and semi-structured data from multiple sources.
- Support the development of scalable ETL/ELT workflows for analytics, reporting, and machine learning use cases.
- Perform data cleansing, validation, and quality checks to ensure accuracy and consistency.
- Optimize Spark jobs and Databricks notebooks for performance, reliability, and cost efficiency.
- Assist in troubleshooting pipeline failures, data issues, and performance bottlenecks.
- Support basic AI/ML data preparation activities, including feature engineering and dataset creation.
Requirements
- 2-6 years of experience with a Bachelor’s degree in Computer Science, Data Engineering, or a related field.
- Hands-on experience with Python for data processing and automation.
- Strong working knowledge of PySpark and distributed data processing.
- Proven experience using Databricks (notebooks, clusters, jobs, Delta tables).
- Experience working with Delta Lake and lakehouse architecture.
- Working knowledge of SQL for querying and transforming data.
- Familiarity with cloud-based platforms such as AWS, Azure, or GCP.
- Understanding of Git or other version control tools.
Skills
- Databricks
- PySpark
- Python
- SQL
- Delta Lake