Role Overview
At Bristol Myers Squibb, we are seeking an early-career Data Engineer passionate about creating high-quality scientific data products and contributing to the development of data fabric to support advanced analytics, AI, machine learning, and scientific decision-making across Pharmaceutical Product Development. The ideal candidate will be an engineer who sees data as a strategic asset that powers scientific discovery and AI innovation, working on complex datasets and modern cloud technologies.
Responsibilities
- Build scientific data products for product development focused on datasets for molecular features, material properties, laboratory data, process and manufacturing parameters, stability and product performance data.
- Contribute to initiatives for data structuring and data contextualization.
- Develop and maintain scalable data pipelines supporting analytics, AI, and scientific modelling initiatives.
- Transform raw scientific and operational data into trusted, model-ready data products.
- Work with cross-functional teams to improve data accessibility, reliability, and quality.
- Support enterprise initiatives involving Databricks, Data Fabric, and cloud-native architectures.
- Automate data workflows and reduce manual effort through engineering best practices.
- Contribute to development of reusable data assets supporting Product Development innovation across US, Europe and India.
Requirements
- Bachelor's or Master’s degree in Computer Science, Chemical Engineering, Information Systems, Bioinformatics, Biotechnology or related field with 2+ years of industry experience.
- Technical hands-on experience with modern data technologies.
- Proficiency in SQL, Python, ETL/ELT Development, Delta Lake, Lakehouse Architecture, Medallion Architecture, Data Modelling, Data Warehousing, Distributed Computing, and Data Validation.
- Experience with Databricks, dbt, and AWS.
- Knowledge of Data Product Design, Vector Databases, Data Quality Engineering, Data Observability, Metadata Management, Master Data Management, Data Lineage, and Data Governance principles.
- Familiarity with AI-Ready Data Foundations, Vector Database fundamentals, Semantic Layer Design, Knowledge Graph Concepts, and Data Foundations for GenAI & Agentic AI Applications.
- Software Engineering & Delivery Practices including Git, API Integration, Workflow Automation, CI/CD Fundamentals, and Agile Delivery.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration abilities.
Skills
- Python
- SQL
- AWS
- Databricks
- dbt