Role Overview
We are seeking a Databricks Engineer to design, build, and operate a Data & AI platform with a strong foundation in the Medallion Architecture (raw/bronze, curated/silver, and mart/gold layers). This platform will orchestrate complex data workflows and scalable ELT pipelines to integrate data from enterprise systems such as PeopleSoft, D2L, and Salesforce, delivering high-quality, governed data for machine learning, AI/BI, and analytics at scale.
Responsibilities
- Design, implement, and optimize end-to-end data pipelines on Databricks, following the Medallion Architecture principles.
- Build robust and scalable ETL/ELT pipelines using Apache Spark and Delta Lake to transform raw data into curated and analytics-ready layers.
- Operationalize Databricks Workflows for orchestration, dependency management, and pipeline automation.
- Connect and ingest data from enterprise systems such as PeopleSoft, D2L, and Salesforce using APIs, JDBC, or other integration frameworks.
- Develop data quality checks, validation rules, and anomaly detection mechanisms to ensure data integrity.
- Integrate monitoring and observability tools like Grafana to track ETL performance.
- Implement Unity Catalog for centralized metadata management, data lineage, and governance policy enforcement.
- Support AIOps/MLOps lifecycle workflows using MLflow for experiment tracking and model registry.
- Architect and manage data lakes on Azure Data Lake Storage (ADLS) or Amazon S3.
Requirements
- Proven experience designing and implementing end-to-end data pipelines using Databricks.
- Strong expertise in Apache Spark and Delta Lake for large-scale data processing.
- Experience with Medallion Architecture (Bronze, Silver, Gold layers).
- Knowledge of data governance and security using Unity Catalog.
- Familiarity with MLOps tools like MLflow.
- Experience with cloud storage solutions such as ADLS or S3.
Skills
- Databricks
- Apache Spark
- Delta Lake
- MLflow
- Python