Role Overview
This role requires extensive experience in data engineering, data platforms, and analytics, with a strong consulting background. The focus is on designing, developing, and optimizing scalable data pipelines using Databricks and Python, ensuring high performance and reliability for both batch and real-time processing.
Responsibilities
- Design, develop, and optimize scalable data pipelines on Databricks using Apache Spark (PySpark/Scala), ensuring high performance and reliability for batch and real-time processing.
- Implement data engineering best practices including Delta Lake, data modeling, partitioning, and performance tuning; manage data workflows using Databricks Workflows or orchestration tools.
- Collaborate with data scientists and analysts to build and deploy machine learning models and analytics solutions, leveraging Databricks notebooks, MLflow, and Unity Catalog for governance.
- Ensure data quality, security, and compliance by applying monitoring, logging, access controls, and CI/CD pipelines (Azure DevOps/Git), supporting end-to-end data lifecycle management.
Requirements
- 7+ years of experience in data engineering, data platforms & analytics, 10+ years of consulting experience.
- Minimum 6-8+ projects delivered with hands-on experience in development on Databricks.
- Strong hands-on experience in Python programming. Expertise in:
- Pandas (data manipulation, transformation, analysis)
- NumPy (array operations, numerical computing)
- Working knowledge of two or more common Cloud ecosystems (AWS, Azure, GCP) with deep expertise in at least one.
- Deep experience with distributed computing with Spark with knowledge of Spark runtime internals.
- Familiarity with CI/CD for production deployments.
- Working knowledge of MLOps.
- Current knowledge across the breadth of Databricks product and platform features.
- Familiarity with optimizations for performance and scalability.
- Databricks certification is an added advantage.