Role Overview
We are seeking a skilled Databricks Developer to design, build, and optimize big data solutions on the Databricks Lakehouse Platform. The ideal candidate will work with large-scale data pipelines, collaborate with data engineering and analytics teams, and help drive data-driven decision-making across the organization.
Responsibilities
- Design, develop, and maintain ETL/ELT pipelines using Databricks, Apache Spark, and Delta Lake
- Build and optimize data workflows using PySpark, Scala, or SQL within Databricks notebooks
- Develop and manage Databricks jobs, clusters, and workflows for batch and streaming data processing
- Implement Delta Lake architecture (bronze/silver/gold layers) for data lakehouse solutions
- Integrate Databricks with cloud platforms (Azure, AWS, or GCP) and other data sources (ADLS, S3, Kafka, etc.)
- Optimize Spark jobs for performance, scalability, and cost efficiency
- Collaborate with data scientists to support machine learning model development and deployment (MLflow)
- Implement data quality checks, validation, and monitoring processes
- Manage version control, CI/CD pipelines for Databricks notebooks and workflows (e.g., using Databricks Repos, Git, Azure DevOps)
- Ensure data security, governance, and compliance using Unity Catalog or similar tools
- Troubleshoot and resolve issues related to data pipelines and cluster performance
- Document technical designs, workflows, and best practices
Requirements
- Bachelor's degree in Computer Science, Information Technology, or related field
- 2–5+ years of experience working with Databricks and Apache Spark
- Strong proficiency in Python (PySpark), Scala, or SQL
- Hands-on experience with Delta Lake, Delta Live Tables, and Databricks Workflows
- Experience with cloud platforms (Azure Databricks, AWS Databricks, or GCP)
- Knowledge of data warehousing concepts, ETL/ELT design, and data modeling
- Familiarity with orchestration tools (Airflow, Azure Data Factory, etc.)
- Experience with version control systems (Git) and CI/CD practices
- Understanding of data governance, security, and Unity Catalog
- Strong analytical and problem-solving skills
- Good communication and collaboration abilities
Nice to Have
- Databricks Certified Data Engineer Associate/Professional certification
- Experience with MLflow, MLOps, or machine learning pipelines
- Knowledge of streaming technologies (Kafka, Structured Streaming)
- Experience with infrastructure-as-code tools (Terraform)
- Exposure to BI tools like Power BI, Tableau, or Looker
Benefits
- Health insurance
- Provident Fund
Skills
- Databricks
- Apache Spark
- PySpark
- Delta Lake
- SQL