Role Overview
We are seeking a skilled Azure Databricks Developer to design, develop, and optimize big data pipelines using Databricks on Azure. The ideal candidate will have strong expertise in PySpark, Azure Data Lake, and data engineering best practices in a cloud environment.
Responsibilities
- Design and implement ETL/ELT pipelines using Azure Databricks and PySpark.
- Work with structured and unstructured data from diverse sources (e.g., ADLS Gen2, SQL DBs, APIs).
- Optimize Spark jobs for performance and cost-efficiency.
- Collaborate with data analysts, architects, and business stakeholders to understand data needs.
- Develop reusable code components and automate workflows using Azure Data Factory (ADF).
- Implement data quality checks, logging, and monitoring.
- Participate in code reviews and adhere to software engineering best practices.
Requirements
- 5+ years of experience in Apache Spark / PySpark.
- 5+ years working with Azure Databricks and Azure Data Services (ADLS Gen2, ADF, Synapse).
- Strong understanding of data warehousing, ETL, and data lake architectures.
- Proficiency in Python and SQL.
- Experience with Git, CI/CD tools, and version control practices.
- Bachelor's degree in Computer Science or related field.
Nice to Have
- Experience with Delta Lake and data lake architecture.
- Exposure to streaming technologies (Azure Event Hub, Kafka, Spark Streaming).
- Knowledge of CI/CD pipelines and DevOps practices.
- Experience with orchestration tools like Airflow.
- Familiarity with Power BI or other visualization tools.
Skills
- Azure Databricks
- PySpark
- Azure Data Factory
- Python
- SQL