Role Overview
RAK Bank is seeking a Databricks Engineer to build a scalable data ingestion and streaming platform. The role requires designing and implementing robust data pipelines capable of handling CDC events from various sources while ensuring data quality, security, and exactly-once processing semantics.
Responsibilities
- Build a scalable data ingestion and streaming platform using Confluent connectors, Databricks Auto Loader, and Delta Lake.
- Design Spark Structured Streaming pipelines to ingest CDC events from sources like SQL Server and Oracle.
- Ensure data integrity through deduplication, schema evolution, and exactly-once semantics.
- Implement monitoring solutions using Prometheus and Grafana.
- Scale ingestion capabilities using config-driven frameworks like Airflow or Delta Live Tables.
- Collaborate with cross-functional teams to enforce data quality and security standards.
Requirements
- 5–8 years of experience with big data frameworks.
- Hands-on expertise in Kafka or Azure Event Hub.
- Proficiency in programming languages: Python, Scala, or Java.
- Deep understanding of Lakehouse architectures.
- Experience with DevOps practices using CI/CD pipelines.
- Preferred: Experience with event-driven architectures, microservices integration, and ingestion frameworks like NiFi or machine learning pipelines on Spark.
- Required Skill: Experience with Data-Azure Databricks with ADF, HDInsight, ADLS, Synapse, and Azure Native Data Services.