Role Overview
We are looking for a Data Platform Engineer (IC2) to build and operate scalable, reliable, and high-performance data platforms. In this role, you will develop streaming and batch data pipelines, manage Databricks environments, and support our event-driven data infrastructure using Kafka and Kafka Connect. This is an excellent opportunity for engineers with a strong foundation in distributed data processing who want to work on modern data engineering technologies at scale.
Responsibilities
- Develop, maintain, and optimize batch and streaming data pipelines using Apache Spark and Apache Flink.
- Build scalable ETL/ELT workflows for processing large datasets.
- Manage and maintain Databricks workspaces, clusters, jobs, notebooks, and workflows.
- Monitor, troubleshoot, and optimize Databricks jobs for performance and cost efficiency.
- Work with Apache Kafka for event-driven data processing and streaming applications.
- Configure, deploy, and troubleshoot Kafka Connect connectors for integrating various data sources and sinks.
- Ensure data reliability through monitoring, alerting, and operational best practices.
- Collaborate with software engineers, platform engineers, and data consumers to design robust data solutions.
- Participate in production support, incident response, and root cause analysis.
- Contribute to automation, documentation, and continuous improvement of the data platform.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 3–4 years of experience in software engineering or data engineering.
- Strong programming skills in Java, Scala, or Python.
- Hands-on experience with Apache Spark for large-scale data processing.
- Knowledge of Apache Flink for stream processing.
- Experience working with Databricks, including cluster management and job orchestration.
- Good understanding of Apache Kafka, including topics, partitions, producers, and consumers.
- Familiarity with Kafka Connect and connector deployment.
- Understanding of distributed systems, data pipelines, and streaming architectures.
- Experience with Git and CI/CD practices.
- Strong problem-solving and debugging skills.
Skills
- Apache Spark
- Databricks
- Apache Kafka
- Apache Flink
- Python