Role Overview
We are looking for an experienced Data Engineering Developer (SDE2) to design and build scalable, high-performance data solutions. The ideal candidate will have strong expertise in Apache Spark, Databricks, and Scala, with a passion for building reliable data pipelines and ensuring data quality at scale.
Responsibilities
- Design, develop, and maintain scalable and reliable ETL pipelines using Apache Spark and Databricks.
- Write high-performance, production-grade Scala code for data processing and transformation.
- Implement and manage observability, monitoring, and alerting frameworks to ensure system health and data integrity.
- Develop and maintain CI/CD pipelines for data applications and infrastructure using tools such as GitHub Actions or Azure DevOps.
- Ensure adherence to data governance, data quality, and security best practices across all data systems.
- Collaborate closely with cross-functional teams including platform, analytics, and product teams to deliver robust data solutions.
- Continuously improve system performance, scalability, and operational excellence through best engineering practices.
- Stay updated with emerging technologies, including advancements in Generative AI (GenAI), and evaluate their applicability to data platforms.
Requirements
- Strong hands-on expertise in Apache Spark and Databricks.
- Proficiency in Scala, with a solid understanding of functional programming concepts.
- Experience in building and maintaining scalable ETL/data pipelines.
- Hands-on experience with CI/CD tools and practices (e.g., GitHub Actions, Azure DevOps).
- Experience in implementing monitoring, logging, and alerting systems for distributed data platforms.
- Good understanding of data governance, security, and data quality principles.
- Strong analytical, problem-solving, and debugging skills.
- Ability to work effectively in a collaborative, Agile environment.
Nice to Have
- Exposure to Generative AI (GenAI) or AI/ML data pipelines.
- Experience working with cloud platforms (Azure/AWS/GCP).
- Familiarity with modern data lake or lakehouse architectures.
- Understanding of containerization and orchestration tools is a plus.
Skills
- Apache Spark
- Databricks
- Scala
- GitHub Actions
- Azure DevOps