Role Overview
We're hiring a Data Engineer in Bangalore to build the platform-level data infrastructure that powers FirstHive's Customer Data Platform across every client integration. You will focus on building connector frameworks, CDC ingestion pipelines, transformation services, and data quality tooling using Java and Spring Boot on Kafka, MongoDB, and analytical warehouse stacks like StarRocks, Snowflake, and BigQuery, deployed on multi-cloud Kubernetes.
Responsibilities
- Design and build pluggable connector frameworks over Kafka and Kafka Connect for databases, APIs, and event streams.
- Develop CDC ingestion pipelines from MongoDB and relational sources via Debezium.
- Create automated schema mapping, detection, and inference tooling.
- Build transformation layers as composable Spring Boot modules for cleaning, deduplication, and identity resolution.
- Implement data quality frameworks including profiling, validation gates, and anomaly detection.
- Design data models for StarRocks, Snowflake, and BigQuery including partitioning and clustering strategies.
- Develop a metadata layer to drive per-client schema definitions and transformation logic via configuration.
Requirements
- 4+ years of experience building framework-level data systems in production.
- Expertise in production-grade Java and Spring Boot for microservices.
- High fluency in SQL and deep knowledge of data modeling (Star schema, SCD types, event sourcing).
- Deep technical knowledge of Kafka and Kafka Connect (custom transforms, consumer group design, DLQ patterns).
- Architectural experience with at least one analytical warehouse (StarRocks, Snowflake, or BigQuery).
- Experience with MongoDB or similar document stores, including schema design and change streams.
- Experience with workflow orchestration tools like Airflow, Argo Workflows, or dbt.
Skills
- Java
- Spring Boot
- Kafka
- Snowflake
- MongoDB