Role Overview
You'll be a core individual contributor on the Fault Management team, building near-real-time streaming pipelines and intelligent fault-processing capabilities within our AIOps platform. You own module delivery end-to-end — from low-level design through production support — and work with platform architects and PMs to ship high-impact features.
Responsibilities
- Own one or more fault-management modules end-to-end and write LLD documents.
- Build event-driven pipelines with Kafka Streams and Kafka Connect for fault ingestion and enrichment.
- Implement Apache Spark batch and micro-batch jobs for large-scale fault analytics.
- Build cloud-native microservices using Java and Spring Boot.
- Use LLM-assisted coding and AI pair-programming tools to accelerate delivery.
- Diagnose issues across Kafka brokers, Streams topologies, Spark executors, and Spring Boot services.
- Maintain observability through metrics, tracing, and structured logging.
Requirements
- 3–8 years of hands-on backend / data engineering experience.
- Production Kafka systems experience including topics, partitioning, and consumer groups.
- Proficiency in Kafka Streams DSL and Processor API.
- Experience with Kafka Connect source and sink connectors.
- Expertise in Apache Spark (Structured Streaming, DataFrames, Spark SQL).
- Strong command of Java 11+ and Spring Boot 3.x.
- Deep debugging skills across multi-threaded, distributed JVM systems.
- Knowledge of distributed systems principles like CAP theorem and eventual consistency.
Skills
- Java
- Spring Boot
- Apache Kafka
- Apache Spark
- Kafka Streams
Nice to Have
- AIOps or observability platform experience.
- Agentic development workflows or LLM-integrated toolchains.
- Telecom fault management standards or OSS/BSS domain knowledge.
- Kubernetes, service meshes, or cloud-native deployments (AWS / Azure / GCP).