Role Overview
At Shaadi.com, we’re not just building a platform — we’re playing Cupid for millions across the world! Behind every successful match, swipe, and connection lies a continuous stream of data. With over 100+ million records generated daily, we rely on rock-solid data engineering to power real human stories, smart recommendations, and seamless user experiences. If you love building heavy-duty pipelines, taming data lakes, and turning raw data streams into matchmaking magic, this is your sign to join our team!
Responsibilities
- Build High-Volume Pipelines: Design, build and maintain scalable data pipelines (real-time and batch) and large-scale distributed systems that ingest, process and organize large volumes of data.
- Architect Data Lakes & Warehouses: Manage scalable data architecture across AWS (S3, Redshift, Glue) so our Data Science, Product, and Analytics teams have clean, queryable data at their fingertips.
- Stream in Real-Time: Work with streaming frameworks like Kafka to move event data instantly.
- Data Zen & Quality Control: Clean, structure, and optimize raw datasets. Monitor data drift, system performance, and maintain immaculate data reliability.
- Cross-Functional Collaboration: Partner with Product, Backend Engineering and Data Science teams to build features that directly impact business metrics and user happiness.
Requirements
- 2–4 years of hands-on experience in Python (Java or Go is a great plus!) with advanced SQL wizardry.
- Solid experience with AWS services (Redshift, S3, Glue, CloudFormation, or ECS) and modern data lake setups.
- Hands-on experience or deep understanding of the Kafka ecosystem (or AWS Kinesis).
- Proven track record of writing efficient, complex ETL/ELT jobs and understanding microservices architectures.
- Comfortable with Git version control, CI/CD pipelines, and writing optimized queries.
Nice to Have
- Familiarity with Elasticsearch, Snowplow, or Docker, and curiosity about how data feeds into ML models & recommender systems.