Role Overview
Sarvam’s Work Agents team builds the harness for developing, evaluating, and serving autonomous agents at scale. This role sits across both sides of infrastructure, building tooling, pipelines, and automation to make the harness self-service and reproducible, while operating day-to-day to keep CI/CD, environments, networks, and deployments running smoothly. It is a mix of engineering and operations, focusing on making agent work shippable, observable, and operable at scale.
Responsibilities
- Kubernetes platform operations including upgrades, node lifecycle, RBAC, and troubleshooting.
- CI/CD and release engineering, owning pipelines from commit to production with rollout patterns like canary and blue-green.
- Cloud infrastructure provisioning and maintenance via Terraform or Crossplane on AWS and Azure.
- Managing platform-level networking, CNI configuration, ingress, and secure connectivity via VPN tunnels.
- Building automation and internal tooling such as CLIs and operators to reduce toil.
- Maintaining observability pipelines including metrics, logging, and tracing.
Requirements
- 3+ years in DevOps, SRE, or platform engineering.
- Fluency in Kubernetes, including the scheduler, API machinery, Helm, and controllers.
- Proficiency in AWS (VPC, EC2, EKS, IAM) and Azure.
- Deep understanding of networking (CNI, ingress, service routing, VPN/tunneling).
- Experience with CI/CD tools like GitLab CI, GitHub Actions, Argo CD, or Flux.
- Strong software engineering fundamentals in Python or Go.
- Experience with observability tools like Prometheus, Grafana, ELK, or Loki.
Skills
- Kubernetes
- AWS
- Terraform
- Python
- CI/CD
Bonus Points
- Experience with on-premise infrastructure and air-gapped environments.
- Experience with multi-cluster or multi-tenant Kubernetes in production.
- Open-source contributions to infrastructure tooling.