Role Overview
We are looking for a highly skilled Site Reliability Engineer (SRE) to manage and scale mission-critical, production-grade distributed systems running on Google Cloud Platform (GCP). The ideal candidate will focus on reliability, automation, observability, and operational excellence while minimizing toil and improving system availability. Maintaining and improving 4 Nines of uptime to 5 Nines with engineering efforts.
Responsibilities
- Own end-to-end production systems reliability, availability, scalability, cost and performance.
- Drive measurable improvements in MTTR, MTTA, and incident response practices using automation and runbook additions.
- Participate in 24x7 on-call rotations and handle high-severity incidents.
- Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics.
- Design, deploy, and manage infrastructure on Google Cloud Platform (GCP) using GKE, Compute, Networking, IAM, and Load Balancers.
- Implement and manage infrastructure using Terraform (Infrastructure as Code).
- Deploy and manage containerized workloads using Kubernetes (GKE) and Docker.
- Manage deployments using Helm, YAML, and rollout strategies like Canary/Blue-Green.
- Build and maintain CI/CD pipelines using Jenkins with Groovy, Shell, or Python scripting.
- Implement and manage monitoring systems using Dynatrace and Grafana.
- Perform deep troubleshooting for distributed systems, microservices, Java, and Golang applications.
Requirements
- 4–8 years of relevant and progressive experience in SRE / DevOps / Cloud Engineering.
- Hands-on experience managing production-grade systems (24x7 environments).
- Deep expertise in Kubernetes (GKE) and Docker.
- Strong hands-on experience with Terraform.
- Strong troubleshooting skills in a distributed environment spread across multiple cloud environments.
Skills
- Google Cloud Platform (GCP)
- Kubernetes
- Terraform
- Python
- Jenkins