Role Overview
We are seeking highly motivated Site Reliability Engineers (SREs) to ensure the availability, reliability, scalability, and performance of our SaaS production environments. The ideal candidate will have hands-on experience with cloud platforms, Kubernetes, Infrastructure as Code (IaC), automation, monitoring, and incident management.
Responsibilities
- Monitor production environments using Datadog, PagerDuty, Grafana, and Prometheus.
- Act as the first responder for alerts and incidents, acknowledging P1 alerts within 5 minutes.
- Perform root cause analysis (RCA) and contribute to post-incident reviews.
- Provision, manage, and optimize cloud infrastructure on AWS and/or GCP.
- Manage Kubernetes clusters (EKS/GKE) and related cloud-native services.
- Execute SaaS application deployments using CI/CD pipelines.
- Manage Kubernetes deployments using Helm Charts and GitOps practices.
- Develop and maintain Terraform modules and infrastructure automation.
- Implement configuration management using Ansible.
- Build automation scripts using Python and Bash to eliminate repetitive tasks.
- Implement metrics, logs, and traces using observability platforms.
- Execute and validate backup and recovery processes using Veeam, AWS Backup, or GCP Snapshots.
Requirements
- 3 to 7 years of experience in an SRE or production support role.
- Strong troubleshooting and analytical skills.
- Experience working in on-call environments and understanding SLI, SLO, and SLA concepts.
- Ability to work under pressure during critical incidents.
- Experience working in Agile, DevOps, or SRE teams.
Skills
- Kubernetes
- AWS
- Terraform
- Python
- Ansible
Benefits
- Competitive salary and benefits package.
- Culture focused on talent development with quarterly growth opportunities.
- Company-sponsored higher education and certifications.