Role Overview
We are looking for a Cloud/SRE professional to maintain platform reliability, performance, monitoring, scalability and incident response.
Responsibilities
- Define and monitor SLOs, error budgets and uptime targets.
- Build monitoring dashboards and alerts.
- Handle production incidents and root-cause analysis.
- Plan autoscaling and capacity requirements.
- Perform load testing and reliability improvements.
- Optimise cloud infrastructure costs.
- Improve system resilience and fault tolerance.
- Maintain incident runbooks.
- Work with backend teams on performance issues.
- Conduct disaster recovery drills.
Requirements
- Experience in SRE, cloud operations or reliability engineering.
- Strong AWS/GCP/Azure knowledge.
- Monitoring and observability experience.
- Real-world incident management/on-call experience.
- Python, Go or Bash scripting.
- Understanding of distributed systems and caching.
- Linux troubleshooting and performance tuning.
Nice to Have
- Kubernetes.
- Terraform/Pulumi.
- k6, Locust or JMeter.
- Cloud architecture certification.
Skills
- AWS
- Python
- Kubernetes
- Terraform
- Linux