Role Overview
We are looking for an experienced Application SRE with 4+ years of experience to ensure the high availability, scalability, and performance of our production applications. The ideal candidate will bridge the gap between application development and operations—leveraging expertise in Application Performance Monitoring (APM), Microservices Architecture, CI/CD pipelines, Cloud Platforms, and incident management. This is a hybrid role with rotational shifts, including night shifts. Immediate joiners are preferred.
Responsibilities
- Monitor, maintain, and optimize application health, performance, and availability across production environments.
- Troubleshoot and resolve complex L2/L3 application incidents, performance bottlenecks, and service disruptions.
- Perform deep-dive Root Cause Analysis (RCA) for application failures and drive permanent remediation.
- Track and improve application-level reliability metrics (SLOs, SLIs, error budgets, and MTTR).
- Manage incident, problem, and release management activities for software deployments.
- Support CI/CD pipelines, blue-green/canary deployments, and application rollout activities.
- Collaborate closely with Application Development, DevOps, and QA teams to fix bugs and enhance application resilience.
- Automate repetitive operational tasks and build automated self-healing scripts using Python or Bash.
- Create and maintain operational documentation, runbooks, and application architecture diagrams.
Requirements
- Bachelor’s degree in Computer Science, IT, Software Engineering, or a related field.
- 4+ years of experience in Application SRE, Application Support Engineering, or Production Support.
- Strong understanding of Microservices, REST APIs, and Web Applications.
- Hands-on experience with APM tools (e.g., Datadog, Dynatrace, New Relic, AppDynamics) and logging stacks (ELK/EFK, Grafana Loki).
- Practical experience with Kubernetes and Docker.
- Hands-on experience with AWS or other major cloud platforms.
- Experience managing CI/CD pipelines (Jenkins, GitHub Actions, GitLab CI) and Git workflows.
- Strong scripting skills in Python, Bash, or Shell.
- Willingness to work in a hybrid mode and rotational shifts, including night shifts.
Nice To Have
- Experience with Caching & Messaging systems (Redis, Kafka, RabbitMQ) and Database performance tuning (PostgreSQL, MySQL, MongoDB).
- Solid knowledge of SLO, SLA, and SLI concepts.
- Familiarity with Infrastructure as Code (Terraform or Ansible).
- ITIL Foundation certification.
- AWS / Azure / GCP certifications.
Skills
- Kubernetes
- AWS
- Python
- Datadog
- Docker