Role Overview
This role focuses on the administration, configuration, and maintenance of critical observability and automation tools to ensure platform reliability and performance.
Responsibilities
- Administration and configuration of PagerDuty including API based integration with other tools including Service Now, Grafana, Zoom etc.
- Design, implement, and maintain observability tools for logging, alerting, and monitoring including Grafana.
- Automate infrastructure provisioning and configuration using Terraform and Ansible.
- Build and manage CI/CD pipelines using Jenkins and GitLab.
- Ensure platform reliability and performance through proactive monitoring and incident response.
- Support compliance remediation and secure cloud configurations.
- Participate in a global 24/7 support model, including on-call rotations.
- Develop self-healing workflows and automate known issue remediation.
Requirements
- 8+ years in DevOps, cloud infrastructure, or systems engineering.
- Experience with Incident management and notification.
- Experience with observability and performance monitoring tools.
- Strong experience with Tomcat and Java Spring boot.
- Proficiency in Terraform, Ansible, Jenkins, GitLab CI.
- Expert-level query writing using PromQL (Prometheus), LogQL (Loki), SQL (PostgreSQL).
- Solid Linux administration and scripting skills (e.g., Python).
- Experience with Java and Angular.
- Familiarity with disaster recovery and compliance frameworks.
- Advanced-level Certifications in PagerDuty, AWS cloud are plus.
- Bachelor’s degree in Computer Science, Engineering, or related field.
- Proven track record managing production-grade systems and automation pipelines.