Role Overview
The Site Reliability Engineer – Incident management has the responsibility of monitoring, maintaining and managing entire Qualys infrastructure and services installed at different datacenters. When there is any malfunction in Product/Services, the technician monitors, troubleshoots, repairs and gets the service/system back up as quickly as possible to ensure maximum possible service availability and performance.
Responsibilities
- Maintain highly available and scalable applications and services.
- Monitor application and infrastructure health using observability tools.
- Respond to incidents, troubleshoot production issues, and perform root cause analysis.
- Participate in on-call rotations and incident response processes.
- Track and improve service reliability, latency, performance, and efficiency.
- Create automation scripts and tooling to reduce manual operational effort.
- Track and document all issues and resolutions in detail using ticketing and documentation tools.
- Escalate issues to management, IT resources, or 3rd party vendors as needed.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience).
- Experience in Site Reliability Engineering, NOC operations, or Cloud Infrastructure roles.
- Knowledge of Linux/Unix systems administration.
- Experience with cloud platforms such as AWS, OCI, or Google Cloud Platform.
- Proficiency in Python, Go, Bash, or Java.
- Familiarity with Docker and Kubernetes.
- Familiarity with Terraform, CloudFormation, or other IaC tools.
- Experience with monitoring and logging tools such as Prometheus, Grafana, AppDynamics, ELK, or Splunk.
- Understanding of networking, security, and distributed systems concepts.
Skills
- AWS
- Kubernetes
- Python
- Terraform
- Prometheus