Role Overview
We are looking for NOC / Observability Engineers to monitor and maintain the availability, performance, and reliability of infrastructure, applications, and services in a 24/7 operational environment. The ideal candidate should have strong hands-on experience with Linux, monitoring and alerting tools, networking, cloud fundamentals, and basic scripting, along with the ability to troubleshoot incidents and escalate issues effectively.
Responsibilities
- Monitor infrastructure, applications, servers, and services 24/7.
- Monitor and respond to alerts related to system availability, performance, and service health.
- Perform initial troubleshooting and identify the probable cause of incidents.
- Analyze system and application logs to identify errors and performance issues.
- Monitor metrics, dashboards, alerts, and service health using observability tools.
- Handle incidents according to defined escalation and response procedures.
- Escalate critical or unresolved incidents to DevOps, Cloud, or Engineering teams.
- Perform basic root-cause analysis and document incident findings.
- Monitor Kubernetes/GKE environments and identify basic pod, node, and service issues.
- Maintain awareness of infrastructure health, capacity, and performance.
- Follow incident management, monitoring, and operational procedures.
- Participate in rotational 24/7 shifts, including nights, weekends, and holidays when required.
Requirements
- Strong knowledge of Linux administration and troubleshooting.
- Fundamentals of GCP, preferably hands-on exposure to GCP services.
- Basic understanding of Kubernetes / GKE.
- Experience with monitoring and observability tools such as: Grafana, Prometheus, Nagios or similar monitoring platforms.
- Understanding of TCP/IP, DNS, HTTP/HTTPS, networking, routing, and basic troubleshooting.
- Basic scripting knowledge in Bash and/or Python.
- Experience with monitoring, logging, metrics, dashboards, and alerting.
- Ability to analyze application and system logs.
- Understanding of incident management and escalation processes.
Skills
- Linux
- GCP
- Kubernetes
- Grafana
- Prometheus