Role Overview
We are building a new 24x7 NOC / Production Support team at Lenskart to monitor and support our production infrastructure, applications and services. The role involves production monitoring, incident detection, L1 troubleshooting, alert management, escalation and coordination with DevOps, Infrastructure, Cloud, Network and Application teams. This is a hands-on operational role and not limited to monitoring or ticket logging.
Responsibilities
- Monitor production infrastructure, applications, services and dashboards in a 24x7 environment.
- Validate alerts, identify genuine incidents and perform first-level troubleshooting.
- Analyse logs, metrics and dashboards to identify the impacted component.
- Troubleshoot common issues across Linux, applications, cloud infrastructure and networks using SOPs/runbooks.
- Troubleshoot basic issues related to CPU, memory, disk, services, connectivity, DNS, HTTP/HTTPS, APIs and load balancers.
- Monitor cloud infrastructure, preferably AWS, and perform basic health checks.
- Support production environments running on Docker/Kubernetes and coordinate with DevOps teams when required.
- Manage incidents, initiate bridge calls for critical issues and follow defined escalation matrices.
- Ensure timely escalation based on SLA and incident priority.
- Maintain shift logs, incident timelines and effective shift handovers.
- Identify recurring incidents, alert noise and opportunities for automation/process improvement.
Requirements
- 1–4 years of experience in NOC, Production Support, Infrastructure Support, Cloud Operations or L1 DevOps.
- Good understanding of Linux and hands-on exposure to monitoring/observability tools.
- Ability to read logs and perform basic technical troubleshooting.
- Basic understanding of networking, DNS, HTTP/HTTPS, APIs and load balancers.
- Basic understanding of AWS/cloud infrastructure.
- Basic knowledge of Docker and Kubernetes.
- Understanding of incident management, escalation and SLA-based support.
- Familiarity with Jira, ServiceNow or similar ticketing platforms.
- Good communication and coordination skills.
- Comfortable working in a 24x7 rotational shift environment.
Skills
- Linux
- AWS
- Docker
- Kubernetes
- Jira
Nice to Have
- AWS – EC2, VPC, IAM, S3, ELB, Auto Scaling, CloudWatch
- Grafana / Prometheus / Datadog / ELK / Kibana
- Jenkins / CI-CD
- Docker / Kubernetes
- Bash or basic scripting
- Ansible
- Production support experience in a high-availability environment
- Basic understanding of SRE, RCA, SLI/SLO