Role Overview
This is a remote L1 Support position focused on AWS infrastructure and application health monitoring. The role involves incident response and operational support with no on-premise infrastructure requirements.
Responsibilities
- Acknowledge and respond to alerts from tools such as Splunk, Datadog, and incident tickets.
- Follow established runbooks and SOPs.
- Escalate unresolved issues to L2/L3 teams.
- Track MTTA and MTTR metrics.
- Perform basic AWS admin functions including EC2, S3, IDM, access control, backups, and scheduled jobs.
- Develop and maintain CI/CD pipelines and Infrastructure as Code (IaC).
- Troubleshoot infrastructure and deployment issues.
- Implement cloud and security best practices for reliability and scalability.
- Automate infrastructure provisioning and configuration management.
- Enhance monitoring and observability using Prometheus, Grafana, and ELK.
- Document processes and perform shift handover reporting.
Requirements
- Experience with cloud-based infrastructure monitoring and incident response.
- Ability to work in a remote environment.
- Knowledge of automated deployment workflows.
Skills
- AWS
- Splunk
- Datadog
- Terraform
- Prometheus