Role Overview
Nerdience Technologies Pvt. Ltd. is looking for a passionate Site Reliability Engineer (SRE) with 2 years of hands-on experience. You will work closely with development teams to build reliable, scalable, and secure platforms while ensuring high availability, performance, and operational excellence using modern cloud-native technologies.
Responsibilities
- Design, implement, and maintain highly available cloud infrastructure on AWS, Azure, or Google Cloud.
- Manage and administer Kubernetes clusters in production environments.
- Build, optimize, and maintain CI/CD pipelines using GitHub Actions, GitLab CI, or Jenkins.
- Automate infrastructure provisioning using Terraform, Ansible, or similar Infrastructure as Code (IaC) tools.
- Monitor applications and infrastructure using Prometheus, Grafana, ELK Stack, or Datadog.
- Investigate production incidents, perform root cause analysis (RCA), and implement preventive measures.
- Develop automation scripts using Python, Bash, or Go to eliminate repetitive operational tasks.
- Implement backup, disaster recovery, and high-availability strategies.
Requirements
- 2 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
- Strong knowledge of Linux system administration.
- Experience with AWS, Microsoft Azure, or Google Cloud Platform.
- Hands-on experience with Docker and Kubernetes.
- Experience building and maintaining CI/CD pipelines.
- Good understanding of Infrastructure as Code using Terraform.
- Experience with monitoring and observability tools such as Prometheus, Grafana, ELK Stack, or Datadog.
- Experience with scripting using Bash or Python.
Skills
- Kubernetes
- Terraform
- AWS
- Python
- Prometheus
Benefits
- Flexible schedule
- Paid sick time
- Competitive salary and benefits