Role Overview
Design, develop, and maintain scalable backend services for Observability and AIOps platforms. This role involves developing robust applications and services to support monitoring, alerting, and anomaly detection within distributed environments.
Responsibilities
- Design and implement microservices and distributed system architectures.
- Develop and integrate APIs and backend components for monitoring and observability solutions.
- Work with Kubernetes (K8s) for deploying, managing, and troubleshooting containerized applications.
- Troubleshoot complex application issues across distributed environments and production systems.
- Analyze logs, metrics, traces, and application behavior to identify performance and reliability issues.
- Support and enhance AIOps capabilities for event management, monitoring, alerting, anomaly detection, and automation.
- Collaborate with DevOps, SRE, and platform engineering teams to improve system reliability.
- Participate in production support, incident investigation, and root-cause analysis (RCA).
- Optimize backend services for performance, scalability, availability, and reliability.
- Implement best practices for coding, testing, version control, CI/CD, and deployment automation.
Requirements
- Proven experience in designing and maintaining scalable backend services.
- Strong understanding of microservices and distributed systems.
- Experience with containerization and orchestration tools.
- Ability to perform deep-dive troubleshooting and root-cause analysis in production environments.
Skills
- Python
- Kubernetes
- Docker
- Microservices
- AIOps