Role Overview
We are looking for an AI DevOps Engineer with expertise in AI/ML, LLMs, observability, and automation to develop intelligent solutions that improve system reliability and operational efficiency. The role involves building AI-driven capabilities for monitoring, anomaly detection, predictive analytics, incident response, log analytics, and root cause analysis across enterprise applications, databases, and cloud infrastructure.
Responsibilities
- Design, develop, and deploy AI-powered solutions to improve system reliability, automate incident response, and optimize enterprise operations.
- Build AI-driven capabilities for anomaly detection, event correlation, predictive analytics, root cause analysis, and intelligent operational insights using AI/ML, LLMs, and AI Agents.
- Design and implement AI-driven automation for infrastructure monitoring, database operations, performance optimization, capacity planning, and operational recommendations.
- Design and optimize enterprise observability solutions using metrics, logs, traces, and events to provide end-to-end visibility across applications and infrastructure.
- Develop and deploy AI-powered log analytics and operational intelligence solutions using Machine Learning, LLMs, RAG, and AI frameworks.
- Develop, deploy, and maintain AI/ML models for observability, predictive operations, and intelligent IT automation.
Requirements
- Bachelor’s degree in computer science, Information Technology, Artificial Intelligence, Data Science, or a related field.
- 3-5 years of experience in AIOps, DevOps or related enterprise operations roles.
- Strong programming skills in Python, SQL, Shell scripting, REST APIs, and automation frameworks.
- Hands-on experience developing and deploying AI/ML solutions for AIOps, observability, automation, or intelligent IT operations.
- Experience building or deploying ML models for anomaly detection, predictive analytics, log analysis, event correlation, and operational intelligence.
- Experience with LLMs, RAG, AI Agents, MCP, LangChain, or similar AI frameworks.
- Knowledge of AI model deployment, MLOps, model lifecycle management, and operational AI pipelines.
- Strong experience with enterprise observability platforms such as Grafana, Prometheus, OpenTelemetry, Elastic/OpenSearch, Splunk, Datadog, or Dynatrace.
- Exposure with Kubernetes, Docker, container platforms, cloud infrastructure (AWS preferred), and on-premises environments.
- Exposure to enterprise databases including Oracle, PostgreSQL, MySQL, and MongoDB.
- Experience with Infrastructure as Code (Terraform, Ansible, or equivalent), CI/CD pipelines, and automation tools.