Role Overview
As an AI Ops Engineer, you will play a critical role in the design, development, and deployment of an AI platform. You will work collaboratively with cross-functional teams, focusing on MLOps & LLMops pipelines, machine learning standardization, GPU and cloud platform setup, and data integrations to ensure scalability, reliability, and efficiency.
Responsibilities
- Designing and delivering a hybrid AI Platform.
- Maintaining scalable, efficient, and reliable data and AI pipelines.
- Maintaining MLOps and LLMOps pipelines across Databricks and on-prem clusters.
- Developing and implementing AI and ML models to automate routine tasks.
- Monitoring system performance and implementing a monitoring strategy across central observability.
- Ensuring AI platform delivery meets security, risk, and operational SLA requirements.
- Collaborating with data scientists, engineers, and governance teams to align with data strategies.
- Integrating AIOps solutions with existing services like ServiceNow and ELK.
- Performing code reviews to minimize technical debt.
Requirements
- 4-6 years of experience in engineering or operation disciplines of AI Ops and ML Ops.
- Hands-on experience with GPU computing and optimization for AI workloads.
- Proven track record of delivering complex AI projects.
- Strong proficiency in Databricks, CI/CD (e.g., Azure DevOps), and machine learning frameworks.
- Experience with MLOps tools such as MLflow or Kubeflow.
- Knowledge of cloud platforms like Azure or AWS and infrastructure as code using Terraform.
- Engineering Degree or equivalent.
Skills
- Python
- Databricks
- MLOps
- Terraform
- PyTorch