Role Overview
The AIOps LLM Platform Engineer works as part of a cross functional team of platform architects, solution architects, engineers, and business stakeholders to apply industry best practices for designing, hosting, automating, securing, and operating enterprise AIOps and LLM platforms.
Responsibilities
- Participate in a global team of AIOps Architects, Platform Engineers, Automation Engineers, and IT Operations teams.
- Automate the provisioning, configuration, and lifecycle management of LLM runtimes, embedding pipelines, vector databases, and knowledge ingestion services.
- Deploy and manage LLM configurations, embedding models, prompt templates, runtime parameters, patching, and platform tools across environments using Ansible playbooks.
- Integrate ITSM and Service Catalog platforms (ServiceNow / First) with Terraform and Ansible to automate AI platform service requests and approvals.
- Provision and operate LLM, vector, and ingestion workloads on Kubernetes platforms (e.g., Tanzu, EKS, or equivalent) using IaC practices.
- Utilize monitoring and observability tools to track performance, availability, and health of AI platforms, including SolarWinds, Prometheus, Grafana, CloudWatch, and custom scripts.
- Support enterprise‑grade backup, recovery, retention, and purge strategies for vector data, metadata, and platform configurations.
- Collaborate with IT, Security, and Business partners to continuously improve platform reliability, scalability, performance, governance, and audit readiness.
Requirements
- Demonstrate strong ownership, willingness to learn, and continuous improvement while supporting AI, automation, and AIOps technologies in an enterprise environment.
Skills
- Kubernetes
- Ansible
- Terraform
- Python
- Prometheus