Role Overview
We are looking for an AI Platform Engineer to build, deploy, optimize, and manage production AI inference systems. The ideal candidate has experience deploying machine learning, computer vision, or LLM models at scale and understands GPU acceleration, inference optimization, and cloud-native AI infrastructure. This role bridges Machine Learning and Infrastructure Engineering. You will work closely with AI engineers to deploy models, optimize inference performance, automate model delivery, and build scalable AI serving platforms.
Responsibilities
- Deploy and manage production AI models including Computer Vision, NLP, and Large Language Models (LLMs)
- Build scalable inference services for real-time and batch workloads
- Deploy and manage NVIDIA Triton Inference Server
- Package models using ONNX, TensorRT, TorchScript, or other optimized formats
- Optimize inference latency, throughput, and GPU utilization
- Build and maintain GPU-enabled Kubernetes clusters
- Configure GPU scheduling and resource allocation
- Deploy inference services using Docker, Kubernetes, and Helm
- Manage model repositories and versioning
- Implement scalable APIs for AI inference
- Profile inference latency and throughput and optimize GPU memory usage
- Benchmark different model formats (PyTorch, ONNX, TensorRT)
- Configure dynamic batching and concurrent inference
- Build CI/CD pipelines for AI models and automate model deployment
- Monitor production AI systems including GPU utilization, latency, and throughput using Prometheus and Grafana
Requirements
- Experience with Triton Inference Server, TensorRT, or ONNX Runtime
- Proficiency in PyTorch or TensorFlow deployment
- Strong knowledge of Docker, Kubernetes, and Helm
- Understanding of CUDA fundamentals and GPU scheduling
- Proficiency in Python, Bash, and Linux
- Experience building REST APIs with FastAPI or similar frameworks
Skills
- Kubernetes
- PyTorch
- Docker
- TensorRT
- Python
Nice to Have
- Experience deploying LLMs using vLLM, TensorRT-LLM, or NVIDIA NIM
- Experience with Ray Serve, KServe, or Seldon Core
- Experience with Kubeflow or MLflow
- Experience with model quantization (FP16/INT8)