Role Overview
We are seeking an experienced MLOps Engineer to help build and operate reliable machine learning platforms and production ML workflows. This role sits at the intersection of software engineering, machine learning, cloud infrastructure, and DevOps. The successful candidate will help data science teams move models from experimentation into production while establishing repeatable processes for deployment, monitoring, versioning, scalability, and operational support.
Responsibilities
- Build and maintain automated workflows for machine learning model development and deployment.
- Collaborate with data scientists and ML engineers to productionize trained models.
- Design CI/CD pipelines for machine learning applications, models, and supporting services.
- Containerize ML workloads using Docker and manage deployments through Kubernetes.
- Implement experiment tracking, model versioning, and lifecycle management using platforms such as MLflow or Kubeflow.
- Automate infrastructure and application provisioning using infrastructure-as-code practices.
- Develop reliable deployment processes across development, testing, staging, and production environments.
- Work with AWS, Azure, or Google Cloud to deploy and operate machine learning workloads.
- Build and maintain data and model pipelines using technologies such as Airflow.
- Configure monitoring and observability for machine learning services and infrastructure.
- Track model performance, system health, resource utilization, and operational metrics.
- Investigate failed pipelines, deployment issues, infrastructure problems, and production incidents.
- Implement appropriate logging, alerting, rollback, and recovery mechanisms.
- Maintain secure and reproducible development and deployment environments.
- Work with engineering and data teams to improve the scalability and reliability of ML platforms.
- Contribute to platform documentation, operational procedures, and engineering standards.
Requirements
- 3+ years of professional experience in MLOps, ML engineering, DevOps, or a closely related field.
- Strong Python programming and scripting skills.
- Good understanding of machine learning development and deployment workflows.
- Hands-on experience with Docker and Kubernetes.
- Experience designing and maintaining CI/CD pipelines.
- Practical experience with Jenkins, Git, or comparable DevOps tooling.
- Familiarity with MLflow, Kubeflow, or similar machine learning lifecycle platforms.
- Experience working with at least one major cloud platform: AWS, Azure, or GCP.
- Good Linux administration and troubleshooting skills.
- Understanding of infrastructure automation and configuration management.
- Experience monitoring production systems and investigating operational issues.
Skills
- Python
- Docker
- Kubernetes
- MLflow
- AWS