Role Overview
We are seeking a specialized QA Engineer to design and execute test strategies for AI/ML models, LLM-based applications, and data pipelines. You will be responsible for ensuring the accuracy, consistency, and ethical compliance of AI-driven features including RAG pipelines and chatbots.
Responsibilities
- Design and execute test strategies for AI/ML models, LLM-based applications, and data pipelines.
- Develop automated test frameworks for model validation, regression testing, and performance benchmarking.
- Evaluate model outputs for accuracy, consistency, relevance, hallucination, and bias.
- Test RAG pipelines, chatbots, and recommendation systems.
- Collaborate with data scientists and ML engineers to define acceptance criteria.
- Build and maintain evaluation datasets and adversarial test cases.
- Monitor models in production for drift and degradation.
- Validate data quality and feature stores.
- Document defects and failure patterns specific to AI behavior.
- Ensure systems meet ethical, fairness, and compliance standards.
Requirements
- 3–6 years of overall QA experience, with 1–2 years specialized in AI/ML QA.
- Bachelor's or Master's degree in Computer Science or a related field.
- Strong proficiency in Python for test automation and data analysis.
- Familiarity with LLM evaluation frameworks like RAGAS, DeepEval, Promptfoo, or LangSmith.
- Hands-on experience with Pytest, Selenium, or Postman.
- Solid understanding of the ML lifecycle.
- Knowledge of data quality tools like Great Expectations or dbt.
Nice to Have
- Experience with prompt engineering and red-teaming LLMs.
- Familiarity with MLOps platforms such as MLflow, SageMaker, or Vertex AI.
- Knowledge of vector databases and embedding quality evaluation.
- Understanding of AI safety and responsible AI principles.
- Experience with A/B testing and shadow deployment.
Skills
- Python
- Pytest
- LLM Evaluation
- MLOps
- RAG