Role Overview
HackerRank is redefining technical talent assessment for the era of AI-assisted coding. As software engineering shifts from manual coding to AI orchestration, we need to solve the unsolved problem of how to fairly and rigorously measure human skill when AI assistance is ambient. This role focuses on building the methodology, infrastructure, and LLM-powered systems that define what meaningful skill evaluation looks like in an agentic world.
Responsibilities
- Build LLM-powered evaluation pipelines that assess AI usage skills consistently, fairly, and at production scale.
- Own the evaluation methodology end to end, including rubrics, model application, measurement, and bias auditing.
- Design and run experiments to determine effective evaluation methodologies in uncharted territory.
- Build RAG pipelines and fine-tuning workflows to ensure evaluation models adhere to set rules.
- Define benchmarking infrastructure to monitor evaluation quality and catch regressions.
- Translate complex model behavior into understandable outcomes for product managers and customers.
Requirements
- Proven experience shipping LLM-powered systems in production with high reliability constraints.
- A rigorous approach to measurement and a research-oriented mindset for inventing new methodologies.
- Systems-level thinking encompassing data pipelines, models, and serving layers.
- Ability to communicate ML judgment and technical concepts to non-technical stakeholders.
Nice to Have
- Experience building evaluation frameworks for generative or conversational AI systems.
- Background in educational assessment, psychometrics, or human-in-the-loop evaluation.
- Publications or open-source contributions in LLM evaluation, benchmarking, or alignment.
- Experience working at the interface of research and product.
Skills
- LLM
- RAG
- Machine Learning
- Python
- MLOps