Role Overview
We are looking for an LLM / Agentic Evaluation Rig Engineer to build the system that decides whether our AI output is good enough to ship. Because our commentary sits next to externally reported financials, we cannot rely on vibes — grounding, faithfulness, and hallucination have to be measured, tracked, and gated before anything reaches a customer. You own the evaluation infrastructure: the datasets, the scorers, the harnesses, and the CI gates that hold the AI and agentic layers to a hard quality bar.
Responsibilities
- Build and curate evaluation datasets, including adversarial and edge-case sets with ground-truth labels
- Build scorers for grounding, faithfulness, hallucination, factual consistency, and structured-output validity
- Combine rule-based checks, reference-based metrics, and LLM-as-judge where appropriate
- Build harnesses that run evaluations reproducibly across model, prompt, and agent versions
- Wire evaluation into CI so grounding / faithfulness regressions block releases
- Evaluate multi-step / agentic flows — routing, tool-use, verification, confirmation
- Build trace capture and step-level scoring for agent runs
- Partner with the Staff AI Engineer and QA to integrate AI evaluation into the broader release process
Requirements
- 4+ years in software / ML engineering, with hands-on work building LLM evaluation or quality tooling
- Real understanding of grounding, faithfulness, and hallucination — and how to measure them rigorously
- Strong Python and solid engineering practices (reproducibility, CI/CD)
- Comfort designing evaluation for non-deterministic systems without producing flaky or meaningless metrics
- Familiarity with LLM eval frameworks and LLM-as-judge patterns
Nice to Have
- Experience evaluating agentic / multi-step LLM systems
- Familiarity with RAG, structured output, and managed LLMs in-VPC
- FinTech / financial-services domain or other high-stakes, correctness-critical AI
- Background in statistics or measurement / metrics design
Skills
- Python
- LangGraph
- GitHub Actions
- Ragas
- LangSmith