Role Overview
GradeLab is building the AI operating system for assessments, helping schools and universities automate handwritten exam evaluation using OCR, Large Language Models, and intelligent grading pipelines. We are seeking a QA Engineer to validate AI behavior, grading accuracy, and system reliability across a fast-moving production platform.
Responsibilities
- AI evaluation and benchmarking: Build golden datasets and measure grading accuracy against human evaluators.
- AI observability: Use Langfuse and evaluation frameworks to monitor accuracy, rubric compliance, cost, and latency.
- OCR validation: Stress-test pipelines against poor handwriting, low-resolution scans, and complex document quirks.
- Grading validation: Test partial marking, step-wise marking, and score calculation consistency.
- Prompt evaluation and AI safety: Run adversarial testing including prompt injection and jailbreak attempts.
- Edge case testing: Test real-world scenarios like crossed-out answers, missing pages, or torn sheets.
- API, load, and release testing: Validate backend APIs and run performance tests under production-scale load.
- Bug investigation: Write actionable reports with reproduction steps and Langfuse trace IDs.
Requirements
- Strong manual testing experience with excellent analytical and debugging instincts.
- Hands-on API testing experience and solid SQL fundamentals.
- Strong documentation skills and real attention to detail.
- Comfort working with bug tracking systems and thinking adversarially.
Nice to Have
- Experience with Langfuse, Promptfoo, OpenAI Evals, Playwright, Cypress, Selenium, k6, Locust, JMeter, Postman, Bruno, Docker, Git, GitHub Actions, Sentry, Grafana, OpenTelemetry, OCR systems, or EdTech/SaaS background.
Skills
- Langfuse
- API Testing
- SQL
- OCR
- LLM Evaluation
Benefits
- Competitive salary and ESOPs for high performers.
- Flexible schedule and work from home options.
- Real ownership and direct work with founders.