Role Overview
We are building LLM evaluation and training datasets to train LLMs to work on realistic software engineering problems. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and quality, specifically focusing on building verifiable SWE tasks based on public repository histories.
Responsibilities
- Analyze and triage GitHub issues across trending open-source libraries.
- Set up and configure code repositories, including Dockerization and environment setup.
- Evaluating unit test coverage and quality.
- Modify and run codebases locally to assess LLM performance in bug-fixing scenarios.
- Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs.
- Opportunities to lead a team of junior engineers to collaborate on projects.
Requirements
- Minimum 3+ years of overall experience.
- Strong experience with at least one of the following languages: C++.
- Proficiency with Git, Docker, and basic software pipeline setup.
- Ability to understand and navigate complex codebases.
- Comfortable running, modifying, and testing real-world projects locally.
- Experience contributing to or evaluating open-source projects is a plus.
Nice to Have
- Previous participation in LLM research or evaluation projects.
- Experience building or testing developer tools or automation agents.
Skills
- C++
- Git
- Docker
- LLM
- Software Engineering