Role Overview
As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. You will curate code examples, provide precise solutions, and refine AI-generated code for efficiency, scalability, and reliability.
Responsibilities
- Working on AI model training initiatives by curating code examples, building solutions, and correcting code primarily in Java.
- Evaluate and refine AI-generated code to ensure that it is efficient, scalable, and reliable.
- Collaborate with cross-functional teams to enhance AI-driven coding solutions against industry performance benchmarks.
- Build agents that can verify the quality of the code and identify error patterns.
- Hypothesize on steps in the software engineering cycle and evaluate model capabilities on them.
- Design verification mechanisms that can automatically verify a solution to a software engineering task.
Requirements
- Several years of software engineering experience.
- Strong expertise in building full-stack applications and deploying scalable, production-grade software.
- Deep understanding of software architecture, design, development, debugging, and code quality/review assessment.
- Excellent oral and written communication skills for clear, structured evaluation rationales.
- Must be based in the US, Canada, or WEU countries.
Skills
- Java
- Python
- React.js
- Software Architecture
- Large Language Models
Engagement Details
- Commitment: flexible engagement, minimum 10 hrs/week, up to 40 hrs/week.
- Type: Contractor.
- Duration: 1 month (potential extensions).