Role Overview
Lexsi Labs is a leading frontier AI lab focused on building aligned, interpretable, and safe superintelligent systems. This role involves building the shared harness, execution substrate, and evaluation system that powers autonomous agents across software engineering, data science, and AI research. You will develop the core agent loop, orchestration, and the measurement infrastructure required to improve agent performance.
Responsibilities
- Develop the core agent loop and execution model, including orchestration, concurrency, and sub-agent coordination.
- Manage context and state for long-running tasks, including retention and compaction.
- Build the shared tool protocol, schemas, and internal libraries.
- Create sandboxed, reproducible, and resource-bounded execution substrates.
- Implement failure semantics, including retry policies, idempotency, and error handling.
- Design trace schemas for debugging, training signals, and audit records.
- Build task suites, verifiers, and scoring infrastructure for agent evaluation.
- Develop regression gates that account for cost, latency, and quality.
- Implement statistical methods to validate results, including variance and pass@k analysis.
- Create tooling for trace inspection, replay, and diffing.
Requirements
- Strong software engineering fundamentals and advanced Python.
- Experience shipping and operating production services.
- Deep understanding of concurrency and distributed systems (async execution, worker pools, queues).
- Proficiency with containers and sandboxing, specifically Docker and OCI internals.
- Experience building test and CI infrastructure and managing large-scale parallel test suites.
- Measurement literacy regarding variance, sampling, and statistical significance.
- Strong observability instincts, including distributed tracing and structured logging.
- Ability to work with loosely specified problems and take full ownership.
Skills
- Python
- Docker
- Distributed Systems
- CI/CD
- Observability