Role Overview
You'll join a small, fast-moving AI engineering team working on next-generation language model applications and agentic systems. This is hands-on product engineering — not research busywork — with direct exposure to production AI infrastructure and real users.
Responsibilities
- Build and ship features for LLM-powered products, including RAG pipelines, structured output layers, and tool-calling agents
- Design and integrate agentic workflows using orchestration frameworks (LangGraph, LlamaIndex, or similar)
- Work with model evaluation frameworks — write evals, run benchmarks, and iterate on prompt engineering
- Build and maintain observability tooling: tracing, logging, and latency monitoring for AI inference in production
- Contribute to model context management — chunking strategies, embedding pipelines, vector store integration
- Participate in code reviews, architecture discussions, and documentation
Requirements
- Currently pursuing a BS or MS in Computer Science, Engineering, or a related field
- Solid Python fundamentals; experience with TypeScript, Go, or Rust is a strong plus
- Familiarity with LLM APIs (OpenAI, Anthropic, Gemini, or open-source equivalents via Ollama/vLLM)
- Exposure to RAG patterns, vector databases (Pinecone, Weaviate, pgvector), or embedding workflows
- Working knowledge of containerisation (Docker) and comfort with cloud environments (AWS / GCP / Azure)
- Understanding of agentic patterns: tool use, function calling, ReAct loops, multi-agent coordination
- Familiarity with model evaluation concepts — LLM-as-judge, benchmark design, RAGAS or similar
- Excellent communicator; can write clearly and discuss tradeoffs in technical reviews
- Curiosity-driven — you read release notes, follow model launches, and have opinions about context windows
Skills
- Python
- LLM APIs
- RAG pipelines
- LangGraph
- Docker