Role Overview
Heu.ai is seeking a hands-on AI / LLM Engineer to design, build, and deploy intelligent system architectures and agentic workflows. You will go beyond simple wrapper scripts to build production-grade Retrieval-Augmented Generation (RAG) pipelines, complex multi-agent orchestrations, and robust backend APIs.
Responsibilities
- Design and deploy stateful, multi-agent orchestrations and dynamic decision-making graphs using LangGraph.
- Build, evaluate, and optimize hybrid retrieval strategies including vector search, BM25, and reranking.
- Architect and iterate on structured prompts and establish evaluation frameworks to monitor quality and hallucination rates.
- Develop scalable, asynchronous REST/gRPC APIs and microservices using FastAPI or Flask.
- Manage and optimize vector databases such as Pinecone, Qdrant, Chroma, or PGVector.
- Monitor latency, token usage, and implement caching strategies for low-latency experiences.
Requirements
- 1-3 years of experience in AI/LLM engineering.
- Strong proficiency in Python with an emphasis on async programming and clean architecture.
- Demonstrated experience with LangGraph, LangChain, or custom agentic execution systems.
- Proven experience implementing end-to-end RAG workflows.
- Solid background in building production APIs via FastAPI or modern web frameworks.
- Experience with structured outputs using JSON schema and Pydantic.
- Experience with fine-tuning open-source models like Llama or Mistral or using Ollama and vLLM.
Skills
- Python
- LangChain
- RAG
- FastAPI
- Vector Databases