Role Overview
This is a founding engineering role for a VC-backed health-tech, AI, and wearable startup. You will work directly with the founders to design, build, deploy, and scale production-grade AI systems used by real customers, moving beyond prototypes to real-world multi-agent architectures and RAG pipelines.
Responsibilities
- Design, build, and deploy production-grade LLM applications and conversational AI systems
- Build multi-agent workflows using frameworks such as LangGraph or CrewAI
- Develop production RAG systems using vector databases and hybrid retrieval architectures
- Optimize inference latency, GPU utilization, memory consumption, and LLM serving costs
- Build reliable model serving infrastructure with monitoring, observability, evaluations, retries, and guardrails
- Work on real-time AI applications involving streaming, speech (STT/TTS), or voice agents
- Design scalable backend services and APIs supporting AI workloads
- Collaborate with product and engineering teams to rapidly ship AI features into production
Requirements
- 2+ years of software engineering or AI engineering experience
- Strong hands-on experience building production LLM applications
- Experience deploying RAG systems into production
- Experience with vector databases such as Pinecone, pgvector, Weaviate, Chroma, or FAISS
- Experience with LangChain, LangGraph, or LlamaIndex
- Strong understanding of transformers, embeddings, prompt engineering, and evaluation methodologies
- Experience building model serving infrastructure and production ML systems
- Strong backend engineering skills (Python preferred)
- Experience with cloud platforms (AWS/GCP/Azure), Docker, and scalable backend architecture
Skills
- Python
- LangChain
- LLM
- AWS
- Vector Databases