Role Overview
This is a founding engineering role for a VC-backed health-tech, AI, and wearable startup. You will work directly with the founders to design, build, deploy, and scale production-grade AI systems, moving beyond prototypes to real-world applications used by customers.
Responsibilities
- Design, build, and deploy production-grade LLM applications and conversational AI systems
- Build multi-agent workflows using frameworks such as LangGraph or CrewAI
- Develop production RAG systems using vector databases and hybrid retrieval architectures
- Optimize inference latency, GPU utilization, and LLM serving costs
- Build reliable model serving infrastructure with monitoring, observability, and guardrails
- Work on real-time AI applications involving streaming, speech (STT/TTS), or voice agents
- Design scalable backend services and APIs supporting AI workloads
Requirements
- 2+ years of software engineering or AI engineering experience
- Strong hands-on experience building production LLM applications
- Experience deploying RAG systems into production
- Experience with vector databases such as Pinecone, pgvector, Weaviate, Chroma, or FAISS
- Experience with LangChain, LangGraph, or LlamaIndex
- Strong understanding of transformers, embeddings, and prompt engineering
- Experience with cloud platforms (AWS/GCP/Azure), Docker, and scalable backend architecture
- Strong backend engineering skills (Python preferred)