Role Overview
You'll build and optimize the AI systems behind our platform - starting with the real-time conversational engine behind Joy, and extending across the surfaces that surround it: call summarization and analytics, multi-modal orchestration, document intelligence, and computer-use agents. This is a hands-on engineering role at the intersection of applied LLMs, real-time audio, and agentic workflows.
Responsibilities
- Build and optimize Joy's real-time voice pipeline (streaming STT → LLM → TTS) on a voice AI orchestration framework.
- Implement and tune latency techniques: sentence-chunked streaming TTS, speculative tool prefetching, filler speech, parallel tool execution, and prompt caching.
- Integrate LLMs for conversational turns - function/tool calling, prompt engineering, structured outputs, and grounded responses.
- Develop and improve the RAG layer over vector search: retrieval quality, knowledge-base grounding, and refusal behavior.
- Benchmark models across multiple frontier LLM providers on latency, cost, and quality.
- Build and maintain evaluation harnesses covering retrieval metrics and generation metrics.
- Integrate telephony and messaging (SIP trunking / SMS) and coordinate with EHR middleware.
- Build call summarization that turns raw transcripts into structured post-call notes.
- Contribute to the orchestration layer that coordinates AI across modalities (voice, SMS, document).
- Build the document intelligence pipeline for healthcare benefits-verification documents.
- Develop computer-use / browser agents that operate EHR portals and web workflows.
- Deploy, monitor, and tune services on a serverless container platform.
- Own the AI/ML technical vision and roadmap across the full agent portfolio.
- Lead, mentor, and grow the AI engineering team.
- Set the end-to-end architecture for the voice pipeline and orchestration layer.
- Define and drive service-level objectives per surface.
- Own model strategy: build-vs-buy decisions, provider and region selection.
- Establish the evaluation and quality framework.
Requirements
- 2–5 years building production software, with meaningful applied LLM / ML work.
- Strong Python, including async and modern async web frameworks.
- Hands-on experience with LLM applications: prompting, tool/function calling, RAG, embeddings, and vector databases.
- Solid grasp of real-time or streaming systems and the discipline to reason about latency budgets.
- Comfort with cloud infrastructure (GCP) and containerized deployment.
- Agentic / computer-use experience - browser or desktop automation, tool-using agents, multi-step task orchestration.
- Streaming STT / TTS systems.
- LLM / RAG evaluation tooling and frameworks.
- A data-driven, benchmark-first instinct.
Nice to Have
- Voice / telephony experience: SIP