Role Overview
We are looking for a skilled Voice AI Engineer to build and improve real-time conversational voice systems. The candidate should have experience with Speech-to-Text (STT), Text-to-Speech (TTS), LLMs, voice agents, APIs, and real-time AI applications.
Responsibilities
- Build and optimise real-time voice pipelines, including Speech-to-Text (STT), Text-to-Speech (TTS), and orchestration layers.
- Improve end-to-end latency and create natural voice interactions with effective turn-taking, barge-in, interruption handling, endpointing, and silence recovery.
- Design and improve the LLM agent layer, including prompt engineering, tool/function calling, conversation state, context management, and fallback behaviour.
- Integrate voice AI systems with telephony platforms, EMR, practice management systems, scheduling tools, and external APIs.
- Develop evaluation and monitoring systems to measure voice quality, transcription accuracy, task completion, latency, and failure rates.
- Analyse call recordings and user feedback to identify opportunities for product improvement.
- Collaborate with clinical, engineering, product, and customer success teams.
- Take ownership of Voice AI features from initial concept and development through testing, deployment, and production support.
Requirements
- 1–5 Years of experience.
- Strong experience with Python, APIs, and backend development.
- Knowledge of STT, TTS, LLMs, conversational AI, and real-time voice systems.
- Experience with prompt engineering, AI agent workflows, RAG, tool calling, and context management.
- Understanding of WebSockets, streaming systems, low-latency architectures, and real-time communication.
- Experience integrating third-party APIs and telephony platforms.
- Knowledge of cloud platforms, databases, monitoring, and production deployment.
- Bachelor's degree preferred.
Skills
- Python
- LLMs
- Speech-to-Text (STT)
- Text-to-Speech (TTS)
- WebSockets