Role Overview
We are hiring a Conversational AI Engineer to architect and ship voice AI agents that handle thousands of live customer conversations daily. You will own real-time audio pipelines, telephony integrations, and the orchestration layer that ties STT, TTS, and LLMs into reliable sub-second voice experiences. This is a hands-on engineering role at the core of our Convogent accelerator.
Responsibilities
- Build voice agents using Pipecat, LiveKit, or equivalent frameworks.
- Engineer real-time audio pipelines to optimize STT and TTS flows for low latency and high concurrency.
- Own telephony integration via WebRTC, SIP, RTP, and WebSocket, integrating with Twilio and Telnyx.
- Handle multi-speaker audio including speaker diarization and segmentation.
- Build the orchestration layer connecting STT, TTS, and LLM agents into resilient workflows.
- Tune for production by implementing Voice Activity Detection, echo cancellation, and codec optimization.
- Operate at scale on AWS, including monitoring, logging, and incident response.
- Partner with product and customer teams to translate requirements into deployable solutions.
Requirements
- Hands-on experience building and deploying voice agents with Pipecat or LiveKit.
- Deep expertise in STT (Deepgram, Whisper, AssemblyAI) and TTS (ElevenLabs, Cartesia, Amazon Polly).
- Strong working knowledge of telephony, SIP, and WebRTC.
- Production fluency in Python and at least one of Node.js, TypeScript, or Go.
- Practical use of AI/ML libraries and integration with LLMs (GPT-4o, Claude, Bedrock).
- Audio codec depth (Opus, G.711, G.729) and real-time streaming.
- Cloud deployment experience on AWS using Docker and Kubernetes.
- Track record of shipping voice AI or telephony systems to production.
Skills
- Python
- AWS
- WebRTC
- LiveKit
- LLMs
Nice to Have
- Experience with media server architectures (SFU/MCU) and Asterisk, FreeSWITCH, or Kamailio.
- Voice analytics and Indian-language STT/TTS models.
- AWS or speech technology certifications.