← Back to Jobs
Senior Voice AI Engineer
Aivar Innovations
Software EngineeringBengaluru, Karnataka
2 months ago
💻 Open SourceDirectly related to vibe coding: integrates LLMs/STT/TTS in real-time pipelines and uses agentic coding tools to measure and improve voice quality.
About the Role
Lead ownership of Convogent's distributed, real-time voice platform to keep thousands of simultaneous calls fast, reliable, and natural. Integrate and optimize STT/ASR, LLM, TTS, RAG, and tool-calling to meet tight latency budgets and production quality standards while mentoring the engineering team.
Job Description
Role
Senior Voice AI Engineer responsible for owning Convogent’s distributed, real-time voice platform that powers thousands of concurrent conversations. The role focuses on performance, reliability, and measurable quality across the full voice pipeline (STT/ASR, LLMs, TTS, RAG, and tool-calling).
Key Responsibilities
- Own voice AI at scale and the core platform that orchestrates parallel voice calls.
- Optimize performance and throughput (e.g., lift calls-per-vCPU) by extending frameworks like Pipecat or rewriting hot paths in Go.
- Manage end-to-end latency budgets (~800ms target) across STT, LLM, TTS, RAG, tool-calling, turn detection, and network hops.
- Diagnose and fix production call failures quickly by interpreting signals and locating faults in the pipeline.
- Build measurable voice-to-voice evaluation systems that gate releases and detect provider drift.
- Make provider tradeoffs for STT/LLM/TTS/RAG/tool-calling balancing quality and latency per conversation flow.
- Extend the platform into adjacent products (live assist, agent assist) and define engineering standards for a growing team.
- Review and mentor other engineers.
Requirements
- 5+ years building production products and systems, with at least 2 years operating real-time voice in production.
- Deep, end-to-end understanding of the voice pipeline and how each stage affects quality and latency.
- Strong production-grade experience in Go and Python; production performance-critical Go is especially valued.
- Proven track record running voice AI at scale and scaling Pipecat or comparable real-time voice frameworks to high concurrency.
- Fluency in provider tradeoffs among STT, LLM, TTS, RAG, and tool-calling.
- Eval-driven approach to quality: experience building or running voice/LLM evals with real signals.
- Comfortable using agentic coding tools and treating quality as a measured signal.
Not a fit
- Candidates who have only called LLM APIs or built prompts without operating real-time pipelines under load.
- Candidates who have only used managed voice platforms (e.g., Vapi, Retell, Bland) without scaling the underlying layers.
- Pure web-backend CRUD engineers without real-time/streaming/voice experience.
- ML researchers focused on training/fine-tuning models (this role integrates provider models rather than building them).