Design, build, and deploy intelligent real-time voice applications by integrating STT, LLM, and TTS technologies. Develop backend services, APIs, and event-driven systems while optimizing voice agents for latency and conversational quality.
This is a remote position.
Shuru is an AI-native engineering services company building products with leading startups, enterprises, and well-funded scale-ups across India, Southeast Asia, the UAE, and global markets. Our engineers work in small, dedicated squads, ship in weekly cycles, and use AI-assisted development tools (Cursor, Claude Code, GitHub Copilot) as a standard part of how we build. We hire for fundamentals, give engineers real ownership, and back them with a culture that respects deep work.
We are looking for a Voice AI Engineer to design, build, and deploy intelligent real-time voice applications. You will work on conversational AI systems that combine Speech-to-Text (STT), Large Language Models (LLMs), Text-to-Speech (TTS), telephony, and real-time audio streaming to deliver natural and responsive voice experiences.
The ideal candidate has strong software engineering fundamentals, experience with AI/LLM applications, and an understanding of low-latency voice systems.
- Design and develop production-ready AI voice agents and conversational voice applications.
- Integrate STT, LLM, and TTS technologies into real-time voice pipelines.
- Build integrations with telephony and communication platforms such as Twilio, SIP, WebRTC, and LiveKit.
- Develop backend services, APIs, WebSocket connections, and event-driven systems.
- Implement RAG, function calling, tool use, memory, and multi-step AI agent workflows.
- Optimize voice applications for latency, interruption handling, turn-taking, accuracy, and conversational quality.
- Integrate voice agents with CRMs, databases, calendars, customer-support systems, and third-party APIs.
- Develop monitoring, logging, analytics, and evaluation frameworks for voice conversations.
- Test and improve transcription accuracy, response quality, TTS naturalness, and overall user experience.
- Deploy and maintain scalable AI services in cloud environments.
- Collaborate with product, engineering, and business teams to translate requirements into voice AI solutions.
Requirements
- Strong programming experience in Python and/or Node.js/TypeScript.
- Experience building applications using LLMs and generative AI APIs.
- Knowledge of Speech-to-Text and Text-to-Speech systems.
- Experience with REST APIs, WebSockets, asynchronous programming, and real-time systems.
- Understanding of prompt engineering, function calling, AI agents, and RAG.
- Experience working with databases and vector databases.
- Familiarity with cloud platforms such as AWS, GCP, or Azure.
- Strong debugging, problem-solving, and software engineering skills.
- Understanding of Git, CI/CD, containerization, and production deployment practices.
Benefits
- Competitive salary and benefits package.
- Work with experienced product and engineering leaders.
- Opportunities for learning, mentorship, and career growth.
- A chance to make a real impact across diverse, innovative projects.