For Employers

indigo.ai

Software Engineer, Voice - Milan

Posted an hour ago
€40000 - €70000 per year
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

The Software Engineer will own the end-to-end real-time voice pipeline, including audio ingress, streaming STT/TTS, and turn-taking logic. They are responsible for optimizing latency, ensuring human-like conversational flow, and maintaining production reliability for AI agents.

If you are here, it is because you know that we are looking for a Software Engineer, Voice for our Product team.

indigo.ai is the leading platform in Italy for building next-generation AI Agents that transform the way companies communicate with their customers. Since 2016, we’ve been helping enterprises in industries such as finance, insurance, utilities, retail, and e-commerce to evolve their Customer Experience through conversational AI. We don’t just “sell software”: we enable a shift in how organizations interact with people, automating millions of conversations every year, reducing operational costs, improving conversion rates, and creating more personalized, scalable, and compliant customer journeys.

Backed by a recent €10 million investment from Azimut, we are on a mission to take this technology global. This is a unique opportunity to join a well-funded, highly ambitious team and play a direct role in shaping the future of enterprise AI.

To make this happen, we are looking for a Software Engineer, Voice to support our Chief Product Development Officer in making talking to AI on the phone feel human.

What are we looking for?

We are looking for a Software Engineer, Voice to own the real-time voice layer of our AI Agents. Voice is where conversational AI is being decided right now, and our voice agents already handle production phone traffic for large companies. The bar is moving fast: we want to make the leap from "works reliably" to "feels human on a real phone line", and we want one person to own that leap. This is a specialist role with end-to-end ownership: the architecture, the model and provider choices, the latency budget, the way a conversation feels. You'll join our Product Engineering team, reporting to our Chief Product Development Officer, as our first full-time engineer dedicated to voice. And if voice grows the way we believe it will, you'll shape the team that grows around it.

Key Responsibilities:

  • Own the real-time voice pipeline end-to-end. From audio ingress on the telephony edge, through streaming STT and turn-taking, to the agent brain and back out through streaming TTS. Every millisecond in between is yours.

  • Engineer how fast the agent feels. Semantic end-of-turn detection, preemptive generation on partial transcripts, eager TTS, filler and backchannel strategies that mask tool calls. All measured on real 8kHz phone audio, not in a browser demo.

  • Make turn-taking human. Barge-in that survives noisy lines. Endpointing policies that know the dialog state, so a caller never gets cut off mid-IBAN. The difference between an IVR and a conversation lives here.

  • Raise voice quality on the channel that actually ships: the phone. Benchmark and A/B STT and TTS providers on real G.711 calls (Italian first: WER, naturalness, numbers and codes read right), exploit wideband/HD voice where the carrier allows it, and experiment with context-aware TTS and conversational speech models as they mature.

  • Build the evaluation harness. Turn "this voice sounds better" into numbers we trust: per-stage latency budgets, turn-taking metrics, regression suites on recorded calls, quality gates before anything reaches a client.

  • Keep production boringly reliable. Per-stage observability, live-call incident debugging (dead air, stuck turns, provider hiccups), graceful degradation when a vendor blinks.

  • Track a weekly-moving ecosystem and turn it into strategy. New STT/TTS/speech-to-speech releases land every month. You decide what we integrate, what we self-host for EU compliance and data residency, and what we skip. And you make provider swaps cheap.

You will need:

The filter is not your degree, and it's not years-of-experience arithmetic. It's having built it. Tell us about a real-time voice or audio system you designed and shipped: the latency budget, where it broke, and what you changed to make it feel right. That tells us more than any title.

  • Real-time audio systems, shipped. You've built voice agents, telephony systems, conferencing or live-streaming products that ran in production. You know what it means to move audio over WebSockets/WebRTC/SIP, through codecs (G.711/μ-law, Opus), against a latency budget.

  • The modern voice AI stack, hands-on. Streaming STT and TTS, VAD and turn detection, voice orchestration frameworks (Pipecat, LiveKit Agents or equivalent), speech-to-speech models. You have opinions on the trade-offs, grounded in things you've actually built, not blog posts.

  • Strong software engineering. TypeScript/Node.js and/or Python, and the maturity to own a production service end-to-end: containers, cloud infrastructure, CI/CD, observability.

  • A latency obsession. You think in milliseconds per stage, you instrument before you optimize, and you know the difference between measured and perceived latency, and how to exploit it.

  • A product ear. You can hear the difference between a demo and a conversation, and you can translate what you hear into engineering priorities and measurable evals.

  • An AI-native way of working. You use agentic coding tools (e.g. Claude Code) daily and you're good at directing them: setting up the problem, judging the output.

  • Language Skills: Fluent English.

We will really like (but they are not required):

  • Italian: our voice market is Italian-first, and you'll be tuning pronunciation, prosody and evals for it every week.

  • Contact-center / CCaaS ecosystem experience: SIP trunking, SBCs, enterprise telephony platforms.

  • ML audio experience: evaluating or fine-tuning ASR/TTS models, working with speech datasets.

  • Elixir: our agent platform is built on it.

  • Open source: contributions to open-source voice/audio projects.

Our values:

At indigo.ai we prize Vision (curiosity and courage to challenge the status quo), Connection (empathy, candor, and trust), Responsibility (ownership, reliability, and follow-through), and Excellence (the habit of raising the bar and refining until it’s right). If you naturally think ahead, build strong relationships, take accountability, and obsess over the quality of what you deliver, we’re probably a great match.

What do we offer?

  • A key role in one of Europe’s fastest-growing AI scale-ups, backed by a €10M investment from Azimut.

  • A competitive salary in the range of 40-70k RAL, commensurate with experience + a performance-based Bonus.

  • A flexible, remote-friendly work environment.

  • Meal vouchers and Welfare programs to support your everyday life.

  • Access to a dedicated education budget for continued learning and growth.

  • Career Development Plan, ensuring a clear path for both personal and professional growth.

  • Top-grade equipment, which may include a MacBook Air, iPhone, and other top-tier devices.

  • Unlimited coffee.

  • Company retreats in stunning locations throughout the year.

Where is the job?

This position is fully remote, so there's no need to be in a specific location to do your work. That said, we have an amazing office at SPACES, Piazza Gae Aulenti 1/Torre B in Milan, available to anyone who wants to use it. We also love getting together and organize various retreats and meetups throughout the year to stay connected.

Why join us?

At indigo.ai you’ll be part of a fast-growing SaaS company where your ideas and work can turn into products used by millions. Join us and you’ll:

  • Work with the most advanced AI technologies, shaping how enterprises across industries engage with their customers.

  • Be part of a dynamic and passionate team, where everyone has a direct impact on company growth.

  • Grow in a culture that values transparency, continuous learning, and career development, with clear paths for personal and professional progression.

  • Enjoy a flexible, people-first environment, with remote-friendly policies, stunning retreats, and a strong focus on well-being.

  • Contribute to our mission of reshaping how companies and people communicate worldwide.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Senior DeFi Engineer — Autheo | Live Mainnet, Staking, Pools & THEO Utility

Part Time United States Software Development

Senior Solutions Architect — Autheo | Live Mainnet, Enterprise & Partner Designs

Part Time United States Software Development

CRM Program Manager - UAE (contract - remote)

Full Time Saudi Arabia Software Development

CRM Program Manager - Turkey (contract - remote)

Full Time Turkey Software Development

CRM Program Manager - Brazil (contract - remote)

Full Time Brazil Software Development

Solution Engineer - FedRAMP

Full Time United States $220K - $250K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 212,037+ Jobs in Software Engineer

Answer easy questions

Answer easy questions

212,037+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified