Software Engineer, Voice - Milan
Own the real-time voice AI stack behind our AI Agents: telephony, turn-taking, speech orchestration. Make talking to AI on the phone feel human. First engineer fully dedicated to voice.
About indigo.ai
Since 2016, indigo.ai has been using AI to transform how companies talk with their customers. All the best stories start with great conversations: the relationships that matter are built on the ability to communicate, listen, and understand each other. Through our AI Agents platform we have automated millions of conversations across voice, chat and messaging, integrating with the CCaaS and enterprise systems of our clients and improving their operational and commercial performance. We are a fast-growing European scale-up, backed by Azimut with a €15M investment.
The role
Voice is where conversational AI is being decided right now, and it's the channel where our enterprise clients feel quality, or its absence, most. Our voice agents already handle real production phone traffic for large companies.
Today our voice runtime is a service we built in-house: a Node.js/TypeScript router that terminates telephony media streams, orchestrates streaming speech-to-text, turn-taking and text-to-speech across multiple speech providers, and talks to our agent platform, the "brain" that owns knowledge, guardrails, routing, integrations and compliance. It works, in production, today.
But the bar is moving fast. Speech-to-speech models, semantic turn detection and context-aware TTS are redefining what callers expect from a machine on the phone. We want to make the leap from "works reliably" to "feels human on a real phone line", and we want one person to own that leap.
This is not a generic backend position. It's a specialist role with end-to-end ownership: the architecture, the orchestration, the model and provider choices, the latency budget, the way a conversation feels. You'll be our first engineer fully dedicated to voice, working directly with the founders and the platform team. And if voice grows the way we believe it will, you'll shape the team that grows around it.
What you'll do
- Own the real-time voice pipeline end-to-end. From audio ingress on the telephony edge, through streaming STT and turn-taking, to the agent brain and back out through streaming TTS. Every millisecond in between is yours.
- Engineer how fast the agent feels. Semantic end-of-turn detection, preemptive generation on partial transcripts, eager TTS, filler and backchannel strategies that mask tool calls. All measured on real 8kHz phone audio, not in a browser demo.
- Make turn-taking human. Barge-in that survives noisy lines. Endpointing policies that know the dialog state, so a caller never gets cut off mid-IBAN. The difference between an IVR and a conversation lives here.
- Raise voice quality on the channel that actually ships: the phone. Benchmark and A/B STT and TTS providers on real G.711 calls (Italian first: WER, naturalness, numbers and codes read right), exploit wideband/HD voice where the carrier allows it, and experiment with context-aware TTS and conversational speech models as they mature.
- Build the evaluation harness. Turn "this voice sounds better" into numbers we trust: per-stage latency budgets, turn-taking metrics, regression suites on recorded calls, quality gates before anything reaches a client.
- Keep production boringly reliable. Per-stage observability, live-call incident debugging (dead air, stuck turns, provider hiccups), graceful degradation when a vendor blinks.
- Track a weekly-moving ecosystem and turn it into strategy. New STT/TTS/speech-to-speech releases land every month. You decide what we integrate, what we self-host for EU compliance and data residency, and what we skip. And you make provider swaps cheap.
Who we're looking for
The filter is not your degree, and it's not years-of-experience arithmetic. It's having built it. Tell us about a real-time voice or audio system you designed and shipped: the latency budget, where it broke, and what you changed to make it feel right. That tells us more than any title.
Must-have: the core of the role
- Real-time audio systems, shipped. You've built voice agents, telephony systems, conferencing or live-streaming products that ran in production. You know what it means to move audio over WebSockets/WebRTC/SIP, through codecs (G.711/μ-law, Opus), against a latency budget.
- The modern voice AI stack, hands-on. Streaming STT and TTS, VAD and turn detection, voice orchestration frameworks (Pipecat, LiveKit Agents or equivalent), speech-to-speech models. You have opinions on the trade-offs, grounded in things you've actually built, not blog posts.
- Strong software engineering. TypeScript/Node.js and/or Python, and the maturity to own a production service end-to-end: containers, cloud infrastructure, CI/CD, observability.
- A latency obsession. You think in milliseconds per stage, you instrument before you optimize, and you know the difference between measured and perceived latency, and how to exploit it.
- A product ear. You can hear the difference between a demo and a conversation, and you can translate what you hear into engineering priorities and measurable evals.
- An AI-native way of working. You use agentic coding tools (e.g. Claude Code) daily and you're good at directing them: setting up the problem, judging the output.
- Fluent English.
Nice to have
- Italian: our voice market is Italian-first, and you'll be tuning pronunciation, prosody and evals for it every week.
- Contact-center / CCaaS ecosystem experience: SIP trunking, SBCs, enterprise telephony platforms.
- ML audio experience: evaluating or fine-tuning ASR/TTS models, working with speech datasets.
- Elixir: our agent platform is built on it.
- Contributions to open-source voice/audio projects.
What we offer
- A key role in a fast-growing European AI scale-up, backed by Azimut (€15M investment), where you own an entire technical domain end to end.
- A competitive package (fixed + variable), aligned with seniority and market benchmarks.
- A flexible, remote-friendly work environment built on trust and autonomy.
- A dedicated education budget and a structured Career Development Plan.
- High-quality equipment (MacBook Pro and iPhone).
- A young, dynamic team with real attention to people's well-being, plus company retreats in inspiring locations throughout the year.
Where is the work?
Milan, hybrid. Our office is at SPACES, Piazza Gae Aulenti 1/Torre B, where we regularly meet to collaborate in person; the rest of the time we work wherever we're most effective.
Why join us?
- Voice is the frontier. 2026 is the year AI phone conversations stop sounding like machines. You'd be the person who makes that happen for real enterprise traffic, not a demo.
- Total ownership. One domain, one owner, direct line to the founders. Your decisions ship, and you literally hear them in millions of phone calls.
- The right architecture to build on. A controllable enterprise brain (knowledge, guardrails, routing, compliance) that pure voice wrappers don't have. Your job is to give it a voice that matches.
- Department
- Product
- Role
- Product Engineer
- Locations
- Milano
- Remote status
- Hybrid