Real-time voice agents can be built on this fully managed, low-latency speech-to-speech service. One WebSocket brings together speech recognition, a generative model and text to speech, with optional avatars and function calling, and replies come back as audio rather than text alone.
Also called Voice Live.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Voice Live API in context, with comparison tables and the common traps.
Terms in this definition
- Fully managed
An Android Enterprise setup for company-owned devices that one person uses only for work, with Intune in control of the entire device.
- Speech recognition
Turning spoken audio into text, for example call transcripts or captions.
- Speech synthesis
Turning text into spoken audio that sounds natural, for example to read messages out loud.
- Function tool
Gives an agent or model access to logic written in your own application. It is listed among the tools as type
functionwith a name, a description and a JSON Schema describing its arguments; when called, your code does the work and passes the result back.
Related terms
- Voice-based prompt agent
A Foundry prompt agent, set up with a model, instructions, tools and audio settings, that users converse with through Voice Live, so you don't run voice orchestration yourself.