Turning text into spoken audio that sounds natural, for example to read messages out loud.
Also called text to speech, TTS.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Speech synthesis in context, with comparison tables and the common traps.
Terms in this definition
- CRUD
Shorthand for create, read, update and delete, the four basic things you do with data. Data-plane roles in Azure Cosmos DB, for instance, authorise those operations on items.
- Agents (classic) API
First-generation Foundry Agent Service API, based on threads, messages and runs. It is deprecated, replaced by conversations and responses, and retires on 31 March 2027.
Related terms
- Batch synthesis API
Asynchronous text to speech REST API suited to big batches of SSML or plain text. Unlike the real-time REST API, which stops at 10 minutes of audio, it can generate longer output such as audiobooks.
- Custom voice
Text to speech feature with limited access that produces a distinctive synthetic voice, for a brand or character, trained on recordings of a real speaker.
- Neural voice
Ready-made text to speech voice built on deep neural networks, usable immediately across more than 100 languages and locales.
- SpeechSynthesizer
The Speech SDK class used for text to speech, offering SpeakTextAsync and SpeakSsmlAsync. Created from a SpeechConfig plus an AudioOutputConfig, it sends audio to a stream, file or speaker.
- SSML
Markup based on XML that controls how text to speech sounds, including pronunciation, pauses, volume, rate, pitch and voice. It is submitted with speak_ssml_async().
- Text to speech
Speech synthesis in Azure Speech: text becomes natural-sounding audio using standard or custom neural voices, adjustable through SSML. Speech recognition works in the reverse direction.
- Voice Live API
Real-time voice agents can be built on this fully managed, low-latency speech-to-speech service. One WebSocket brings together speech recognition, a generative model and text to speech, with optional avatars and function calling, and replies come back as audio rather than text alone.
- Voices list API
A GET request in the text to speech REST API returning standard voices with their locale, gender and styles; the path is /cognitiveservices/voices/list on the regional tts endpoint or /tts/cognitiveservices/voices/list on the resource endpoint. No speech is synthesised.