Speech synthesis in Azure Speech: text becomes natural-sounding audio using standard or custom neural voices, adjustable through SSML. Speech recognition works in the reverse direction.
Also called speech synthesis.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Text to speech in context, with comparison tables and the common traps.
Terms in this definition
- Speech synthesis
Turning text into spoken audio that sounds natural, for example to read messages out loud.
- Azure Speech in Foundry Tools
Foundry Tool for spoken language, able to transcribe audio into text, synthesise voice from text, translate what is said and run live voice conversations.
- Standard deployment type
A Foundry deployment type billed per token that keeps processing of prompts and responses inside the Azure geography of the resource, meeting data residency needs at lower volumes.
- SSML
Markup based on XML that controls how text to speech sounds, including pronunciation, pauses, volume, rate, pitch and voice. It is submitted with speak_ssml_async().
- Speech recognition
Turning spoken audio into text, for example call transcripts or captions.