Markup based on XML that controls how text to speech sounds, including pronunciation, pauses, volume, rate, pitch and voice. It is submitted with speak_ssml_async().
Also called Speech Synthesis Markup Language.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains SSML in context, with comparison tables and the common traps.
Terms in this definition
- XML
A text-based format for data, used for instance in message bodies.
- Speech synthesis
Turning text into spoken audio that sounds natural, for example to read messages out loud.
- Volume
Holds non-tabular files sitting in cloud storage under Unity Catalog governance, addressed as
/Volumes/<catalog>/<schema>/<volume>. A volume can be external or managed, andWRITE VOLUMEorREAD VOLUMEdecide who may use it.
Related terms
- Batch synthesis API
Asynchronous text to speech REST API suited to big batches of SSML or plain text. Unlike the real-time REST API, which stops at 10 minutes of audio, it can generate longer output such as audiobooks.
- mstts:express-as
SSML extension from Microsoft for choosing a speaking style for a voice, such as calm, gentle or advertisement_upbeat. If the style is not valid, the entire element is disregarded.
- speak_text_async()
Method of SpeechSynthesizer that synthesises plain text without blocking, returning a future that resolves to a SpeechSynthesisResult. For SSML input, speak_ssml_async() is used.
- SSML emphasis element
Element in SSML that increases or reduces stress on words (strong, moderate, none or reduced). Only a small number of neural voices support it.
- SSML prosody element
Element in SSML for changing the volume, rate, range, contour and pitch of synthesised speech.
- SSML voice element
Element in SSML that chooses, through its name attribute, which voice speaks the text inside it. Accent follows the voice's locale, so any two en-US voices sound alike in accent.
- Text to speech
Speech synthesis in Azure Speech: text becomes natural-sounding audio using standard or custom neural voices, adjustable through SSML. Speech recognition works in the reverse direction.