Foundry Tool for spoken language, able to transcribe audio into text, synthesise voice from text, translate what is said and run live voice conversations.
Also called Azure Speech, Azure AI Speech.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Azure Speech in Foundry Tools in context, with comparison tables and the common traps.
Terms in this definition
- Chat message roles
Labels on chat messages: instructions go under system, the person's input under user, the model's previous answers under assistant, and results returned by a called tool under tool (or function).
Related terms
- Professional voice fine-tuning
Azure Speech feature with limited access, created in Speech Studio or the Foundry portal, that builds a bespoke brand voice from the recordings and consent statement of a consenting voice talent.
- Speaker recognition
Capability of Azure Speech that verifies or identifies a speaker based on their voice. It is a Limited Access feature and differs from speech recognition, which transcribes the words spoken.
- Speech SDK
The Azure Speech client library, installed in Python as azure-cognitiveservices-speech and imported as speechsdk, providing SpeechSynthesizer, SpeechRecognizer, SpeechConfig and classes for audio configuration.
- Speech Studio
A web portal offering no-code Azure Speech tools, including batch speech to text, audio content creation, custom voice and custom speech.
- Speech to text
Capability of Azure Speech that turns spoken audio into text for voice commands, captions and transcripts of calls or meetings. Identifying the speaker is not part of it.
- Speech translation
Takes speech as it is spoken and renders it, in real time, as text or synthesised audio in one target language or more; an Azure Speech capability.
- Text to speech
Speech synthesis in Azure Speech: text becomes natural-sounding audio using standard or custom neural voices, adjustable through SSML. Speech recognition works in the reverse direction.
- Whisper
An OpenAI model that transcribes speech and can translate it into English, with files capped at 25 MB; it is offered through Azure Speech and Azure OpenAI. Azure OpenAI's whisper (001) model is due to retire on 15 December 2026.