Turning spoken audio into text, for example call transcripts or captions.
Also called speech to text, STT.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Speech recognition in context, with comparison tables and the common traps.
Related terms
- Captioning
Turning speech in live or recorded video into timed text shown on screen; a speech to text use case that can also filter profanity and display partial results.
- Fast transcription
Converts recorded audio files to text in one synchronous REST call that finishes quicker than the audio's duration and has predictable latency; part of Speech to text.
- Intent recognition (Speech)
Retired on 30 September 2025, this Speech SDK feature mapped spoken utterances to intents; use speech to text and then CLU or an Azure OpenAI model instead.
- Language identification
Speech capability that works out which of a set of possible languages is being spoken: up to 4 when detecting at the start of audio, or up to 10 when detecting continuously. It runs with speech to text or speech translation.
- Real-time transcription
Speech to text mode that converts streamed audio, from a microphone or file, into text as it is recognised, with interim results along the way. Typical uses are voice input and live captioning.
- Speaker recognition
Capability of Azure Speech that verifies or identifies a speaker based on their voice. It is a Limited Access feature and differs from speech recognition, which transcribes the words spoken.
- Speech captioning
Generating captions with timestamps, as SRT or WebVTT, from live or recorded audio using speech to text; profanity can be filtered.
- Speech profanity filter
Setting in speech to text that decides whether profanity in captions and transcripts is masked (the default), removed or left as is (raw).