Speech to text, text to speech, translation, batch transcription and speaker recognition compared.
From Ultra Transcenders AI-901 by Tony Rough (publishing soon)
Every speech workload is defined by its direction: audio in and text out, text in and audio out, or audio in and something else (a translation or an identity) out. Recognising which direction a scenario needs is usually enough to pick the right capability.
| Capability (Azure Speech in Foundry Tools) | Direction | Scenarios |
|---|---|---|
| Speech recognition (speech to text) | Audio → text | Closed captions for recorded or live video, transcribing calls and meetings, voice assistant input |
| Speech synthesis (text to speech) | Text → audio | Reading messages aloud, public-address announcements, reading back keypad digits, a game character’s voice, audio commentary |
| Speech translation | Audio → translated output | Spoken audio in another language |
| Speaker recognition | Who is speaking (voice biometrics) | Voice-activated security key (a separate Limited Access feature) |
Common trap: calling a voice-activated security key “speech recognition” - speech recognition works out what was said; recognising whose voice it is (voice biometrics) is speaker recognition. Captioning a video, on the other hand, is speech recognition.
This note is one section of Ultra Transcenders AI-901: Microsoft Azure AI Fundamentals, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-901 glossary · All AI-901 study notes
Fairness, reliability and safety, privacy and security, inclusiveness, transparency and accountability, and how to tell them apart in a scenario.
How guardrails, system messages, grounding and user experience design reduce harm in a generative AI solution.
What happens between a prompt and a response, and how inference differs from training.
Which model setting controls randomness, which controls length and cost, and which ones are not set at deployment.
How to recognise each AI workload from a scenario.
Agents as model plus instructions, knowledge and tools, and the auto, required and none tool_choice values.
How analyzers turn documents, images, audio and video into structured JSON, and how to call them from code.