FREE STUDY NOTES · AI-901

Azure Speech capabilities at a glance

Speech to text, text to speech, translation, batch transcription and speaker recognition compared.

From Ultra Transcenders AI-901 by Tony Rough (publishing soon)

Every speech workload is defined by its direction: audio in and text out, text in and audio out, or audio in and something else (a translation or an identity) out. Recognising which direction a scenario needs is usually enough to pick the right capability.

Capability (Azure Speech in Foundry Tools) Direction Scenarios
Speech recognition (speech to text) Audio → text Closed captions for recorded or live video, transcribing calls and meetings, voice assistant input
Speech synthesis (text to speech) Text → audio Reading messages aloud, public-address announcements, reading back keypad digits, a game character’s voice, audio commentary
Speech translation Audio → translated output Spoken audio in another language
Speaker recognition Who is speaking (voice biometrics) Voice-activated security key (a separate Limited Access feature)

Common trap: calling a voice-activated security key “speech recognition” - speech recognition works out what was said; recognising whose voice it is (voice biometrics) is speaker recognition. Captioning a video, on the other hand, is speech recognition.

Get the whole book

This note is one section of Ultra Transcenders AI-901: Microsoft Azure AI Fundamentals, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.

Amazon.co.ukKindle: coming soonPaperback: coming soon
Amazon.comKindle: coming soonPaperback: coming soon

Publishing soon on Amazon in Kindle and paperback editions.

About the book · Free AI-901 glossary · All AI-901 study notes

More AI-901 study notes