The Azure Speech client library, installed in Python as azure-cognitiveservices-speech and imported as speechsdk, providing SpeechSynthesizer, SpeechRecognizer, SpeechConfig and classes for audio configuration.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Speech SDK in context, with comparison tables and the common traps.
Terms in this definition
- Azure Speech in Foundry Tools
Foundry Tool for spoken language, able to transcribe audio into text, synthesise voice from text, translate what is said and run live voice conversations.
- SpeechSynthesizer
The Speech SDK class used for text to speech, offering SpeakTextAsync and SpeakSsmlAsync. Created from a SpeechConfig plus an AudioOutputConfig, it sends audio to a stream, file or speaker.
- SpeechRecognizer
The Speech SDK class used for speech to text. It is created from a SpeechConfig plus an AudioConfig (stream, WAV file or microphone) and returns transcriptions either once or continuously.
Related terms
- AudioConfig
Object in the Speech SDK specifying where audio comes from or goes to, for example FromStreamInput, FromWavFileInput or FromDefaultMicrophoneInput.
- AudioOutputConfig
Speech SDK class choosing a single destination for synthesised speech: a file, a named device, a custom output stream or the default speaker.
- AudioStreamFormat
Speech SDK class describing how an audio stream is encoded, namely channels, bits and sample rate. It does not act as an output destination.
- Compressed audio input
Speech SDK ability to accept MP3, OPUS/OGG, FLAC, ALAW and MULAW through push or pull streams, with GStreamer doing the decoding;
GetWaveFormatPCMapplies to uncompressed audio instead. - GStreamer
Open-source multimedia library that the Speech SDK relies on to decode compressed audio streams into PCM. It is not bundled, so you install it yourself and add it to the path.
- Intent recognition (Speech)
Retired on 30 September 2025, this Speech SDK feature mapped spoken utterances to intents; use speech to text and then CLU or an Azure OpenAI model instead.
- KeywordRecognizer
Class in the Speech SDK that runs on the device, waits for a particular custom keyword defined by a .table model and fires an event on detection. It isn't meant for general transcription.
- recognize_once()
Single-shot recognition call on the Speech SDK's SpeechRecognizer. It returns one utterance, stopping at silence or a time limit of around 15 to 30 seconds, which makes it good for short commands.