The Speech SDK class used for speech to text. It is created from a SpeechConfig plus an AudioConfig (stream, WAV file or microphone) and returns transcriptions either once or continuously.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains SpeechRecognizer in context, with comparison tables and the common traps.
Terms in this definition
- Speech SDK
The Azure Speech client library, installed in Python as azure-cognitiveservices-speech and imported as speechsdk, providing SpeechSynthesizer, SpeechRecognizer, SpeechConfig and classes for audio configuration.
- Speech recognition
Turning spoken audio into text, for example call transcripts or captions.
- AudioConfig
Object in the Speech SDK specifying where audio comes from or goes to, for example FromStreamInput, FromWavFileInput or FromDefaultMicrophoneInput.
Related terms
- recognize_once()
Single-shot recognition call on the Speech SDK's SpeechRecognizer. It returns one utterance, stopping at silence or a time limit of around 15 to 30 seconds, which makes it good for short commands.
- start_continuous_recognition()
Suits long audio with many utterances: this SpeechRecognizer method raises recognizing, recognized and canceled events with results, carrying on until you call stop_continuous_recognition().
- start_keyword_recognition()
Method of SpeechRecognizer that uses a keyword model to listen on the device for a wake word, after which audio starts streaming to the service. It is not meant for general transcription.