Speech to text mode that converts streamed audio, from a microphone or file, into text as it is recognised, with interim results along the way. Typical uses are voice input and live captioning.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Real-time transcription in context, with comparison tables and the common traps.
Terms in this definition
- Speech recognition
Turning spoken audio into text, for example call transcripts or captions.
- Captioning
Turning speech in live or recorded video into timed text shown on screen; a speech to text use case that can also filter profanity and display partial results.