Turning speech in live or recorded video into timed text shown on screen; a speech to text use case that can also filter profanity and display partial results.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Captioning in context, with comparison tables and the common traps.
Terms in this definition
- Speech recognition
Turning spoken audio into text, for example call transcripts or captions.
- FILTER
Returns just those rows of a table that meet a condition. In CALCULATE it handles conditions too complex for a Boolean filter argument, though a Boolean filter is faster whenever one will work.
Related terms
- Alt text (Image Analysis)
Image description for accessibility, produced by Image Analysis captioning above a confidence threshold. New solutions should prefer Content Understanding or a multimodal model.
- Batch transcription
Asynchronous REST API that transcribes large quantities of recorded audio held in storage and delivers the text later; it is unsuitable for live captioning.
- Image Analysis 4.0
More recent release of Image Analysis, whose improved models handle Read, captioning (including dense captions), tagging, object and people detection and smart cropping. It is deprecated, with retirement set for 25 September 2028.
- Real-time transcription
Speech to text mode that converts streamed audio, from a microphone or file, into text as it is recognised, with interim results along the way. Typical uses are voice input and live captioning.