Azure Vision API that returns several visual features in a single request: for instance tags, a caption, text recognised by Read OCR and the bounding boxes of detected objects.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Image Analysis in context, with comparison tables and the common traps.
Terms in this definition
- API
Short for application programming interface: a contract that client code calls programmatically, for example a web API secured with tokens or the Files, Images or Responses APIs.
- Document Intelligence add-on capabilities
Extra Document Intelligence options switched on with the features query parameter: barcodes, formulas, keyValuePairs, languages, ocrHighResolution, queryFields and styleFont.
- Image tagging
An Image Analysis feature producing single-word tags, each with a confidence score, for actions, scenery, objects and living things in an image.
- CRUD
Shorthand for create, read, update and delete, the four basic things you do with data. Data-plane roles in Azure Cosmos DB, for instance, authorise those operations on items.
- OCR
Optical character recognition. It reads handwritten or printed text from documents and images, making it available for analysis, something Azure Language cannot do with scanned images by itself.
Related terms
- AI enrichment
Indexing-time process in Azure AI Search where skills such as OCR, key phrase extraction, entity recognition and image analysis turn raw content into fields that can be searched.
- Alt text (Image Analysis)
Image description for accessibility, produced by Image Analysis captioning above a confidence threshold. New solutions should prefer Content Understanding or a multimodal model.
- Area of interest
Smart-crop capability of Image Analysis that, for a requested aspect ratio, returns a bounding box around the most important part of the image, useful for thumbnails and cropping.
- Azure Vision in Foundry Tools
Foundry Tool for analysing images, reading text from them (OCR) and detecting faces. Image Analysis API versions 3.2 and 4.0 are both deprecated, with retirement set for 25 September 2028.
- Image Analysis 3.2
The earlier version of Image Analysis, offering the broadest range of features: tags, objects, brands, faces, landmarks, celebrities, Describe captions, image type, colour scheme and adult content.
- Image Analysis 4.0
More recent release of Image Analysis, whose improved models handle Read, captioning (including dense captions), tagging, object and people detection and smart cropping. It is deprecated, with retirement set for 25 September 2028.
- imageAction
A parameter in indexer configuration. Setting it to
generateNormalizedImages, withdataToExtractset tocontentAndMetadata, pulls images out and normalises them into /document/normalized_images ready for OCR and image analysis. - Normalized images
Node of the enrichment tree (
/document/normalized_images/*) containing images resized and rotated during document cracking. OCR, Image Analysis and file projections accept no other image input.