How a Content Understanding schema field is filled: extract takes the value as written (documents only), classify chooses from categories, and generate infers or summarises it.
Also called extract, classify, generate.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Field generation method in context, with comparison tables and the common traps.
Terms in this definition
- Azure Content Understanding in Foundry Tools
Foundry Tool whose generative AI analyzers turn documents, images, audio and video into structured fields by extracting, classifying or generating values.
- Schema
The middle part of a Unity Catalog name (
catalog.schema.table), grouping tables, views, volumes, functions and models inside a catalog. A grant on it covers everything in it now and later, and nothing inside can be reached withoutUSE SCHEMA. - ELT
Extract, load, transform: raw data lands in the target system first and is transformed there. See ETL for the opposite order.
Related terms
- Batch synthesis API
Asynchronous text to speech REST API suited to big batches of SSML or plain text. Unlike the real-time REST API, which stops at 10 minutes of audio, it can generate longer output such as audiobooks.
- Benign true positive
Used to classify an alert caused by activity that really happened but was approved, a penetration test for example. Instead of remediating, close the alert or risk and narrow the policy's scope.
- DSPM for AI
Purview solution securing data used by Copilot, agents and AI apps. Its one-click recommendations generate policies, for example DLP scoped to Microsoft 365 Copilot and Copilot Chat, which you then adjust in the solution that owns them rather than in DSPM.
- External model
A database object, made with CREATE EXTERNAL MODEL, that stores the details AI_GENERATE_EMBEDDINGS needs to call an AI model: endpoint, API format, model type (EMBEDDINGS), model name and credential. It is changed with ALTER, and dropping it leaves the credential in place.
- Field schema
The list of fields a Content Understanding analyzer extracts, each defined by name, type, description and method: extract, classify or generate.
- Harm categories
The four risk types used to classify content, namely violence, sexual, self-harm and hate and fairness. For both images and text, each receives a severity of safe, low, medium or high.
- Max completion tokens
Caps how many tokens a model may generate, reasoning tokens included, which limits both the length of the response and what it costs.
- Next-token prediction
Step in a transformer where a probability distribution across the vocabulary is calculated, one token is chosen and added, and the process loops to generate the output.