uri_file, uri_folder and mltable data assets, and how mltable.load() finds the MLTable file.
From Ultra Transcenders AI-300 by Tony Rough (publishing soon)
A data asset is a named, versioned reference to data, built on a URI that points at a datastore or storage endpoint. The URI scheme must match the storage endpoint, and the asset type must match what the path points at.
| Scheme / type | Meaning |
|---|---|
wasbs://account.blob.core.windows.net/container/... |
Blob Storage (blob endpoint) |
abfss://...dfs.core.windows.net/... |
ADLS Gen2 (dfs endpoint) |
azureml:// |
Registered datastore |
AssetTypes.URI_FOLDER |
Folder or container (many files) |
AssetTypes.URI_FILE |
One file |
AssetTypes.URI_MLTABLE |
Folder containing an MLTable file |
Common trap: Using
abfssfor a data asset over all the blobs in a container -abfssneeds the ADLS Gen2 dfs endpoint; for the blob endpoint, usewasbswithURI_FOLDER.
mltable.load("./iris"), then call to_pandas_dataframe() to get the data. save() writes a blueprint, take() returns the first rows, and take_random_sample() returns a random sample.Common trap: Passing the MLTable file itself to
load(), as inmltable.load("./sample_data/MLTable")-load()takes the folder that contains the MLTable file, not the file.
This note is one section of Ultra Transcenders AI-300: Operationalizing Machine Learning and Generative AI Solutions, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-300 glossary · All AI-300 study notes
When to use a workspace, a registry, models, environments, components and data assets.
Experiments, runs, parameters, metrics, artifacts and autologging.
Search spaces, sampling methods and early termination policies.
Blue-green deployments, traffic splitting, mirroring and instant rollback.
Where prompts are processed for each deployment type, and which one meets data residency needs.
When to use provisioned deployments, how 429s and spillover work, and how PTUs are billed.
Groundedness, relevance, coherence, fluency and risk and safety evaluators, and what each needs.