Search spaces, sampling methods and early termination policies.
From Ultra Transcenders AI-300 by Tony Rough (publishing soon)
When you choose the algorithm yourself, a sweep job tunes its hyperparameters by running the same training many times over a search space. The choices that matter are the sampling method, the early termination policy and the limits that bound cost.
command_job(param=<expression>) and then .sweep(...), or with a type: sweep YAML.primary_metric must exactly match the name the script logs (for example mlflow.log_metric("accuracy", ...)). goal is maximize or minimize.Choice, plus QUniform, QLogUniform, QNormal and QLogNormal.Uniform, LogUniform, Normal and LogNormal.| Sampling | Supports | Use when |
|---|---|---|
Random (optionally rule="sobol" with a seed) |
Discrete and continuous; early termination | Starting out or finding promising regions |
| Grid | choice only; early termination |
You can afford every combination |
| Bayesian | choice, uniform, quniform |
Enough budget: total trials of at least 20 × the number of hyperparameters. Lower concurrency converges better because each trial learns from completed ones |
Common trap: Using grid sampling over a continuous range such as
Uniform(0.01, 0.1)- grid sampling takeschoicevalues only; use random or Bayesian sampling for continuous distributions.
evaluation_interval counts metric reports, and delay_evaluation skips the first N intervals.
| Policy | Ends a trial when |
|---|---|
| Bandit | Its metric is outside slack_factor (ratio) or slack_amount (absolute) of the best trial. With best 0.8 and slack_factor=0.2, trials under 0.8/1.2 ≈ 0.67 stop |
| Median stopping | Its best metric is worse than the median of the running averages of all trials |
| Truncation selection | It’s in the lowest truncation_percentage (1–99) at that interval |
| None (default) | Never; all trials run to completion |
Median stopping with evaluation_interval=1 and delay_evaluation=5 is the conservative choice (about 25–35% savings with no loss in the primary metric). Bandit with small slack or truncation with a high percentage saves more aggressively.
max_total_trials is 1–1000.max_concurrent_trials defaults to max_total_trials, subject to the compute’s capacity.timeout applies to the whole sweep, in seconds; trial_timeout applies to each trial.max_total_trials and timeout is reached first ends the sweep.az ml job download --name <sweep> --output-name model or ml_client.jobs.download.This note is one section of Ultra Transcenders AI-300: Operationalizing Machine Learning and Generative AI Solutions, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · Free AI-300 glossary · All AI-300 study notes
When to use a workspace, a registry, models, environments, components and data assets.
uri_file, uri_folder and mltable data assets, and how mltable.load() finds the MLTable file.
Experiments, runs, parameters, metrics, artifacts and autologging.
Blue-green deployments, traffic splitting, mirroring and instant rollback.
Where prompts are processed for each deployment type, and which one meets data residency needs.
When to use provisioned deployments, how 429s and spillover work, and how PTUs are billed.
Groundedness, relevance, coherence, fluency and risk and safety evaluators, and what each needs.