Continuing to train a pretrained base model on task examples, through supervised fine-tuning, DPO or RFT, so its weights shift towards a style, format or task. It neither acts as a safety control nor adds new knowledge well.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Fine-tuning in context, with comparison tables and the common traps.
Terms in this definition
- DPO
Direct preference optimization, a fine-tuning technique that trains on pairs of better and worse responses without needing a reward model. Records contain input, preferred_output and non_preferred_output, and the two outputs can't be identical.
- Model parameters
The weights a model learns in training, which together hold what it knows. How many there are is what distinguishes a small language model from a large one.
- FORMAT
A DAX function that turns a value into text according to a format string, for instance "MMMM" to show a month's name. Since the output is text, numeric operations can't use it, and dynamic format strings were introduced to get around that.
- CONTROL
Granting this on a securable gives all other permissions on it too, making it the most powerful SQL permission. At database scope that includes UNMASK and ALTER ANY MASK. Warehouse access through the Admin, Member or Contributor workspace roles carries it.
Related terms
- AI Runtime
Microsoft Learn steers GPU deep learning towards this rather than Databricks Runtime ML. It provides serverless GPUs, currently in Public Preview, for fine-tuning and bespoke deep learning models.
- BOM
Bytes placed first in a text file to signal its encoding. Fine-tuning data should be saved as UTF-8 with a BOM, keeping each file below 512 MB; ASCII, UTF-16 or ISO-8859-1 files risk failing validation.
- Cognitive Services OpenAI Contributor
Azure OpenAI role that adds dataset uploads, fine-tuning and deployment creation to everything OpenAI User allows, yet can neither create resources nor see keys.
- Cognitive Services OpenAI User
The narrowest Azure OpenAI role for calling models through Entra ID; it can see the endpoint and deployments and use the playgrounds, but not deploy, view keys or upload fine-tuning data.
- Custom speech
Improves recognition of domain vocabulary by fine-tuning a base speech model on your text, or on audio paired with transcripts. You start a project and choose the base model, then upload data and train it, and finally test and deploy.
- get_openai_client()
A method on
AIProjectClientthat hands back an OpenAI client already authenticated to the project. You use it forresponses.create(readingresponse.output_text), plus conversations, evaluations and fine-tuning. - JSONL
Format in which each line holds a single JSON object. Fine-tuning training data and batch input files must use it, so CSV, TSV or one big JSON array may be rejected; ordinary REST calls don't involve it.
- Model mitigation layer
Bottom layer of generative AI risk mitigation, concerned with selecting, understanding and possibly fine-tuning the model. Swapping models does not by itself count as a safety control.