Improves recognition of domain vocabulary by fine-tuning a base speech model on your text, or on audio paired with transcripts. You start a project and choose the base model, then upload data and train it, and finally test and deploy.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Custom speech in context, with comparison tables and the common traps.
Terms in this definition
- Domain
A way of grouping workspaces by area of the business, in support of a data mesh approach. Items take on their workspace's domain, letting you filter the OneLake catalog by it, and certain tenant settings can be passed to domain admins; domains have no effect on access permissions.
- Fine-tuning
Continuing to train a pretrained base model on task examples, through supervised fine-tuning, DPO or RFT, so its weights shift towards a style, format or task. It neither acts as a safety control nor adds new knowledge well.
- Foundry project
A container beneath a Foundry resource that keeps one team's agents, data, files, evaluations and project connections separate. Model deployments and resource-level connections are shared from the parent, and governance is not configured at this level.
Related terms
- Audio + human-labeled transcript data
Training material for custom speech: recordings in mono 16-bit PCM WAV (RIFF) format at 8 or 16 kHz, zipped up together with a plain-text file listing each recording's name alongside what was said.
- Custom speech endpoint
Needed for real-time recognition with a custom speech model and identified by its
endpointId. Batch transcription doesn't require one, and the model behind it can be replaced without any downtime. - Custom speech model expiration
What happens once a custom speech model expires: real-time endpoint calls drop back to the newest base model for that locale, whereas batch jobs naming the model fail with a 4xx error; the model itself remains.
- Speech Studio
A web portal offering no-code Azure Speech tools, including batch speech to text, audio content creation, custom voice and custom speech.