Text to speech feature with limited access that produces a distinctive synthetic voice, for a brand or character, trained on recordings of a real speaker.
Also called custom neural voice, brand voice.
Read more: Microsoft Learn
In the Ultra Transcenders books
Each book explains Custom voice in context, with comparison tables and the common traps.
Terms in this definition
- Speech synthesis
Turning text into spoken audio that sounds natural, for example to read messages out loud.
- Limited Access
Under this Microsoft policy, sensitive AI features such as custom neural voice or Face identification and verification can only be used after registering and being approved as a managed customer with an approved use case.
Related terms
- Consent statement
Before a custom voice is built, the voice talent records set wording giving permission for their voice to be used; the .wav or .mp3 must match the training data's language and recording conditions.
- Custom voice endpoint
Synthesises speech from a custom voice you have deployed, at voice.speech.microsoft.com/cognitiveservices/v1?deploymentId=...; it offers no way to list available voices.
- Professional voice fine-tuning
Azure Speech feature with limited access, created in Speech Studio or the Foundry portal, that builds a bespoke brand voice from the recordings and consent statement of a consenting voice talent.
- Speech Studio
A web portal offering no-code Azure Speech tools, including batch speech to text, audio content creation, custom voice and custom speech.