A five-step VALL-E example from text and a three-second voice prompt to speech, evaluation, and safeguards.

VALL-E: text + a 3-second voice prompt

Text-to-speech (TTS) turns text into speech. Zero-shot TTS uses a short recording to produce a voice not seen during training.

This is a fixed, offline schematic; no speech model runs on the page.