Text to speech, or TTS, is the technology that converts written text into spoken audio, producing a narration track from a script without anyone recording it.
TTS is the engine behind an AI voiceover. The two terms get used interchangeably, but they are different levels of the same thing: TTS is the underlying conversion, and a voiceover is what you do with it in a video.
The recognisable platform voices carry a format signal of their own. Viewers know what kind of video they are about to watch before a word finishes.
Script, then TTS, then captions built from the returned timings, then the render. Getting that order right is what keeps the words on screen in sync with the voice. See auto captions.
MakeViral runs this pipeline end to end from a prompt. See the AI voiceover video generator.
The narrator-over-gameplay format: a written script becomes a voiced, captioned vertical video with a background clip, in one pass.
Open itAn AI voiceover is narration produced by a speech model from a written script, used in place of a recorded human voice, and usually returned with per-word timing that captions can be built from.
Read moreGlossaryAuto captions are subtitles generated automatically from a video's audio by speech recognition, or from the script that produced the audio, rather than typed and timed by hand.
Read morePick a format, paste a prompt or a product URL, and download a 9:16 video with a voiceover and timed captions.
Cancel anytime. 14-day refund window on unused credits.