Auto captions are subtitles generated automatically from a video's audio by speech recognition, or from the script that produced the audio, rather than typed and timed by hand.
There are two ways to get them, and they fail differently.
Speech recognition listens to the finished track and writes what it hears. It works for any video, including one you filmed. It also makes the mistakes recognition always makes: names, acronyms, jargon and any word spoken over music.
Every faceless niche has a vocabulary that trips it. A history channel gets place names wrong, a finance channel gets tickers wrong. Those words sit on screen, large, for everyone to read.
When the audio came from text to speech, the words are already known. The model returns the timing of each one, and the captions are built from that rather than guessed. No recognition step, no misheard words, and the highlight lands on the spoken syllable.
This is the better path whenever it is available, and it is why script-first pipelines produce cleaner captions than recording first and transcribing afterwards.
Auto captions are a first draft that is right most of the time. The ten seconds it takes to scan them is the difference between a professional video and one with a misspelled name in 90-point type.
MakeViral builds captions from the voice model's character timings. See the AI voiceover video generator.
The narrator-over-gameplay format: a written script becomes a voiced, captioned vertical video with a background clip, in one pass.
Open itSubtitle burn-in means rendering captions permanently into the video frames, so the words are part of the picture and cannot be switched off, moved or restyled by the player.
Read moreGlossaryText to speech, or TTS, is the technology that converts written text into spoken audio, producing a narration track from a script without anyone recording it.
Read morePick a format, paste a prompt or a product URL, and download a 9:16 video with a voiceover and timed captions.
Cancel anytime. 14-day refund window on unused credits.