Auto caption generator FAQ
Is my video uploaded to make the captions?
No. The soundtrack is decoded by your browser and transcribed by the Whisper speech model running in this page on your own device, and the captioned video is recorded locally too. The only download is the speech model itself, once.
Is there a watermark or a length limit?
No watermark, no sign-up and no minute cap. The practical limits are your device's memory, because the whole soundtrack is decoded before transcription, and time, because exporting the captioned video happens in real time.
How long does it take?
Transcription runs at roughly 3.5 to 7.5 times real time with the Accurate setting on a recent laptop, so a 5-minute video is transcribed in about one to two minutes. Exporting the captioned video then takes as long as the video itself, because it is recorded as it plays.
Can I edit the captions?
Yes. Every caption line is editable: fix a name or a misheard word and the new text is spread across the same time span. You can also delete a caption. Switching the style regroups the lines but keeps your corrections.
Which languages does it support?
All 99 languages Whisper knows, detected automatically from the first stretch of speech, or chosen by you. You can also have the captions translated into English directly from the speech.
Can I get an SRT file for YouTube instead of burned-in captions?
Yes. SRT and VTT exports are instant and work with YouTube, Vimeo, Premiere, DaVinci Resolve, CapCut and most HTML5 players. They keep captions switchable and searchable, which burned-in captions are not.
Why is the exported video WebM on some browsers?
The browser records the video, and each browser supports different formats. Recent Chrome and Edge record MP4 when possible, Safari records MP4, and Firefox records WebM. Instagram, TikTok and YouTube accept both.
What if the captions are slightly early or late?
Timing comes from Whisper's word alignment, which is usually within a fifth of a second. Captions are shown from the moment their first word starts and held long enough to read, so small offsets are rarely noticeable; the karaoke style shows alignment most clearly, so check it there first.