Free · No watermark · No upload · No sign-up

Auto caption generator

Drop a video and get captions made automatically — classic subtitles, bold TikTok-style words or karaoke highlighting. Fix any line, then export the captioned video or an SRT file. Everything runs on your own device.

Quality

Drop a video here

MP4, MOV, WebM — phone videos, Reels, interviews, talks — or choose a file

🔒 Your video stays on this device. The first run downloads the speech model once (about 40–140 MB).

What people caption with it

📱

Reels, Shorts and TikToks

Most people scroll with the sound off. Bold word-by-word captions keep a talking-head clip watchable on mute — the style that took over short-form video.

🎤

Interviews, talks and courses

Classic two-line subtitles for longer videos, plus an SRT or VTT file for YouTube, Vimeo or your own player, so captions stay searchable and switchable.

🌍

Reaching another language

Caption a video in the language that was spoken, or have Whisper translate the speech straight into English subtitles.

Why captions matter more than they used to

Social feeds autoplay video without sound, and most people never tap to turn it on. A talking-head clip without captions is, for the majority of the people it reaches, a person moving their lips. Captions turn that same clip into something that works silently — and they help everyone else too: viewers in a noisy place, people who are deaf or hard of hearing, and anyone watching in a second language.

Captioning by hand is slow, which is why automatic captioning became a standard feature of every video app. The catch is that almost all of them either upload your footage to process it, stamp a watermark on the free export, or both. This generator does neither: the speech recognition runs in your browser on your own hardware, and the export is a clean file.

Picking a caption style

StyleLooks likeBest for
SubtitlesTwo lines in a dark box at the bottomInterviews, talks, tutorials, anything longer than a minute
BoldOne to three heavy words, outlined, lower middleReels, Shorts and TikToks — the short-form look
KaraokeA short phrase with the spoken word lit upHooks, quotes and music-backed clips where rhythm matters
MinimalClean text with a soft shadow, no boxCinematic footage where a box would get in the way

For vertical video, keep captions out of the bottom fifth of the frame — that is where Instagram, TikTok and YouTube Shorts draw their own buttons and description. "Lower middle" places them above that zone automatically.

How accurate is it, and how do I fix mistakes?

The captions come from OpenAI's Whisper model, the same family of speech recognition used by many paid tools, running at three sizes: Fast, Accurate and Best. On clear speech in a common language the text is usually right word for word; names, jargon and heavy background noise are where it slips. Every caption is editable — change a word and the new text is re-timed across the same span, so the rhythm of the captions stays intact. Switching style afterwards keeps your fixes.

Each word carries its own timing, which is what makes karaoke highlighting and short bold captions possible: the tool knows when each word starts, not just each sentence.

How the captioned video is made

The soundtrack is decoded by your browser, resampled and transcribed by Whisper in a background worker. The export then plays the video through a canvas, drawing each frame with the captions on top, and records that canvas together with the original audio. That recording happens in real time, so a two-minute video takes about two minutes to export — keep the tab in front while it runs. Chrome and Edge produce MP4 where they can and WebM otherwise; Safari produces MP4. If you only need captions for YouTube or an editor, the SRT and VTT exports are instant.

Auto caption generator FAQ

Is my video uploaded to make the captions?

No. The soundtrack is decoded by your browser and transcribed by the Whisper speech model running in this page on your own device, and the captioned video is recorded locally too. The only download is the speech model itself, once.

Is there a watermark or a length limit?

No watermark, no sign-up and no minute cap. The practical limits are your device's memory, because the whole soundtrack is decoded before transcription, and time, because exporting the captioned video happens in real time.

How long does it take?

Transcription runs at roughly 3.5 to 7.5 times real time with the Accurate setting on a recent laptop, so a 5-minute video is transcribed in about one to two minutes. Exporting the captioned video then takes as long as the video itself, because it is recorded as it plays.

Can I edit the captions?

Yes. Every caption line is editable: fix a name or a misheard word and the new text is spread across the same time span. You can also delete a caption. Switching the style regroups the lines but keeps your corrections.

Which languages does it support?

All 99 languages Whisper knows, detected automatically from the first stretch of speech, or chosen by you. You can also have the captions translated into English directly from the speech.

Can I get an SRT file for YouTube instead of burned-in captions?

Yes. SRT and VTT exports are instant and work with YouTube, Vimeo, Premiere, DaVinci Resolve, CapCut and most HTML5 players. They keep captions switchable and searchable, which burned-in captions are not.

Why is the exported video WebM on some browsers?

The browser records the video, and each browser supports different formats. Recent Chrome and Edge record MP4 when possible, Safari records MP4, and Firefox records WebM. Instagram, TikTok and YouTube accept both.

What if the captions are slightly early or late?

Timing comes from Whisper's word alignment, which is usually within a fifth of a second. Captions are shown from the moment their first word starts and held long enough to read, so small offsets are rarely noticeable; the karaoke style shows alignment most clearly, so check it there first.