Free · No upload · No sign-up · No minute limits

Transcribe audio & video to text

Drop an interview, lecture, podcast, voice memo or video and get the words back with timestamps, in 99 languages. The speech model runs on your own device — the recording never leaves it.

Quality

Drop audio or video here

MP3, M4A, WAV, OGG, FLAC, MP4, MOV, WebM — or choose a file

🔒 Transcribed on this device. The first run downloads the speech model once (about 40–140 MB); after that it works offline.

What people transcribe with it

🎤

Interviews and research

Journalists, students and researchers turn recorded interviews into searchable text — privately, which matters when a source or participant was promised confidentiality.

🎓

Lectures and meetings

A two-hour lecture or a recorded call becomes notes you can skim and search, with timestamps that jump straight back to the moment in the recording.

🎬

Podcasts and video

Show notes, blog posts, quotes for social, and SRT or VTT subtitle files for YouTube — straight from the episode or the video file.

Why this transcriber runs on your device

Almost every online transcription service works the same way: you upload the recording, it is processed on their servers, and the free tier is rationed by minutes or files per day because server time costs money. That model has two problems for a lot of recordings. The first is privacy — an interview, a therapy-adjacent conversation, a board meeting or a medical appointment is exactly the kind of audio you should not be sending to a company you know nothing about. The second is the limit, which always seems to run out on the one long recording that matters.

This tool takes the other route. It downloads OpenAI's open-source Whisper speech model into your browser once, and from then on every transcription runs on your own processor or graphics chip. The recording is decoded by your browser and handed to the model inside the page; there is no server to upload it to. Because the work happens on hardware you already own, there is nothing to ration either: no minute cap, no file cap, no account.

Choosing Fast, Accurate or Best

SettingModelOne-time downloadSpeed (measured)Use it for
FastWhisper tiny≈ 40 MB≈ 6× real time on a laptop CPUClear speech, quick drafts, older phones
AccurateWhisper base≈ 70–140 MB≈ 3.5× on CPU, ≈ 7.5× with WebGPUThe default — interviews, meetings, most languages
BestWhisper small≈ 140–300 MB≈ 4.5× with WebGPUAccents, noisy rooms, non-English audio

"4× real time" means an hour of audio takes about fifteen minutes. Best needs WebGPU — current Chrome and Edge on desktop, and recent Safari — because the larger model is too slow on a single processor thread to be worth offering; where WebGPU is missing, Best quietly runs as Accurate. The download happens once per quality level and is then kept by the browser, so the second transcription starts in seconds and works without an internet connection.

Getting a cleaner transcript

  • Set the language if you know it. Auto-detect listens to the first stretch of speech and picks the language from it, which is reliable for a single-language recording. If the recording opens with music, a foreign-language intro or silence, choosing the language yourself avoids a wrong guess.
  • Use Best for accents and noise. The small model is noticeably better on strong accents, overlapping background sound and languages other than English.
  • Silence is skipped on purpose. Whisper is known to invent phrases like "Thank you for watching" over silence, so stretches with no sound at all are not sent to the model.
  • Speakers are not labelled. The transcript is split into paragraphs at pauses and sentence breaks, but it does not tell you who said what.

What you can export

TXT is the transcript as paragraphs, with or without a timestamp at the start of each one — the format for notes, quotes and articles. SRT and VTT are subtitle files: short timed captions that YouTube, Vimeo, most video editors and HTML5 players accept directly. Every word also carries its own timing inside the page, which is why clicking a word jumps the player to it — handy for checking a quote against the recording before you publish it.

Transcriber FAQ

Is my recording uploaded anywhere?

No. Your browser decodes the file and the Whisper speech model runs inside this page on your own device. The only thing ever downloaded is the model itself, once; your audio and the text it produces never leave the device.

How long does a transcription take?

It depends on the quality setting and your hardware. Measured on a recent laptop, Accurate runs at about 3.5 times real time on the processor and about 7.5 times with WebGPU, so a 30-minute interview takes roughly 4 to 9 minutes. Fast is quicker, Best is slower but more accurate. The progress bar shows the time remaining once the first section is done.

Which languages does it support?

All 99 languages Whisper was trained on, including English, Spanish, French, German, Portuguese, Italian, Romanian, Polish, Russian, Ukrainian, Arabic, Hindi, Chinese, Japanese and Korean. Accuracy is highest for widely spoken languages; for the rest, the Best setting makes the biggest difference.

Can it translate the audio into English?

Yes. Choose "Translated to English" under Output and Whisper produces English text directly from speech in any of its languages, with timestamps, so the SRT and VTT exports become English subtitles for a foreign-language video.

Is there a limit on file length or number of files?

There is no minute cap and no daily limit, because the work runs on your own device. The practical limit is memory: the whole file is decoded before transcription, so a multi-hour recording is fine on a laptop but may be too large for an older phone. Very long recordings are best split into parts there.

Why does the first run download so much?

The speech model has to be on your device to run there. Fast is about 40 MB, Accurate about 70 to 140 MB and Best about 140 to 300 MB, depending on whether your browser uses the graphics chip. It downloads once, is kept by the browser, and later transcriptions start in seconds and work offline.

Does it tell me who is speaking?

No. The transcript is split into paragraphs at pauses and sentence ends, but speakers are not identified. For interviews, paragraph breaks usually fall at the handover between speakers, which makes adding names by hand quick.

Which file formats work?

Anything your browser can play: MP3, M4A, WAV, OGG, FLAC and AAC audio, and the soundtrack of MP4, MOV and WebM video. Only the audio track is used, so large video files are fine as long as they fit in memory.