Transcriber FAQ
Is my recording uploaded anywhere?
No. Your browser decodes the file and the Whisper speech model runs inside this page on your own device. The only thing ever downloaded is the model itself, once; your audio and the text it produces never leave the device.
How long does a transcription take?
It depends on the quality setting and your hardware. Measured on a recent laptop, Accurate runs at about 3.5 times real time on the processor and about 7.5 times with WebGPU, so a 30-minute interview takes roughly 4 to 9 minutes. Fast is quicker, Best is slower but more accurate. The progress bar shows the time remaining once the first section is done.
Which languages does it support?
All 99 languages Whisper was trained on, including English, Spanish, French, German, Portuguese, Italian, Romanian, Polish, Russian, Ukrainian, Arabic, Hindi, Chinese, Japanese and Korean. Accuracy is highest for widely spoken languages; for the rest, the Best setting makes the biggest difference.
Can it translate the audio into English?
Yes. Choose "Translated to English" under Output and Whisper produces English text directly from speech in any of its languages, with timestamps, so the SRT and VTT exports become English subtitles for a foreign-language video.
Is there a limit on file length or number of files?
There is no minute cap and no daily limit, because the work runs on your own device. The practical limit is memory: the whole file is decoded before transcription, so a multi-hour recording is fine on a laptop but may be too large for an older phone. Very long recordings are best split into parts there.
Why does the first run download so much?
The speech model has to be on your device to run there. Fast is about 40 MB, Accurate about 70 to 140 MB and Best about 140 to 300 MB, depending on whether your browser uses the graphics chip. It downloads once, is kept by the browser, and later transcriptions start in seconds and work offline.
Does it tell me who is speaking?
No. The transcript is split into paragraphs at pauses and sentence ends, but speakers are not identified. For interviews, paragraph breaks usually fall at the handover between speakers, which makes adding names by hand quick.
Which file formats work?
Anything your browser can play: MP3, M4A, WAV, OGG, FLAC and AAC audio, and the soundtrack of MP4, MOV and WebM video. Only the audio track is used, so large video files are fine as long as they fit in memory.