How it works
Whisper in your browser: how the transcription runs
The page runs an open speech recognition model, Whisper, inside your browser tab. Here is what it downloads, how it reads a long file without holding all of it, how the paragraphs are made, and why the time differs from one device to the next.
Three steps, all inside your browser
1. Your browser reads the file
Where your browser can, a worker in your tab takes the sound track a 30-second slice at a time, and turns it into the 16 kHz mono sound the model expects.
2. Whisper transcribes it in windows
The engine, whisper.cpp compiled to WebAssembly, transcribes about three minutes of sound at a time and hands each window's text to the page.
3. The page sets out the text
The words are grouped into paragraphs at pauses, with a timestamp for each if you want it, ready to search, copy, or save as TXT.
The model: Whisper, base or small
Whisper is a speech recognition model that OpenAI published in 2022 with its code and weights under the MIT license. It listens to 30 seconds of sound at a time and writes out the words, with a time for each segment. This site offers two sizes of it:
- The smaller model (the default) is Whisper base, stored with 5 bits per weight. It is the lighter download.
- The larger model is Whisper small, stored with 8 bits per weight: a bigger download that needs more memory.
Storing weights with fewer bits (quantization) shrinks a model at some cost in accuracy. How the two sizes compare on this engine, per language, is not published yet: we measure before we describe. The language menu lists the languages Whisper can be set to; with automatic detection, the engine picks the language in the first window that has speech and keeps it for the rest of the file.
The engine: whisper.cpp compiled to WebAssembly
The model runs in whisper.cpp, an open-source C/C++ version of Whisper, which we compile to WebAssembly (a program format every current browser can run) with a small interface of our own. There are three builds, and the page tries them in this order:
- WebGPU, when your browser offers a graphics adapter with the features the engine needs. We have not measured this build on a real graphics card yet, so we make no claim about it.
- WebAssembly on several threads, one per logical core your browser reports. Threads need the tab to be cross-origin isolated, which the site's headers provide.
- WebAssembly on one thread, when threads are not available or your browser reports two logical cores or fewer. It gives the same kind of result.
If a build fails to start, or the WebGPU build fails during a run, the page restarts the engine on the next build and carries on. The line above the transcript names the build that did the work.
What the page downloads, and when
Opening the page downloads only the page and its own files. The engine and a model come from this site when you add your first file:
| File | Size | When |
|---|---|---|
| Engine, multi-threaded or single-thread build (one of the two) | about 1.5 MB | Your first file |
| Smaller model (Whisper base, 5-bit), the default | 59.7 MB in 3 parts | Your first file with it |
| Larger model (Whisper small, 8-bit) | 264.5 MB in 13 parts | Only if you choose it |
| Voice detector (Silero VAD) | 0.9 MB | Your first file |
| Media reader (Mediabunny), part of the page's scripts | about 0.3 MB | Your first file |
| Engine, WebGPU build | about 3.4 MB | Only if the browser offers a usable WebGPU adapter |
The models come in parts of at most 20 MiB, because the host serves no single file over 25 MiB. Before the engine uses a part, your browser checks it against its SHA-256 fingerprint, so a damaged download is refused instead of producing a wrong transcript. The verified parts stay in your browser's Cache Storage when it has room for them, so your next file, and your next visit, skip the download.
Why your recording is never uploaded
Every step runs inside your tab, the site has no endpoint that could accept a file, and the page is allowed to connect only to this site. The check that your recording is not uploaded is on the privacy and security page, with the steps to watch the requests yourself.
How a long file is read and transcribed
A one-hour recording is about 115 MB of sound once it is decoded to 16 kHz mono (3,600 seconds x 16,000 samples x 2 bytes, or twice that as the 32-bit numbers the engine works with). So, wherever the browser allows it, the page does not decode the whole file at once. A second worker opens the file with Mediabunny and decodes its sound track with your browser's own decoder, 30 seconds at a time, and hands over the next slice only when the engine is ready for it.
The slices are joined into windows of about three minutes. Each window ends at the quietest half second in its last 15 seconds, so a cut rarely falls inside a word, and the windows follow each other with no gap and no overlap. A small voice detector marks the stretches with speech in each window, so the model skips silence instead of writing words over it. Each window is transcribed with the end of the text before it as a prompt, which helps a sentence that runs across a cut, and its times are placed back on the file's own clock.
When a window is done, its paragraphs appear on the page, so you can start reading and searching while the rest runs. Where your browser has no decoder for a file's sound, the page falls back to decoding the whole file at once, and then refuses a file over 500 MB or longer than 1 hour.
How the paragraphs and timestamps are made
Whisper returns segments of text with a start and an end time, and a time for each word. The page joins them into paragraphs: a new one starts at a pause of 2 seconds or more, measured from the times of the words on each side, and a paragraph that grows past about 600 characters ends at the next sentence end. There are no speaker labels, so a paragraph break marks a pause or a long stretch of speech, never a change of voice.
The timestamp of a paragraph is the time its first word starts, written as [hh:mm:ss] and rounded down to the second. It is approximate: good for finding a passage in the recording, not for subtitles. Copy text and Download TXT write the same paragraphs, with a blank line between them and the timestamps when they are shown.
Why the time varies
The work happens on your processor, so the same file takes a different time on a new laptop and on an old one, and the model and the length of the recording change it too. That is why the page promises no speed. Once the engine reports its first progress, the page measures how many seconds of work each second of sound takes on your device, smooths that over the windows so one slow moment does not swing it, and shows the time left.
What your device needs
- A current browser with WebAssembly. Chrome, Edge, Firefox, and Safari all have it; our automated test runs in Chromium.
- Enough free memory for the model and one window of sound. The larger model needs more; we will publish measured figures once we have them.
- A connection for the first file. The model is kept in the browser's storage when there is room for it, so later visits skip the download.
The open-source components of the engine, the models, and the media reader are listed with their licenses on the credits page.
Questions about the engine
Is this the same Whisper that OpenAI published?
It is the same model: OpenAI released Whisper's code and weights under the MIT license. This page runs those weights in the file format of the open-source whisper.cpp project, quantized (stored with fewer bits per weight) so the download is smaller. The engine is whisper.cpp compiled to WebAssembly, not OpenAI's own Python code, and no request goes to OpenAI.
Does it use my graphics card?
It tries. When your browser offers WebGPU with the features the engine needs, the page starts the WebGPU build first, and if that fails it moves to the CPU builds on its own. Our test machines have no graphics card, so the WebGPU build has only been checked for falling back cleanly. The line above the transcript names the build that ran.
Why did the model download again?
Your browser keeps the model in its Cache Storage for this site until you clear the site's data, or until the browser clears it to free space. After that, the next file downloads it again. What your browser keeps and how to remove it has the details.
What happens when I press Cancel?
The page stops the engine at once by closing the worker it runs in, and drops the paragraphs made so far. Your next file starts a fresh engine and takes the model from your browser's storage, so the model does not download again.
Sources
- OpenAI, Whisper repository: The model's code, its MIT license, and the table of model sizes (base and small).
- Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision (2022): The paper that introduces Whisper.
- whisper.cpp: The C/C++ port of Whisper that the engine compiles to WebAssembly (version 1.9.4), and its quantized model files.
- Silero VAD: The voice activity detector that finds the stretches with speech.
- Mediabunny: The media library that reads the sound track from your file in slices.
- MDN, AudioDecoder (WebCodecs): The browser's own decoder that turns the compressed sound into samples.
- MDN, CacheStorage: Where the verified model parts are kept between visits.
- web.dev, Making your website cross-origin isolated using COOP and COEP: Why the site sends these headers: WebAssembly threads need SharedArrayBuffer.
- Cloudflare Workers, Limits: The 25 MiB limit per static file, the reason the models come in parts.
Try it on a recording
The transcription runs in your browser. Your file never leaves your device.