About
What this audio to text tool does
Transcribe Audio turns an audio or video recording into a readable transcript inside your browser. This page says what it does, what it leaves out, and how we check it; how well it transcribes waits until we have measured it.
What the tool does
You add an audio or video file, and the page turns the speech in it into a readable transcript inside your browser tab. The file is not uploaded: the site is a set of static files with no upload endpoint, and its security policy lets the page connect only to this site. The privacy and security page shows how to check that.
- Language: detected automatically, or the one you choose from the list.
- Model: a smaller model, the default, and a larger one, which is a bigger download and needs more memory. How the two compare will be published once they are measured on this engine.
- Paragraphs: the text is set out in paragraphs, a new one at a pause of 2 seconds or more and, in a long stretch of talk, at a sentence end.
- Timestamps: if you want them, the approximate time at the start of each paragraph.
- Search: find a word or a phrase in the transcript and step through every place it occurs.
- Copy and download: copy the whole text, or save it as a plain text (TXT) file named after your recording.
How it runs
The transcription uses an open-source speech recognition model and an engine compiled to WebAssembly, so it runs in a web page on your own device. Your browser downloads both from this site, the model in parts, the first time you add a file, and keeps the model so later files skip the download. Nothing about your file goes the other way.
Long recordings are read a slice at a time wherever the browser allows it, and the transcript appears part by part as the work goes on. While it runs, the page shows an estimate of the time left, measured on your device, rather than a promised speed. How the transcription works lists every file the page downloads and each step it takes.
What it does not do
- Tell speakers apart. There are no speaker labels: an interview or a meeting comes out as one text, and you add the names yourself.
- Translate. The transcript is in the language that is spoken in the file.
- Fetch a recording from a link. The file has to be on your device; the page does not download videos or podcasts from other sites.
- Take dictation. It works on recorded files, not on a live microphone.
- Summarize. You get the transcript itself, not notes or a summary written by a model.
- Work offline. The page needs a connection to load, and your first file downloads the engine and the model.
Measured before it is described
How accurate a transcript is, and how long it takes, depends on the recording, the language, the model, and your device. So no page on this site gives an accuracy, speed, or language figure until we have measured it on this engine. When we publish one, it will carry its date, the method, and the device it was measured on.
How we test it
An automated browser test runs against a built copy of the site with its production security headers. It transcribes short sample recordings (a WAV file, an MP4 video, and a WebM file), checks the words, the paragraphs, the timestamps, the copied text, and the TXT file, and fails if any request leaves this site or carries data. It also checks that Cancel stops a run, that an empty or silent file gets a clear message, and that a second file uses the model the browser already keeps.
Who builds it
Transcribe Audio is built and operated by Fernbrook Transcription. The site shows no ads and runs no analytics, and its pages set no cookies. The details are in the privacy policy.
Contact
Questions or suggestions: write to support@transcribe-audio.online. Please describe what you need rather than attaching a private recording.
Try it on a recording
It runs in your browser tab. Your file never leaves your device.