Skip to content
Convertto

Transcribe Audio in Your Browser — Nothing Uploaded

Run Whisper locally to turn speech into text, SRT, VTT or a timestamped transcript. Works offline after the first run.

This transcriber runs OpenAI's Whisper model in your browser using WebGPU or WebAssembly. Audio is decoded locally, resampled to 16 kHz and processed in 30-second windows, producing a transcript with segment timings that can be exported as plain text, SRT, WebVTT or JSON. The file is never uploaded, and after the first model download the tool works offline.

Runs in your browser
Uploads
None — the audio never leaves your device
Languages
99 on every model except tiny, which is English only
Best accuracy
Whisper large-v3-turbo — a 1 GB download, cached after the first run
Offline
Works with no network after the first model download
Privacy
Runs entirely in your browser — nothing is uploaded
Cost
Free, unlimited, no sign-up

Frequently asked questions

Is my recording uploaded?

No. The file is decoded in your browser and the model runs on your device. This is the whole reason the tool exists — the recordings people most need transcribed are usually the ones they are not allowed to upload anywhere.

How long does it take?

With WebGPU, roughly real-time or better on the base model — a ten-minute recording in a few minutes. On the WebAssembly fallback expect several times the audio duration. The first run also includes the model download.

How accurate is it compared with a paid service?

At the large-v3-turbo setting this runs the same family of model most hosted services do, so on clear speech the gap largely closes. The hosted services still win on diarisation, on very noisy audio, and on turnaround for long files — but the model here is no longer the small end of the family. Move up a size before concluding the audio is the problem.

Can it tell speakers apart?

No. Whisper transcribes speech but does not do diarisation, so a two-person interview comes back as one continuous transcript. The segment timings make it straightforward to label afterwards.

How to use the ai audio transcription

  1. 1Select your audio or video — the file stays on your device and is never uploaded.
  2. 2Choose the model.
  3. 3Choose the language.
  4. 4Choose the task.
  5. 5Choose the output.
  6. 6Set the characters per subtitle line.
  7. 7Press Run, then download the result when it is ready.

Sources & specifications

Embed this tool

Put the working ai audio transcription on your own site. It runs in your visitors' browsers exactly as it does here — free, no account, nothing uploaded.

Share this tool

Last updated

More on-device ai