Skip to content
Convertto

Video to Text Transcriber

Pull a transcript out of a video's speech without uploading the file.

Video to Text Transcriber runs OpenAI's Whisper model inside your browser through WebGPU or WebAssembly. The audio track is decoded out of the video locally, resampled to 16 kHz and processed in overlapping 30-second windows, producing a transcript with segment timings. The file is never uploaded, and after the model downloads once it works offline.

Runs in your browser
Model
Whisper, running on your own device
Offline
Works with no network after the first model download
Privacy
Runs entirely in your browser — nothing is uploaded
Cost
Free, unlimited, no sign-up

Frequently asked questions

Is my recording uploaded?

No. The model is downloaded to your browser and the audio is processed there. This is the reason to use it for an interview, a medical consultation or a legal recording — the audio never leaves the device.

Why is it slow the first time?

The model has to download before anything can run. After that it is cached by the browser and subsequent runs start immediately. Transcription itself is many times faster with WebGPU than on the CPU.

Which model should I choose?

Base is the right default. Move to small or large-v3-turbo when the audio has accents, background noise, crosstalk or technical vocabulary — the accuracy difference on difficult recordings is large, and the cost is download size and time.

How to use the video to text transcriber

  1. 1Select your video file — the file stays on your device and is never uploaded.
  2. 2Choose the model.
  3. 3Choose the language.
  4. 4Choose the output.
  5. 5Set the characters per subtitle line.
  6. 6Press Run, then download the result when it is ready.

Embed this tool

Put the working video to text transcriber on your own site. It runs in your visitors' browsers exactly as it does here — free, no account, nothing uploaded.

Share this tool

Last updated

More audio & video tools