Skip to content
Convertto

Audio to Subtitles

Generate timed SRT or WebVTT subtitles from an audio recording.

Audio to Subtitles runs OpenAI's Whisper model inside your browser through WebGPU or WebAssembly. The audio is decoded locally, resampled to 16 kHz and processed in overlapping 30-second windows, producing properly timed subtitle cues wrapped to a readable width. The file is never uploaded, and after the model downloads once it works offline.

Runs in your browser
Model
Whisper, running on your own device
Offline
Works with no network after the first model download
Privacy
Runs entirely in your browser — nothing is uploaded
Cost
Free, unlimited, no sign-up

Frequently asked questions

Is my recording uploaded?

No. The model is downloaded to your browser and the audio is processed there. This is the reason to use it for an interview, a medical consultation or a legal recording — the audio never leaves the device.

Why is it slow the first time?

The model has to download before anything can run. After that it is cached by the browser and subsequent runs start immediately. Transcription itself is many times faster with WebGPU than on the CPU.

Which model should I choose?

Base is the right default. Move to small or large-v3-turbo when the audio has accents, background noise, crosstalk or technical vocabulary — the accuracy difference on difficult recordings is large, and the cost is download size and time.

How to use the audio to subtitles

  1. 1Select your audio file — the file stays on your device and is never uploaded.
  2. 2Choose the model.
  3. 3Choose the language.
  4. 4Choose the output.
  5. 5Set the characters per subtitle line.
  6. 6Press Run, then download the result when it is ready.

Embed this tool

Put the working audio to subtitles on your own site. It runs in your visitors' browsers exactly as it does here — free, no account, nothing uploaded.

Share this tool

Last updated

More audio & video tools