# Text to Speech — Generated On Your Device, Downloadable as WAV

> Turn text into natural speech with a neural voice running in your browser, and download the audio as a WAV file.

This tool synthesises speech from text using Kokoro, an 82-million-parameter neural voice model that runs in your browser on WebGPU or WebAssembly. Eleven American and British voices are available, speed is adjustable, and the result downloads as a standard WAV file. The text is never sent anywhere and the output has no watermark or usage limit.

**URL:** https://convertto.tech/t/ai-text-to-speech
**Category:** On-Device AI (https://convertto.tech/c/local-ai-tools)
**Privacy:** Runs entirely in the browser; no upload
**Cost:** Free, no sign-up
**Last updated:** 2026-08-01

## Key facts

- **Voices:** 11 American and British, male and female
- **Output:** WAV, 24 kHz — no watermark, no limit
- **Uploads:** None — the text never leaves your device
- **Privacy:** Runs entirely in your browser — nothing is uploaded
- **Cost:** Free, unlimited, no sign-up

## How to use

1. Enter or paste your text.
2. Choose the voice.
3. Set the speed.
4. Press Run, then download the result when it is ready.

## FAQ

### How does this compare with the browser's built-in speech synthesis?

The Web Speech API is instant and free but uses whatever voices the operating system ships, which vary wildly and generally sound robotic — and critically, it plays audio without giving you a file. This produces a downloadable WAV from a neural model, which is what you need for a video, a podcast or a voiceover.

### Can I use the audio commercially?

The Kokoro model is released under Apache 2.0, which permits commercial use. Check the model card yourself before shipping something that depends on it — licences change and you should not take a tool page's word for it.

### Why is there a length limit?

Synthesis time scales with text length and it all happens on your device. Eight thousand characters is a few minutes of speech and a reasonable single run; beyond that a browser tab is the wrong place for the job.

## Sources

- [Kokoro-82M model card](https://huggingface.co/hexgrad/Kokoro-82M) — hexgrad

## Related tools

- [AI Audio Transcription](https://convertto.tech/t/ai-audio-transcription): Run Whisper locally to turn speech into text, SRT, VTT or a timestamped transcript. Works offline after the first run.
- [Audio Converter](https://convertto.tech/t/audio-converter): Convert audio between MP3, WAV, AAC, OGG and FLAC in your browser.
- [Sample Audio File Generator](https://convertto.tech/t/sample-audio-generator): Generate WAV test audio: sine, square, sweep, noise or silence at any rate, depth and duration.
- [Sample File Generator](https://convertto.tech/t/sample-file-generator): Create valid test files of an exact size in 22 formats — text, data, code, images, audio, PDF and ZIP.
- [Sample File Library](https://convertto.tech/t/sample-file-library): Download a ready-made ladder of 15 sample files: size steps, page counts, aspect ratios, durations or row counts.
- [Audio to Subtitles](https://convertto.tech/t/audio-to-subtitles): Generate timed SRT or WebVTT subtitles from an audio recording.
- [Audio to Text Transcriber](https://convertto.tech/t/audio-to-text): Transcribe speech in an audio file to text with Whisper, running on your own device.
- [AI Background Remover](https://convertto.tech/t/ai-background-remover): Cut the subject out of a photo with a real segmentation model that runs in your browser. No upload, no sign-up, no watermark.
