ToolsTray

Star a tool to keep it here.

Text to Speech

Text to speech two ways: the voices your device already has, or a 13-voice English model that hands you a file.

Runs entirely in your browser — your files and text never leave your device. The AI model downloads the first time you run this tool, then stays on this device — so later runs work offline.

Downloads the speech synthesis model (~92 MB) the first time you run it, then works offline.

System voices play through your device and aren't downloadable; switch the Engine to Natural to generate audio you can save.

Start here

  1. Put your text in the box at the top. An article you're proofing, or a script you want to hear out loud.
  2. Choose an Engine. System voices speak straight away in any language your device has. Natural fetches its model the first time, then gives you the English voices and the download button, and Heart is the safest voice to start on.
  3. Set the Speed, press Speak, and follow the word readout. On Natural, let the bar finish, then pick a format and press Download audio.

About this tool

Paste something in and it gets read back to you. The Engine dropdown decides how. System voices hand the text to whatever speech engine your operating system already has, which starts playing at once and covers every language you have installed. Switch to Natural and an 82-million-parameter model called Kokoro runs on your own machine instead, with 13 English voices: six American women, three American men, two British women and two British men.

The difference that matters most is whether you end up with a file. System voices play out through the OS and can't be captured, so there's nothing to save. Natural builds the waveform itself at 24 kHz, and when it finishes a format picker and a Download audio button appear, giving you speech.m4a or the same thing as WAV, Opus or FLAC. That's the route to a voiceover for a video, or an audio copy of your notes for the commute.

On both engines the word currently being spoken appears under the text box. That little readout is what makes proofreading by ear practical, and Speed a notch or two above 1.0 suits the job. Pitch only moves system voices. Length isn't a problem on Natural either, because it breaks the input into pieces of about 450 characters at sentence boundaries and starts playing the first piece while the rest is still being made.

Meeting notes for the drive home, as a file

The notes run to 235 words, 1,320 characters, four paragraphs. Set the Engine to Natural, leave Heart selected, nudge Speed to 1.2 and press Speak. On a first visit the status reads Downloading model with a percentage while the 92 MB model and a 2.8 MB pronunciation dictionary of about 126,000 entries come down, then the 510 KB style file for Heart. The text is split at sentence ends into four pieces of 404, then 295, then 312, then 306 characters, each kept under the 450-character target.

Playback starts the moment the first piece is ready, and the current word appears under the box while the remaining three are still being generated. Each piece is levelled to a peak of 0.9 before it plays, a quiet one lifted by up to four times, and a 0.12-second gap is left at every join, so the finished recording is the four pieces plus 0.36 seconds of silence. When the bar completes, the format picker appears; pick M4A and Download audio writes speech.m4a at 24 kHz, ready for the car.

Good to know

  • Natural is English only. The voices are American and British, and handing it another language gets you an English speaker mangling it. Anything else needs the System engine and whatever your OS has installed.
  • System voices can't be saved. They play through the operating system, out of the page's reach, which is why the download row only appears once you're on Natural.
  • The Natural model is a one-time fetch, and it comes in two precisions of the same model: a machine with WebGPU gets the full-precision version at 326 MB because the graphics card runs it quickly, everything else gets an 8-bit version at 92 MB that suits the WebAssembly path. The first half of the progress bar is that fetch and the second half is the pieces of speech arriving one by one. Your browser caches it, but the first run on a slow connection is a real wait.
  • Without WebGPU the first chunk takes noticeably longer to arrive, so there's a gap between pressing Speak and hearing anything on a long passage.
  • Some browsers cut a long System-voice request off partway through. It's a quirk of their own speech engines. Split the text into paragraphs, or move to Natural, which chunks long input by itself.
  • Natural reads the text as English prose, whatever it contains. A pronunciation dictionary covers around 126,000 entries, and anything outside it, such as a product name or an acronym, falls back to letter-to-sound rules that guess from the spelling, with no way to teach it a correction.
  • Speed is fixed for the run. On Natural the value is handed to the model with each piece as it is generated, so moving the slider mid-sentence changes nothing until the next press of Speak. System voices restart the whole passage from the top when Voice, Speed or Pitch is changed while they are talking.

Frequently asked questions

Can I save the narration as an audio file?

On the Natural engine, yes. It generates the samples on your machine, so when the run finishes you can write them out as M4A, WAV, Opus or FLAC. System voices go straight to your speakers with nothing for the page to grab, which is why that engine has no download row.

Which engine should I pick?

System voices for a quick listen, a language other than English, or text you just want checked by ear. Natural when you care what the voice sounds like, want the same voice on every device you use, or need a file at the end of it.

Which of the 13 Natural voices is worth starting on?

Heart, the American female voice selected by default, is the strongest all-rounder. The dropdown groups the rest by accent and gender, and Fable on the British male side and Michael on the American male side are the two most people settle on next.

Can I pause a long Natural reading and pick it up later?

Pause holds the audio where it is, and Resume carries on from the same word, for as long as the tab stays open. Stop throws the run away. Nothing is stored between visits, so a reading you want to come back to tomorrow is better saved with Download audio and played from the file.

Why do the system voices change from one device to another?

They belong to the operating system and the browser, not to this page. A phone, a laptop and two different browsers each ship a different set, with different quality. Getting one consistent voice everywhere is a large part of why the Natural engine is here.

Missing something in Text to Speech? Suggest a feature →

Tell someone who needs this.

LinkedInXWhatsAppEmail