TTSbox
Text to Speech Built-in voices 8 languages · 54 voices WAV download No signup

Free Kokoro Text to Speech Online

Enter text, choose one of 54 official Kokoro voices across 8 languages, and generate speech locally in your browser. Download the result as a WAV file with no signup.

1
2
3

Kokoro-82M

Q8 model · local browser inference

Official Kokoro samples · previews play instantly

1.00×
0.7×1.3×

Quick facts

Text to Speech at a Glance

Price
Free
Signup
Not required
Input
Up to 1500 characters per generation
Voice sources
54 Kokoro built-in voices
Supported languages
English, Japanese, Chinese, Spanish, French, Hindi, Italian, Portuguese
Processing
Runs locally in supported browsers
Output
WAV audio download
First use
Model download required
Recommended browser
Desktop Chrome or Edge

Workflow

How Text Becomes Voice

Steps

From text to downloadable audio

Four steps to generate voice audio from written text in your browser.

View steps
  1. Enter text: Type or paste your script (up to 1500 characters).
  2. Choose a language and voice: Select a language and pick a built-in sample voice.
  3. Generate speech: Kokoro creates voice audio locally in your browser.
  4. Download WAV audio: Play the generated audio in the browser and download it as a WAV file.

The model is cached in your browser after the first download. Future visits load faster.

Privacy

Text to Speech Privacy

Your text and generated audio stay in your browser.

View details
  • No server upload of your text or generated audio.
  • All generation runs locally via WebAssembly on your device.
  • No account signup or API keys required.
  • Generated WAV files are saved directly to your device.

The model is downloaded from the internet on first use. After caching, it loads from your browser storage.

Output

WAV audio download

Generate and download WAV audio files with no watermarks.

View details
  • Output format: WAV (uncompressed audio).
  • Play audio directly in the browser before downloading.
  • Download the file to use in video editing, presentations, or other projects.
  • No watermarks or branding on the generated audio.

Generated audio quality depends on the voice sample and model capabilities.

Voice Source

Text to Speech and Voice Cloning Are Separate

Tool Best For Requires Upload or Recording
Kokoro Text to Speech Quick text-to-speech generation No
Voice Cloning Matching a custom voice Yes

Need a custom voice? Open the separate Voice Cloning tool and only use a voice you own or have explicit permission to clone.

Use Cases

Best Uses for Text to Speech

Great for

  • YouTube and TikTok draft voiceovers
  • Product demo narration and walkthroughs
  • Audiobook drafts and preview chapters
  • Language localization for indie projects
  • Private voice experiments and prototyping

Not for

  • Impersonation or deceptive content
  • Production dubbing requiring professional QA
  • Commercial use without proper voice licensing

Commercial use depends on your rights to the text, voice source, and generated audio.

FAQ

Frequently asked questions

What is text to speech?
Text to speech (TTS) is technology that converts written text into spoken audio. TTSBox runs Kokoro-82M in your browser and lets you choose from its 54 built-in voices.
Is voice cloning part of this text-to-speech tool?
No. Text to speech and voice cloning are separate tools. This page uses Kokoro built-in voices and never asks for a voice sample. Use the dedicated Voice Cloning page when you need to upload or record an authorized reference voice.
Can I use my own voice?
Use the separate Voice Cloning tool to record your own voice or upload an authorized reference clip. The Kokoro text-to-speech tool only uses built-in voices.
What languages are supported?
Kokoro text to speech supports all 8 official language groups: English, Japanese, Mandarin Chinese, Spanish, French, Hindi, Italian, and Brazilian Portuguese, with 54 voices in total.
Can I download the generated voice?
Yes. After generating speech, you can play the audio directly in the browser and download it as a WAV file. The audio is generated and processed locally on your device. No watermarks are added to the downloaded file.
How long can my text be?
You can enter up to 1500 characters of text per generation. The resulting audio length depends on the text length and speaking pace. For longer content, generate multiple segments and combine them in your preferred audio editor.
Is generated audio processed locally?
Yes. After the initial model download, audio generation runs locally in your browser with WebAssembly. Your text is not sent to a TTSBox server, and the generated WAV file is created on your device.

Related

Related Voice Tools

Need the opposite workflow? Try Speech to Text.