TTSbox

Technology and Licensing Transparency

TTSbox is designed as a local browser AI audio tool. We believe in transparency regarding how data is handled and the ethical boundaries guiding this technology.

What TTSbox discloses

We document the parts users need to evaluate trust: how local mode handles voice samples, how browser caching works, what rights are required for uploaded voices, what uses are prohibited, and our commitment to privacy and transparency.

What TTSbox does not disclose

TTSbox does not publicly disclose every implementation detail of its model pipeline. This helps protect product integrity, reduce misuse, and avoid exposing implementation-specific behavior that could be used to bypass safety boundaries. Instead, TTSbox documents the trust boundaries users need to evaluate: local-mode data handling, browser caching, voice rights, commercial use limits, sample voice policy, and prohibited uses.

Local-mode data handling

In local mode, TTSbox does not upload your voice sample or generated audio to a TTSbox server. Local inference requires WebGPU and a compatible desktop browser. Generated audio is exported as WAV directly from your device.

Browser model loading and caching

Model files are loaded by the browser for local inference. To improve performance on subsequent visits, model files may be stored in browser-managed Cache Storage or IndexedDB. Users retain control over this data and can clear it through browser settings at any time.

Text-to-speech model licenses

Text to speech uses the complete Kokoro-82M voice set across its 8 official language groups through the browser ONNX distribution. Kokoro-82M is provided under the Apache License 2.0. The multilingual pronunciation pipeline uses eSpeak NG (GPL-3.0-or-later), Kuroshiro (MIT), and Kuromoji (Apache-2.0). The separate voice-cloning tool uses Kyutai Pocket TTS with its ONNX export; the model and export retain their respective attribution and license requirements.

Model and runtime licenses apply independently from the TTSbox website terms. See the Kokoro-82M model card, eSpeak NG repository, Kuroshiro repository, Kuromoji repository, and Pocket TTS model card for the source terms.

Voice rights and consent

Users must only use voices they own, have explicit permission to use, or properly licensed synthetic/sample voices. Because processing happens offline on the user's browser, TTSbox cannot independently verify every uploaded sample.

Sample voice policy

The platform includes sample voices strictly for responsible testing and experimentation. Users must adhere to safe usage guidelines when interacting with these samples.

Commercial use boundaries

Commercial use requires rights clearance by the user. You are fully responsible for ensuring you have the legal right to generate and distribute the audio.

Prohibited uses

Prohibited uses include impersonation, fraud, scams, phishing, political deception, non-consensual voice cloning, misleading public audio, and identity bypass.

Abuse reporting

Safety and abuse reporting remain documented. If you encounter audio generated using TTSbox that violates these policies, please report it immediately to: abuse@ttsbox.xyz

Not legal advice

This transparency document provides technical and ethical boundaries but does not constitute legal advice. Users must consult legal professionals regarding laws applicable in their jurisdiction.

Last reviewed: July 2026

FAQ

Frequently asked questions

Does TTSbox disclose its underlying model pipeline?
TTSbox does not publicly disclose every implementation detail of its model pipeline. This helps protect product integrity, reduce misuse, and avoid exposing implementation-specific behavior that could be used to bypass safety boundaries.
How does TTSbox handle local-mode data?
In local mode, TTSbox does not upload your voice sample or generated audio to a TTSbox server. Execution happens securely on your own device.
How does browser model caching work?
Model files are downloaded to your browser and stored in browser-managed Cache Storage or IndexedDB to speed up future sessions. You can clear this data through your browser settings at any time.
What is the policy on sample voices?
TTSbox provides sample voices built for responsible testing. Users must adhere to licensing constraints and are encouraged to only use voices they own or have explicit permission to clone.
Can I use generated audio commercially?
Commercial use requires rights clearance by the user. You should only use voices commercially if you own them, have explicit permission to use them, or they are properly licensed synthetic voices.
What uses are strictly prohibited?
Prohibited uses include impersonation, fraud, scams, phishing, political deception, non-consensual voice cloning, misleading public audio, and identity bypass.
Can TTSbox verify my uploaded voice sample?
No. Because TTSbox operates largely as a local browser tool, TTSbox cannot independently verify every uploaded sample. Users bear full responsibility for rights clearance.