Browser TTS workspace

Kokoro TTS

Free AI text to speech with Kokoro TTS. Browser local TTS audio synthesis with 54 voices across 9 languages and WebGPU or WASM inference.

Private generation WAV + MP3 export Default TTS workspace
305-326
q4 / fp32
MB model
54
9 languages
voices
82M
StyleTTS 2
params
WebGPU
+WASM fallback
GPU+CPU

TTS works best on desktop

You can still try lightweight engines on mobile, but desktop Chrome or Edge remains the most reliable setup for large model downloads and long-form generation.

About Kokoro TTS

Kokoro TTS is the default engine on OfflineTTS, offering 54 voices across 9 language variants including English, Japanese, Chinese, Spanish, French, Hindi, Italian, and Portuguese. It makes OfflineTTS a practical Local TTS workspace for browser audio synthesis without requiring an account or hosted synthesis API.

English uses the local kokoro-js phonemizer after required assets download. Non-English Kokoro sends entered text to the documented phonemization service, receives pronunciation tokens, and then generates the waveform on your device.

It supports two model sizes: q4 (~305MB, recommended) and fp32 (~326MB, full precision). q8 quantization is not available as it produces garbled audio with this model.

Compare engines: Kitten TTS (8 voices, 24MB, lightest) · Piper TTS (25 voices, fastest CPU) · Supertonic TTS (31 languages, local inference)

Kokoro Data Path and Practical Limits

Language determines the network path

Model and voice files must be downloaded before synthesis. English text is phonemized in the browser after those assets are available. Japanese, Chinese, Spanish, French, Hindi, Italian, and Brazilian Portuguese text is sent to api.offlinetts.com for phonemization; audio inference and export remain local. Hosting, analytics, and model delivery are documented separately in the Privacy Policy.

Review before a long or public export

Generate a sample containing the final script's names, numbers, abbreviations, quotations, and longest sentence. Compare q4 and fp32 only if the audible result justifies the larger precision choice. Listen for pronunciation, clipped chunk boundaries, repeated words, and pacing, then keep the voice ID, precision, speed, browser, and reviewed source script with an important production file.

Getting Started with Kokoro TTS

New to AI text to speech? Use this sequence to create and review a representative Kokoro sample before generating a longer file.

1. Choose Your Model Size

Start with q4 (~305MB) when download size matters. FP32 (~326MB) retains full numerical precision. Compare both on the same text before assuming an audible benefit.

2. Pick a Voice

Filter by language and compare the same passage in several voices. Catalog grades and traits are discovery labels, not a guarantee for every script.

3. Write Your Script

Use proper punctuation — commas add pauses, periods create full stops, question marks raise pitch. Well-punctuated text produces the most natural speech.

4. Generate & Download

Click generate, wait for the audio to play, then download as WAV (lossless) or MP3 (compressed). WAV is recommended for further editing.

Tips for Best TTS Quality

1.

Compare the available backend. WebGPU can be faster on compatible hardware, but results vary by browser, model, GPU, driver, and memory. Start with automatic selection and record the environment when performance matters.

2.

Punctuate properly. This is the single most important factor for natural-sounding speech. Commas, periods, question marks, and exclamation marks all create distinct prosodic effects.

3.

Break long text into paragraphs. The tool handles up to 50,000 characters, but shorter paragraphs with clear punctuation produce better rhythm and pacing.

4.

Try multiple voices. Different voices suit different content types. Heart excels at warm narration, Bella at energetic delivery, Michael at professional reviews.

5.

Use WAV for production. WAV preserves full audio quality for editing. MP3 is fine for quick sharing, but use WAV if you plan to mix, master, or further process the audio.