Browser TTS workspace

Kitten TTS

Free AI text to speech with Kitten TTS. 8 expression-based voices, ~24MB model, WebGPU + WASM.

Private generation WAV + MP3 export 8 voices · Lightest
24
lightest
MB model
8
expressions
voices
15M
compact
params
WebGPU
+WASM
GPU+CPU

TTS works best on desktop

Audio generation uses WebGPU/WASM. Desktop Chrome or Edge gives the most reliable result.

Sponsored

Ads help keep OfflineTTS free to use.

About Kitten TTS

Kitten TTS is the smallest TTS model option currently integrated into OfflineTTS at approximately 24MB before related assets and browser overhead. Actual load time and device compatibility depend on the connection, browser, available memory, and backend.

The 8 expression-labeled presets cover cheerful, serious, sad, whisper, excited, gentle, calm, and neutral. Each preset selects a compact embedding that changes the model output; the label is not an analysis of the text or a guaranteed emotion.

The interface offers output sample-rate choices from 16kHz to 48kHz. Values above the model's native output are resampled and do not create additional source detail.

Compare engines: Kokoro TTS (54 voices · q4 or fp32) · Piper TTS (25 voices · Fastest CPU) · Supertonic TTS (31 languages · Local)

The 24 MB Tradeoff

Kitten is the compact option in the OfflineTTS engine switcher. The approximate model size can reduce first-load bandwidth and cache use compared with Kokoro, but it does not predict generation speed or compatibility by itself. Browser memory, the selected backend, device thermal limits, extensions, and other open tabs can still affect a run. Test the intended browser and keep a short fallback script before relying on it in a live workflow.

The model produces English speech from eight preset embeddings rather than a broad catalog of named speakers. Choose it when a lightweight draft, a simple expression comparison, or a smaller cached asset is the priority. Use Kokoro or Piper when their speaker selection better fits the text, and use Supertonic when a supported non-English language is required.

Expression Presets Are Not Emotion Detection

Cheer, Serious, Sad, Whisper, Excited, Gentle, Calm, and Neutral are interface labels attached to fixed voice embeddings. The application does not read the semantic meaning of a paragraph and choose an emotion, detect a real speaker, or guarantee that listeners will perceive the label. Punctuation, wording, speed, and the model can make the same preset sound different across passages.

Compare all candidate presets with identical text and settings. Listen beyond the opening sentence for intelligibility, level, artifacts, and whether the delivery stays usable. Record the preset ID and output rate with production files. The page generates locally after assets load and does not send the text to an OfflineTTS synthesis API; normal model delivery and website requests remain documented separately.

Getting Started with Kitten TTS

1

Load the Compact Model

The approximately 24MB model is smaller than the other integrated TTS options. Load time still depends on the connection and related assets; the browser may reuse cached files on later visits.

2

Choose an Expression

Select from 8 expressions: cheerful, serious, sad, whisper, excited, gentle, calm, or neutral. Each expression shapes the emotional character of the output.

3

Set Your Sample Rate

Choose 16kHz, 22.05kHz, 24kHz, 44.1kHz, or 48kHz output. Higher resampled rates produce larger files but do not add detail that was absent from the generated source.

4

Generate & Iterate

Click generate to hear your text spoken. Try different expressions for the same text to find the right tone. Download as WAV or MP3 when you're satisfied.

Tips for Kitten TTS

1

Use expressions creatively. Whisper expression is perfect for ASMR-style content and intimate narration. Serious expression suits business and formal content. Cheerful works for welcome messages and children's content.

2

Treat sample rate as an output setting. Use a rate required by the next tool in your workflow. Resampling to 48kHz changes container compatibility and file size, not the underlying model information.

3

Use it for quick comparisons. The compact download can make Kitten convenient for drafts. Compare the final expression on representative text instead of assuming a smaller or larger engine is automatically better.

4

Test the target device. A smaller model can reduce download and memory pressure, but mobile and low-memory browser behavior still varies. Confirm generation and export on the actual device.