Is this text to speech online tool free?
Yes. You can generate speech on the homepage without an account, API key, subscription, or per-character charge. The voice model runs on your device, so OfflineTTS does not need to charge for server inference.
How do I convert text to speech online?
Paste up to 1,500 characters into the homepage tool, choose an English AI voice and speed, then select Generate Speech. When generation finishes, use the waveform player to listen or download the result.
Can I download text to speech audio as MP3 or WAV?
Yes. Every successful homepage generation can be downloaded as WAV or encoded to MP3 in your browser. The advanced TTS workspace provides the same audio formats for longer scripts.
Does text to speech work offline?
The first English generation downloads and caches the Kokoro model and selected voice. After those files are available, English text to speech can work offline in a compatible browser. Non-English voices may need an online phoneme-conversion step.
Which AI voices can I use?
The homepage offers five curated American and British English voices for quick generation. Open the advanced workspace to browse all 54 Kokoro voices across nine languages or compare Kitten, Piper, Supertonic, and Pocket TTS.
Does OfflineTTS upload my text or generated audio?
No. The homepage generates and plays English speech in your browser. Your text and generated waveform are not uploaded to an OfflineTTS account or generation server.
Can I use generated speech commercially?
Many creator workflows permit commercial use, but licensing depends on the selected upstream model and voice. Review the relevant model terms before publishing or distributing audio commercially.
Can I edit an audio file without uploading it?
Yes. The audio tools decode your file with the browser's own codecs, apply the effect with an offline render, and encode the result as WAV or MP3 in the same tab. The file is never uploaded to an OfflineTTS processing server.
Which audio effects are available?
A ten-band equalizer, bass boost, volume boost and peak normalization, convolution reverb, spectral noise removal, a sixteen-preset voice changer, pitch shifting, and tempo changes with or without pitch lock. Trim, join, reverse, and sample-rate conversion cover the editing side, and BPM, key, Camelot code, and loudness analysis cover measurement.
What audio formats can these tools read and write?
Inputs are MP3, WAV, OGG, Opus, FLAC, M4A, AAC, and the audio track of MP4 or WebM video. Outputs are WAV at 16-bit PCM and MP3 at 128, 192, or 320 kbps. FLAC, OGG, and M4A export are not offered because encoding them in the browser would need a large WASM codec.
Is the tempo and key detection accurate?
It is a good starting point rather than a studio reference. Tempo comes from an onset-interval histogram folded into a 70 to 190 BPM range, which is reliable on music with a clear beat and can lock onto a subdivision on syncopated or free-tempo material. The key estimate correlates a chromagram against major and minor profiles, and the tool reports the harmonic relationships it found rather than claiming certainty.
Do these audio tools work offline?
They work without a network round trip because all processing is local JavaScript and Web Audio. The page itself still needs to load once, and the tools only need the browser's built-in audio codecs, so no model download is required.