TTS Model Quality Ranking 2026: Speech Arena Results
- tts
- ranking
- leaderboard
- comparison
- models
- quality
TTS quality has improved rapidly, but a ranking depends on the exact date, voice, provider, language, prompt set, and listener pool. Open-weight and proprietary systems can both perform well; a public preference score does not show that most users cannot distinguish them in every workflow.
This article preserves an editorial 2026 snapshot assembled from public preference results and product research. The original numeric table mixed sources and should be treated as a dated comparison to verify—not as an official Artificial Analysis order.
2026 TTS Quality Ranking
| Rank | Model | Type | Elo/Score | License | Parameters | Hardware |
|---|---|---|---|---|---|---|
| 1 | ElevenLabs Turbo v2.5 | Proprietary | 1350+ | Commercial | Unknown | API only |
| 2 | Zonos2 8B | Open-weight | 1320+ | Apache 2.0 | 8B MoE | GPU 16GB+ |
| 3 | CosyVoice 3 | Open-weight | 1280+ | Apache 2.0 | 0.5B | GPU 8GB+ |
| 4 | Fish Speech 1.6 | Open-weight | 1260+ | CC-BY-NC-SA | ~500M | GPU 6GB+ |
| 5 | Chatterbox Turbo | Open-weight | 1240+ | MIT | ~1B | GPU 6GB+ |
| 6 | Step Audio EditX | Open-weight | 1230+ | Apache 2.0 | ~1B | GPU 8GB+ |
| 7 | Google Cloud Neural2 | Proprietary | 1220+ | Commercial | Unknown | API only |
| 8 | Azure Neural HD | Proprietary | 1210+ | Commercial | Unknown | API only |
| 9 | F5-TTS | Open-weight | 1180+ | CC-BY-NC | 330M | GPU 6GB+ |
| 10 | Kokoro 82M | Open-weight | 1150+ | Apache 2.0 | 82M | Any CPU |
| 11 | GPT-SoVITS | Open-weight | 1130+ | MIT | ~1B | GPU 8GB+ |
| 12 | OuteTTS 1.0-1B | Open-weight | 1100+ | Apache 2.0 | 1B | CPU/GPU |
| 13 | MeloTTS | Open-weight | 1050+ | MIT | Small | Any CPU |
| 14 | Piper | Open-weight | 950+ | MIT/GPL | Varies | Any CPU |
Elo scores are approximate, based on the Artificial Analysis Speech Arena (Q3 2026). Updated rankings available at artificialanalysis.ai/speech-arena.
Top Tier: Studio Quality (Elo 1300+)
1. ElevenLabs Turbo v2.5
The gold standard for proprietary TTS. Turbov2.5 produces remarkably natural speech with excellent prosody, emotion, and pacing. Voice cloning is best-in-class.
- Best for: Premium content, voice cloning, audiobooks
- Cost: $5–$330/mo
- Hardware: API only (cloud)
- Limitations: Character caps, requires internet, per-character pricing
2. Zonos2 8B
Zyphra’s Zonos2 is the strongest open-weight TTS model. The 8B Mixture-of-Experts architecture delivers quality competitive with ElevenLabs. Apache 2.0 licensed.
- Best for: Self-hosted premium TTS, production deployment
- Cost: Free (self-hosted), GPU cloud ~$0.50–$1.00/hr
- Hardware: 16GB+ VRAM (FP16), or GGUF quantized for CPU
- Limitations: Large GPU requirement, relatively new ecosystem
3. CosyVoice 3
Alibaba’s CosyVoice 3 packs remarkable quality into just 0.5B parameters. Supports 9 languages + 18 Chinese dialects. Zero-shot voice cloning. Apache 2.0 licensed.
- Best for: Multilingual content, Chinese-focused applications, voice cloning
- Cost: Free (self-hosted)
- Hardware: GPU 8GB+ VRAM
- Limitations: Installation complexity
Mid Tier: Excellent Quality (Elo 1150–1300)
4. Fish Speech 1.6
Fish Audio’s model trained on 1M+ hours of speech. Excellent multilingual support with emotion tags. Voice cloning from 10 seconds.
- Best for: Expressive multilingual TTS, podcast/dubbing
- Cost: Free (self-hosted, CC-BY-NC-SA)
- Hardware: GPU 6GB+ VRAM
- Limitations: Non-commercial license, GPU required
5. Chatterbox Turbo
Resemble AI’s Chatterbox Turbo offers MIT-licensed zero-shot voice cloning. Fast inference, good quality, commercially friendly license.
- Best for: Commercial voice cloning projects
- Cost: Free (self-hosted, MIT)
- Hardware: GPU 6GB+ VRAM
- Limitations: Smaller community than established projects
10. Kokoro 82M — Best Quality per Parameter
Kokoro 82M is remarkable for its size. With just 82 million parameters, it produces speech that rivals models 10x its size. Apache 2.0 licensed. The practical choice for most users.
- Best for: General-purpose TTS, CPU inference, batch processing
- Cost: Free (self-hosted, Apache 2.0)
- Hardware: Any modern CPU, 4GB+ RAM
- Voices: 54 across 9 languages
- Limitations: No voice cloning, 9 languages only
Lightweight Tier: Good Quality (Elo below 1150)
13. MeloTTS
Fast multilingual CPU inference. Great for quick prototyping. MIT licensed. Supports English, Chinese, Japanese, Korean, French, Spanish.
14. Piper
The fastest neural TTS on CPU. 900+ English voices. Ideal for Home Assistant and embedded systems. Forked to GPL-3.0 (OHF-Voice) from the original MIT archive.
Open-Source vs Proprietary: The Gap is Closing
In 2024, the gap between open-source and proprietary TTS was significant — ElevenLabs was clearly ahead of anything you could run locally. In 2026, the gap has nearly closed:
| Aspect | Open-Source (2026) | Proprietary (2026) |
|---|---|---|
| Voice quality | Near parity | Slightly ahead |
| Voice cloning | Good (F5-TTS, CosyVoice) | Excellent (ElevenLabs) |
| Languages | 9–13 (Kokoro, CosyVoice) | 30–140+ |
| Latency | Device-dependent | Network-dependent |
| Cost at scale | $0 | $4–$220 per 1M chars |
| Privacy | Full (local) | None (cloud) |
| License | Apache 2.0 / MIT / CC-BY-NC | Commercial |
For most use cases — YouTube voice-overs, podcasts, e-learning, audiobooks — open-weight models like Kokoro or CosyVoice deliver quality that’s indistinguishable from cloud APIs in blind testing. The main remaining advantage of proprietary TTS is language breadth and maximum quality in long-form content.
Best Model by Use Case
| Use Case | Best Model | Why |
|---|---|---|
| General TTS (CPU) | Kokoro 82M | Best quality-to-size, any CPU, Apache 2.0 |
| General TTS (GPU) | CosyVoice 3 | 0.5B, excellent quality, voice cloning |
| Voice cloning | F5-TTS | 5-second reference, good quality, easy setup |
| Premium self-hosted | Zonos2 8B | Best quality, Apache 2.0, needs 16GB GPU |
| Embedded/Home Assistant | Piper | Fastest CPU inference, 900+ English voices |
| Multilingual (cloud) | Google Neural2 | 50+ languages, full SSML |
| Commercial API (cheap) | Amazon Polly | $4/1M chars, SSML, Speech Marks |
| Professional cloud | ElevenLabs | Best overall quality, voice cloning |
| Browser (zero setup) | OfflineTTS (Kokoro) | Free, private, works offline |
How We Tested
This ranking draws from:
- Artificial Analysis Speech Arena — community blind A/B testing (Elo ratings)
- Editorial auditions — useful for identifying issues, but not a published controlled MOS study
- Workflow checks — representative narration and production examples where the model was accessible
- Community benchmarks — Hugging Face TTS leaderboard, Reddit discussions, GitHub issues
Quality is subjective — your mileage depends on voice selection, text content, and specific use case. We recommend testing 2-3 models for your specific content before committing to one.
Verification Rules for Model Ranking Claims
This page was reviewed on August 1, 2026. The “Elo/Score” values in the editorial table are not all rows from one current leaderboard and therefore must not be cited as official live Elo ratings. For procurement, replace each candidate with its current upstream model card and a dated row from the live benchmark, or label the value as unavailable.
Run a project-specific blind test with identical normalized text and output loudness. Record model and provider versions, voice, language, settings, hardware or API, latency, failed renders, listener recruitment, sample count, confidence or disagreement, and pronunciation corrections. Score quality, speed, rights, cost, and data path separately instead of adding them into one unsupported ordinal rank.
Try the Top Open-Source Model
OfflineTTS provides a browser integration of Kokoro 82M without an account or provider usage bill. The local engine can work after caching on a compatible device; verify language routing and network behavior. Its site grade is an editorial audition aid, not a standardized MOS result or permanent global rank.
Related articles
Try OfflineTTS
Four local TTS engines, Whisper transcription, and private browser audio tools.
Open TTS Tool