TTS Arena Leaderboard 2026: How to Read the Results
- tts
- comparison
- benchmark
- leaderboard
- open-source
The Artificial Analysis text-to-speech leaderboard can help you discover speech systems worth auditioning. It should be read alongside the current methodology and the exact view you selected. A result summarizes a comparison pool; it does not certify a voice for every language or production requirement.
Correction, September 30, 2026: Earlier versions of this article presented detailed scores, vote counts, category results, and a reconstructed historical trend without a retained source export sufficient to reproduce them. Those tables and related claims have been removed. We do not publish a current numeric order here. Follow the primary source for its current results; the guidance below explains how to record and interpret them responsibly.
Identify the Benchmark before Comparing Numbers
Several speech evaluations and services use similar “arena” language. Confirm the publisher, URL, task, and methodology before combining results. Text-to-speech listening preference differs from speech recognition error rates, latency benchmarks, and a service’s own demo gallery. A number from one task cannot be substituted for a number from another.
Check what listeners heard and how they made their choices. Where a benchmark uses anonymous paired listening, its result reflects preference among the presented samples. It may capture naturalness or expressive delivery more strongly than exact pronunciation. A conversational prompt set may reward qualities different from technical narration. Consult the publisher’s current explanation rather than assuming every evaluation uses the same categories or accents.
A benchmark can change its prompt pool, participating models, filtering, or scoring method. Record those conditions if you intend to compare a snapshot with a later one. A changed order is a reason to inspect the new samples and documentation, not proof that a deployed model’s weights changed.
Read the Full Row Label
Distinguish a model family from a version, provider, and voice. A hosted service can wrap an open-weight model with different normalization, sampling, or infrastructure. A browser package may quantize the same weights or expose different voices. The result for one configuration does not transfer automatically to another configuration with a familiar model name.
If you plan to evaluate Kokoro in a browser, record the implementation, precision, voice, language, and acceleration backend. Kokoro’s model card establishes its architecture and upstream license. It does not certify that a specific OfflineTTS render matches an externally hosted benchmark sample. Use a matching script to determine whether the differences matter for your use case.
The same distinction applies to proprietary services. A provider may offer several speech models with different latency, language support, or expressive controls. Save the exact product and version rather than a company name. When a row name is ambiguous, leave your interpretation unresolved until the source documentation clarifies it.
Treat Preference and Uncertainty Together
A leaderboard’s first row is easy to quote, but the distance between nearby candidates needs context. Inspect confidence information or sample participation where the publisher provides it. A small nominal difference does not necessarily support a strong claim that one candidate will be preferred by your audience. Participation counts are not a guarantee of representative language coverage.
Scores are relative to a particular comparison pool. They are not percentages of pronunciation accuracy or universal audio quality. A win rate is affected by the opponents and cases represented in the evaluation. Do not convert an Elo-style rating into an invented “studio quality” threshold or assume the same numerical scale can be compared across independent leaderboards.
A relative rank also omits hard requirements. A highly preferred sample can still contain an incorrect name, omit a word, or use an accent your audience does not understand. Listen to the relevant examples and test difficult vocabulary rather than relying only on the aggregate number.
Editorial Snapshot Checklist
For a reproducible citation, retain a dated source capture or permitted export, the full source URL, selected filters, model/provider/voice labels, and the visible values you quote. Record uncertainty or participation information where available. Separate the time you retrieved the source from the model’s release date and from your article’s original publication date.
If a source is inaccessible or a historical export is missing, state that the number is unavailable. An approximate score written from memory is not a snapshot. Linking to a live page does not make an old table reproducible, because the live page can change. Avoid silently refreshing only the article date while leaving an older order underneath it.
Use a correction note when an earlier claim cannot be substantiated. Keep the useful explanation and replace the unsupported table with a direct source link. This page follows that approach. We can explain how to investigate a benchmark without inventing a permanent rank for a model we want readers to try.
Open Weights and Browser Deployment Are Different Questions
“Open-weight” describes availability of model parameters under a stated license. It does not mean that a model can run on every device, that all assets share a permissive license, or that its hosted API is private. Large models may need a server environment, while smaller models may have suitable browser integrations. Check the actual runtime and resource requirements.
OfflineTTS has five engine integrations, each with its own tradeoffs. The site can synthesize audio locally, but non-English Kokoro uses a separate text phonemization service. Model downloads, page hosting, and optional analytics or advertisements are additional network paths. Review our privacy policy before using confidential material.
The browser’s cold start matters as well. A listener comparison usually does not represent the delay for downloading assets, the storage quota of a mobile device, or the behavior after a suspended tab. Test these operational boundaries separately. Keep warm render time and first-use setup time in different fields.
What the Leaderboard Does Not Purchase for You
An audio preference score does not grant a commercial license, verify speaker permission, or clear the source script’s copyright. Review the model card, voice-specific terms, and any reference recording rights. A reference-based cloning feature needs the speaker’s authorization even when its software can be downloaded without payment.
Similarly, cost belongs to the deployment configuration. A hosted row’s price is not the operating cost of running that model yourself. Provider prices can change and may use incompatible billing units. Local inference removes a provider generation bill in supported workflows, while hardware, electricity, maintenance, and review still take resources.
A benchmark also does not settle privacy or compliance. A hosted provider has policies and contracts; a local application has network routes and browser storage. Evaluate those against your actual obligations. The answer may differ for a public tutorial and for a confidential client document even when the same voice sounds appropriate.
Turn a Result into an Audition
Shortlist a few candidates that meet language, rights, deployment, and data-handling requirements. Prepare a fixed script with the names, numbers, abbreviations, pauses, and sentence lengths that occur in your work. Save each output with its configuration and use neutral labels during listening. Alternate playback order to reduce the effect of hearing one sample first.
Record intelligibility, corrections, fatigue, pacing, and suitability for the audience. For a short video, inspect the result inside the actual video edit. For a long document, review a continuous passage rather than stitching together flattering isolated sentences. Measure export reliability and the time needed to fix mistakes alongside listener impressions.
Repeat an audition when the deployed version or voice changes. Keep the old result so a regression can be investigated. Your production record can be much more useful than a leaderboard position because it describes the output you actually shipped.
Our model comparison guide provides a fuller recording checklist. You can also open the workspace to try the integrated engines. We offer these as practical evaluation steps, without claiming that a catalog label establishes a global winner.
Sources
- 1. Artificial Analysis Text-to-Speech Leaderboard — Artificial Analysis
- 2. Kokoro-82M model card — hexgrad
Related articles
Try OfflineTTS
Five local TTS engines, Whisper transcription, and private browser audio tools.
Open TTS Tool