Benchmarks

How our models measure up,
with the workings shown.

We publish the datasets, the pinned revisions, the scoring rules, the audio, the artifact hashes and the limitations for every run, including the results where another system comes out ahead. If you want to check us, everything you need is on these pages or in the archive.

Published runs

Addis Voice 2 on Amharic

Addis Voice 2 against Azure, Gemini, OpenAI, ElevenLabs and Meta MMS on 100 Amharic prompts from three open datasets, with audio, judge transcripts and confidence intervals.

Listening composite
4.25 / 5
Word error rate
13.87%
Date
12 August 2026
Category
Text to speech
Read the benchmark
How we publish

What every run on this page includes

Fixed inputs

Public test-split prompts pinned to exact dataset revisions and drawn once with a fixed seed, so the prompt set can be rebuilt independently and hash-checked against ours.

Full results

Confidence intervals on every error rate, per-slice breakdowns, latency, and the rows where a competitor leads, highlighted for them rather than left out.

Named conditions

Systems without documented support for the language are labelled as unsupported-language conditions, and the model and voice used for every system is published.

Every run is published as an open dataset with the prompts, the generated audio, the judge transcripts, the scorer and a checksum for every file. The Amharic run is at addisai/amharic-tts-benchmark.