Benchmarks

How Addis Scribe measures up
on live Amharic speech.

A public, reproducible evaluation of realtime Amharic speech recognition on held-out natural and read speech, scoring both accuracy and how fast text reaches the screen.

Word error rate, natural speech
28.4%
Final text after speech ends
30 ms
First words on screen
1.56 s
Released
5 October 2026
Summary

Three results, including the one we lose

Addis Scribe Streaming is the most accurate live system in this run and the fastest by a wide margin. An offline-only model, Hohe, is more accurate on natural speech, and we publish that alongside the rest.

Result 1

Most accurate live system on natural speech

On 100 held-out WAXAL clips, streaming in 0.32 s steps, Addis Scribe records 28.4% WER and 13.1% CER. Google Chirp 3 records 32.5% WER streaming and 29.2% WER when given the whole file at once.

Read the full accuracy tables
Result 2

Final text 30 ms after the speaker stops

First words appear after 1.56 s and the transcript is final 30 ms after speech ends. Google Chirp 3 streaming takes 6.67 s and 1,411 ms on the same clips.

Read how latency was measured
Result 3

Hohe leads on natural speech, offline

Hohe, an open wav2vec2-BERT model, records 25.4% WER on the same WAXAL clips, about 3 points ahead of Addis Scribe. It needs the full recording before it can answer, so it cannot run on live audio.

Read the limitations
Systems

What was tested, and how each handles live audio

Streaming here means the model reads each new piece of audio once, keeps what it learned in a cache, and updates the text without re-reading earlier audio.

SystemArchitectureLive audioHow it handles live audio
Addis Scribe StreamingCache-aware FastConformer, hybrid RNNT + CTC, 0.6BNative streamingEach 0.32 s (or 1.12 s) chunk is read once and cached, so cost and delay stay flat however long someone talks.
Google Chirp 3 (am-ET)Cloud APINative (cloud API)Server-side streaming over the network. Also tested as a batch request on the whole file.
Shook (Whisper medium)Whisper medium, Addis AI's previous open ASR modelNot supported, run liveWhisper transcribes complete audio. Run live with the WhisperLiveKit method: once per second it re-transcribes all audio so far, so delay grows with length.
Hohe (wav2vec2-BERT)wav2vec2-BERT, third-party open modelOffline onlyNeeds the whole recording before it can answer. Scored for accuracy only.

Shook is Addis AI's previous open speech recognition model. Hohe is a third-party open model, scored for accuracy only because it cannot run on live audio.

Protocol

Test sets and scoring rules

Both test sets are public and neither was used in training. WAXAL test, FLEURS and the evaluation speakers were held out of every training stage.

Test setRepositorySplitWhat it measuresSelection
WAXALgoogle/WaxalNLPamh testNatural speech with human transcripts100 clips, 22 speakers, 29 minutes, 2 to 30 s each, drawn at random with a fixed seed
FLEURSgoogle/fleursam_et testRead sentencesFirst 150 clips for streaming, all 516 offline

Scoring

Every system is scored with the same text normaliser: Unicode NFC, punctuation and symbols removed, whitespace collapsed. WER is word error rate and CER is character error rate, both lower is better. Addis Scribe outputs text without punctuation, so removing it from every system keeps the comparison even.

Accuracy

Word and character error rate

Natural conversational speech is harder than read sentences, so WAXAL error rates run higher than FLEURS for every system.

WAXAL: natural speech

100 held-out clips, word error rate, lower is better

0%10%20%30%40%Hohe (offline only)Hohe (offline only): 25.4%25.4%Addis Scribe (streaming)Addis Scribe (streaming): 28.4%28.4%Google Chirp 3 (batch)Google Chirp 3 (batch): 29.2%29.2%Google Chirp 3 (streaming)Google Chirp 3 (streaming): 32.5%32.5%Shook (live)Shook (live): 35.0%35.0%Word error rate
Addis ScribeOther live systemsOffline only

FLEURS: read speech

First 150 test clips, word error rate, lower is better

0%10%20%30%40%Addis Scribe (streaming)Addis Scribe (streaming): 19.8%19.8%Hohe (offline only)Hohe (offline only): 20.6%20.6%Previous Addis AI modelPrevious Addis AI model: 35.0%35.0%Word error rate
Addis ScribeOther live systemsOffline only

WAXAL natural speech, every system and mode

SystemModeWERCERFinal text after speech endsFirst words
Addis Scribe StreamingStreaming, 0.32 s chunks28.4%13.1%30 ms1.56 s
Addis Scribe StreamingStreaming, 1.12 s chunks27.3%12.4%34 ms2.21 s
Google Chirp 3 (am-ET)Streaming, realtime32.5%16.5%1,411 ms6.67 s
Google Chirp 3 (am-ET)Batch, whole file29.2%13.5%n/an/a
Shook (Whisper medium)Live, WhisperLiveKit method (1 s updates)35.0%16.6%10,839 ms3.66 s
Shook (Whisper medium)Offline35.0%16.6%n/an/a
Hohe (wav2vec2-BERT)Offline only25.4%11.2%n/an/a

Bold marks the best value in each column across every mode, including offline ones. The 1.12 s chunk setting trades about 0.65 s of extra delay for 1.1 points lower WER.

FLEURS read speech

SystemModeClipsWERCER
Addis Scribe StreamingStreaming, 0.32 s chunksFirst 15019.8%8.2%
Hohe (wav2vec2-BERT)OfflineFirst 15020.6%7.9%
Previous Addis AI streaming modelStreaming, 0.32 s chunksFirst 15035.0%17.0%
Addis Scribe StreamingOfflineAll 51620.6%7.7%
Hohe (wav2vec2-BERT)OfflineAll 51620.2%7.3%

On the first 150 clips Addis Scribe streaming has the lowest WER and Hohe the lowest CER. On all 516 clips offline, Hohe leads on both, by 0.4 points of WER and CER.

Latency

How fast the words reach the screen

Audio is fed at real-time speed, as from a live microphone. We report two numbers: how long until the first words appear, and how long after the speaker stops until the transcript is final.

First words on screen

WAXAL, 100 clips, lower is better

0 s2 s4 s6 s8 sAddis ScribeAddis Scribe: 1.56 s1.56 sShook (live)Shook (live): 3.66 s3.66 sGoogle Chirp 3Google Chirp 3: 6.67 s6.67 sSeconds
Addis ScribeOther live systems

Final text after speech ends

WAXAL, 100 clips, lower is better

0 s0.5 s1 s1.5 s2 sAddis ScribeAddis Scribe: 30 ms30 msGoogle Chirp 3Google Chirp 3: 1,411 ms1,411 msShook (live)Shook (live): 10,839 ms10,839 msSeconds
Addis ScribeOther live systems

Addis Scribe and Shook were measured on one NVIDIA A100 with no network hop. Google's figures include the network round trip to its EU endpoint from the benchmark machine, so part of its delay is transport rather than the model. Shook's bar is cut at 2 s; its delay grows with utterance length because every update re-transcribes all audio so far.

Limitations

What this run does not show

  • On natural conversational speech Addis Scribe is about 3 WER points behind Hohe, which is offline only.
  • Addis Scribe outputs no punctuation. Its training text had punctuation removed.
  • Tested on Amharic only. Other languages are out of scope for this model.
  • WAXAL results come from 100 clips and 22 speakers. That is enough to rank systems with wide gaps, not to separate systems a point or two apart.
  • Latency is measured on a dedicated GPU with one stream at a time. It is not a capacity or concurrency guarantee.
Reproduce

Run this yourself

Every per-clip transcript, every latency measurement and the scoring script are published next to the model. Running python score.py rebuilds the summary exactly.

benchmark/clips.json

The 100 WAXAL Amharic test clips: ids, speakers, durations and human reference transcripts.

benchmark/transcripts.json

Every system's transcript and latency for every clip, in every mode tested.

benchmark/results.json

The scored summary behind every number on this page.

benchmark/score.py

The scorer, including the shared text normaliser applied to every system.

Published at addisai/addis-scribe-streaming/benchmark. Questions to contact@addisassistant.com.

The model is open

Addis Scribe Streaming is published on Hugging Face with the checkpoint, a streaming helper and the benchmark records.