Open source

Addis Scribe Streaming.
Amharic, transcribed as it is spoken.

An open realtime speech recognition model for Amharic. Text appears while people speak, in 0.32 second steps, and is final about 30 ms after they stop. Fine-tuned from NVIDIA Nemotron 3.5 ASR Streaming on more than 5,000 hours of quality-checked Amharic speech.

Parameters
0.6B
Word error rate, natural speech
28.4%
Final text after speech ends
30 ms
Training audio
5,000+ hrs
Why streaming

Built for live audio, not finished recordings

Offline recognisers, including our earlier Whisper-based Shook, need the whole recording before they answer. Addis Scribe reads each new piece of audio once, keeps what it learned in a cache, and updates the text without re-reading earlier audio, so cost and delay stay flat however long someone talks.

Live

Words on screen while people speak

First words appear after 1.56 s on natural speech. Google Chirp 3 streaming takes 6.67 s on the same clips.

Accurate

Lower error than Google Chirp 3

28.4% WER streaming on 100 held-out WAXAL clips, against 32.5% for Chirp 3 streaming and 29.2% for Chirp 3 given the whole file.

Tunable

Two chunk sizes

Use 0.32 s chunks for the fastest response, or 1.12 s chunks for 27.3% WER at the cost of about 0.65 s more delay.

Full method, every system and the results where another model leads are on the benchmark page.

Quickstart

Transcribe a live stream in a few lines

Download the repository from Hugging Face, then feed audio frames to the streaming helper. Requires NVIDIA NeMo 3.0 with numba-cuda[cu12], numpy<2.4 and nvidia-nvjitlink-cu12>=12.9.

python
from stream_infer import AmharicStreamer
scribe = AmharicStreamer("addis-scribe-streaming.nemo", chunk="320ms") # or "1120ms"
for frame in microphone_frames(): # float32 mono, 16 kHz, any frame size
print(scribe.feed(frame)) # transcript so far
final_text = scribe.finish() # at end of speech; scribe.reset() for the next utterance

From the command line

shell
python stream_infer.py addis-scribe-streaming.nemo speech_16k.wav 320ms
Model

What is in the release

ArchitectureCache-aware FastConformer encoder, hybrid RNNT + CTC, 0.6B parameters
Base modelNVIDIA Nemotron 3.5 ASR Streaming 0.6B
DecodingGreedy RNNT, language prompt am-ET
Chunk sizes0.32 s (default) or 1.12 s
Input / output16 kHz mono audio in, Amharic text without punctuation out
ToolkitNVIDIA NeMo 3.0
LicenceOpenMDW-1.1
addis-scribe-streaming.nemo

The model checkpoint, loadable with NVIDIA NeMo.

stream_infer.py

A streaming helper. Fed 100 ms frames like a live microphone, it reproduces the reference transcripts exactly on all 100 benchmark clips.

serving.json

Sample rate, language prompt and attention context for each chunk size.

benchmark/

Every per-clip transcript and latency, the scored summary and the scoring script.

Every file is listed with its SHA-256 in SHA256SUMS in the repository.

Training and limits

How it was trained, and where it falls short

  • Fine-tuned on more than 5,000 hours of quality-checked Amharic speech, then refined on human-transcribed speech. FLEURS, the WAXAL test split and the evaluation speakers were never used in training.
  • On natural conversational speech it is about 3 WER points behind Hohe, an open model that runs offline only.
  • Outputs no punctuation. Its training text had punctuation removed.
  • Tested on Amharic only. Other languages are out of scope.

Build on it.

Addis Scribe Streaming is free to download under OpenMDW-1.1. Browse the rest of our open Amharic and Afan Oromo models and datasets.