Words on screen while people speak
First words appear after 1.56 s on natural speech. Google Chirp 3 streaming takes 6.67 s on the same clips.
An open realtime speech recognition model for Amharic. Text appears while people speak, in 0.32 second steps, and is final about 30 ms after they stop. Fine-tuned from NVIDIA Nemotron 3.5 ASR Streaming on more than 5,000 hours of quality-checked Amharic speech.
Offline recognisers, including our earlier Whisper-based Shook, need the whole recording before they answer. Addis Scribe reads each new piece of audio once, keeps what it learned in a cache, and updates the text without re-reading earlier audio, so cost and delay stay flat however long someone talks.
First words appear after 1.56 s on natural speech. Google Chirp 3 streaming takes 6.67 s on the same clips.
28.4% WER streaming on 100 held-out WAXAL clips, against 32.5% for Chirp 3 streaming and 29.2% for Chirp 3 given the whole file.
Use 0.32 s chunks for the fastest response, or 1.12 s chunks for 27.3% WER at the cost of about 0.65 s more delay.
Full method, every system and the results where another model leads are on the benchmark page.
Download the repository from Hugging Face, then feed audio frames to the streaming helper. Requires NVIDIA NeMo 3.0 with numba-cuda[cu12], numpy<2.4 and nvidia-nvjitlink-cu12>=12.9.
from stream_infer import AmharicStreamerscribe = AmharicStreamer("addis-scribe-streaming.nemo", chunk="320ms") # or "1120ms"for frame in microphone_frames(): # float32 mono, 16 kHz, any frame sizeprint(scribe.feed(frame)) # transcript so farfinal_text = scribe.finish() # at end of speech; scribe.reset() for the next utterance
python stream_infer.py addis-scribe-streaming.nemo speech_16k.wav 320ms
| Architecture | Cache-aware FastConformer encoder, hybrid RNNT + CTC, 0.6B parameters |
| Base model | NVIDIA Nemotron 3.5 ASR Streaming 0.6B |
| Decoding | Greedy RNNT, language prompt am-ET |
| Chunk sizes | 0.32 s (default) or 1.12 s |
| Input / output | 16 kHz mono audio in, Amharic text without punctuation out |
| Toolkit | NVIDIA NeMo 3.0 |
| Licence | OpenMDW-1.1 |
The model checkpoint, loadable with NVIDIA NeMo.
A streaming helper. Fed 100 ms frames like a live microphone, it reproduces the reference transcripts exactly on all 100 benchmark clips.
Sample rate, language prompt and attention context for each chunk size.
Every per-clip transcript and latency, the scored summary and the scoring script.
Every file is listed with its SHA-256 in SHA256SUMS in the repository.
Addis Scribe Streaming is free to download under OpenMDW-1.1. Browse the rest of our open Amharic and Afan Oromo models and datasets.