data/prompts.jsonl
100 rows. Source dataset, pinned repository revision, upstream row id, split, licence, reference text, normalised text, length stratum, character count, human audio hash and duration.
data/judge_rows.jsonl
800 rows. System, raw and normalised transcript, character and word edit counts, reference lengths, CER, WER, judge confidence, plus the audio path and SHA-256 of the exact file that produced the row.
audio/
800 clips. One per prompt per system, plus the original human recordings, with a metadata index listing each clip against its transcript, CER and WER.
score.py and checksums.sha256
The scorer, which rebuilds the CER/WER tables and figures from the published judge transcripts with no arguments, and a SHA-256 for every shipped file.