OmniASR Yiddish · 1 hour
198 held-out test utterances joined into one upload-ready MP3 · 6,899 reference words
OPEN BENCHMARK · 2026
A direct comparison of five commercial and open-source Yiddish transcription systems across shared test sets.
CURRENT RESULTS
Raw and normalized WER are entered manually after each service returns its transcript.
| # | System | Type | Overall norm. | Omni 1h raw | Omni 1h norm. | Multi-genre raw | Multi-genre norm. | Status |
|---|
SYSTEMS IN SCOPE
Three hosted commercial services and two reproducible open models. Exact versions are recorded with every result.
THE TEST SETS
A one-hour single-file stress test plus a multi-genre collection reveal different ASR strengths.
Download the 1-hour MP3 ↗198 held-out test utterances joined into one upload-ready MP3 · 6,899 reference words
Six Yiddish24 clips covering shiur, news, interview, podcasts, and vlog speech
Long-form religious discourse
Three joined bulletins
Multi-speaker conversation
A single-speaker podcast episode
Informal program audio
Casual on-location speech
SCORING CONTRACT
01WER first. Report raw and normalized Word Error Rate for every dataset, then a macro average. Lower is better.
02Preserve the language. Normalize Unicode and whitespace consistently while retaining meaningful Yiddish letters and diacritics.
03Show the work. Record the exact revision, engine, decoding configuration, transcripts, and run artifacts.
04Keep the test clean. Do not train or tune on the benchmark clips.
ABOUT THE PROJECT
ייִדיש האָט פֿאַרשיידענע דיאַלעקטן, אויסלייג־סיסטעמען, לשון־קודש־ווערטער און אָפֿט ענגלישע ווערטער אינמיטן זאַץ. דער לוח שאַפֿט אַן אָפֿענעם און איבערחזרלעכן וועג צו פֿאַרגלײַכן מאָדעלן אויף אמתע רעקאָרדירונגען.
This is a community benchmark, not a product ranking. Missing measurements are never converted to zero, and no score appears merely because a model exists.
Built under Yiddish-AI with the public OmniASR one-hour derivative and Kohn-AI multi-genre benchmark. Inspired by the open-source ivrit-ai Hebrew leaderboard.
CONTRIBUTE
Submit the exact service version, raw and normalized WER for each dataset, and the returned transcripts.
Submit a result ↗