YI/ASR Submit a run ↗

OPEN BENCHMARK · 2026

Yiddish speech,
measured clearly.

A direct comparison of five commercial and open-source Yiddish transcription systems across shared test sets.

ייִדיש
BENCHMARKS2 datasets · ≈113 min
SYSTEMS3 commercial · 2 open
STATE0 / 5 scored

CURRENT RESULTS

Leaderboard

Raw and normalized WER are entered manually after each service returns its transcript.

#SystemTypeOverall norm.Omni 1h rawOmni 1h norm.Multi-genre rawMulti-genre norm.Status

SYSTEMS IN SCOPE

The comparison field

Three hosted commercial services and two reproducible open models. Exact versions are recorded with every result.

    THE TEST SETS

    Two datasets.
    One scoring rule.

    A one-hour single-file stress test plus a multi-genre collection reveal different ASR strengths.

    Download the 1-hour MP3 ↗
    A

    OmniASR Yiddish · 1 hour

    198 held-out test utterances joined into one upload-ready MP3 · 6,899 reference words

    60:00
    B

    Yiddish Multi-Genre

    Six Yiddish24 clips covering shiur, news, interview, podcasts, and vlog speech

    ≈53m
    01

    Torah shiur

    Long-form religious discourse

    תורה
    02

    News

    Three joined bulletins

    נײַעס
    03

    Interview

    Multi-speaker conversation

    שמועס
    04

    Monologue

    A single-speaker podcast episode

    רעדע
    05

    General podcast

    Informal program audio

    פּאָדקאַסט
    06

    Vlog

    Casual on-location speech

    ווידעאָ

    SCORING CONTRACT

    Comparable by design.

    01WER first. Report raw and normalized Word Error Rate for every dataset, then a macro average. Lower is better.

    02Preserve the language. Normalize Unicode and whitespace consistently while retaining meaningful Yiddish letters and diacritics.

    03Show the work. Record the exact revision, engine, decoding configuration, transcripts, and run artifacts.

    04Keep the test clean. Do not train or tune on the benchmark clips.

    ABOUT THE PROJECT

    Progress needs
    a common measure.

    פֿאַר וואָס אַ ייִדישער ASR לוח?

    ייִדיש האָט פֿאַרשיידענע דיאַלעקטן, אויסלייג־סיסטעמען, לשון־קודש־ווערטער און אָפֿט ענגלישע ווערטער אינמיטן זאַץ. דער לוח שאַפֿט אַן אָפֿענעם און איבערחזרלעכן וועג צו פֿאַרגלײַכן מאָדעלן אויף אמתע רעקאָרדירונגען.

    Transparent by default.

    This is a community benchmark, not a product ranking. Missing measurements are never converted to zero, and no score appears merely because a model exists.

    Built under Yiddish-AI with the public OmniASR one-hour derivative and Kohn-AI multi-genre benchmark. Inspired by the open-source ivrit-ai Hebrew leaderboard.

    CONTRIBUTE

    Help establish the first baseline.

    Submit the exact service version, raw and normalized WER for each dataset, and the returned transcripts.

    Submit a result ↗