MeshioMeshio
News

ArmBench-ASR Brings Standardized Testing to Armenian Speech Recognition

A new benchmark tests nearly 30 Armenian ASR systems across five speech domains, exposing gaps hidden by single-dataset evaluations.

Meshio Newsroom
Meshio NewsroomAug 23, 2026
ArmBench-ASR Brings Standardized Testing to Armenian Speech Recognition

Armenian speech recognition now has a broader yardstick. ArmBench-ASR, a new benchmark from Metric-AI and contributors, evaluates almost 30 open-weight and closed systems across 10,113 clips totaling about 20.7 hours of audio.

Rather than relying on one test set, the benchmark combines five sources: crowdsourced and public read speech from Common Voice 26 and FLEURS, plus expressive poetry, movie dialogue, and narrated news. The latter three collections include more varied conditions, including background music or noise. References were also reviewed and corrected for the evaluation.

What the results show

Google’s Gemini 2.5 Pro leads the initial leaderboard on strict combined word error rate (WER) at 14.31%, followed by HiSpeech’s Armenian conversational model at 16.81%. Closed systems occupy the eight lowest combined WER positions, while NVIDIA’s Armenian FastConformer is the highest-ranked open model at 20.21%.

The results underline why AI builders should avoid treating a single benchmark score as a complete picture. Movie audio was the hardest category for every evaluated model, with a median strict WER of 61.87%, compared with roughly 17% on Common Voice 26 and FLEURS. Rankings also shift by domain, meaning a model tuned for read speech may not transfer cleanly to conversational or expressive audio.

ArmBench-ASR reports both strict and more heavily normalized WER and character error rate (CER). That distinction helps separate recognition failures from differences in punctuation, capitalization, spacing, and orthographic conventions. Its interactive leaderboard lets developers inspect results by dataset, metric, and model availability—useful for choosing systems for real Armenian voice applications rather than optimizing for one narrow test.

Source: Hugging Face Blog

Comments

Log in to join the discussion