The research identifies a key issue with existing word error rate (WER) metrics in medical speech recognition: the denominator, often reliant on named-entity recognition (NER) models, introduces variability and obscures clinically relevant errors. The protocol introduces MedWER, an evaluation tool utilizing a fixed term list of 19,373 drug, diagnosis, symptom, and injury-mechanism entries derived from public sources. This list serves as the denominator, ensuring consistency across evaluations. The protocol combines a pinned text normalizer with a phrase-aware term-restricted WER, focusing on accuracy within the defined term list. Validation against an independent provincial drug-benefit file confirms the list’s accuracy. Results are reported with 95% confidence intervals from resampled per-utterance scores, comparing Moonshinebase, Whisperbase.en, and MedASR on two open benchmarks.
Source: https://arxiv.org/abs/2609.05728