SEA-SpeechBench is the first large-scale multitask benchmark for speech understanding in 11 Southeast Asian languages. The benchmark contains 97,194 samples across 99 evaluation sets and 597 hours of audio data. It includes nine diverse tasks across three categories: speech processing, paralinguistic analysis, and temporal understanding. Temporal understanding involves timestamped content queries and localization within audio sequences up to 3 minutes. The benchmark utilizes multilingual prompting in native SEA languages and English. Evaluation results show performance gaps across all models, particularly in temporal understanding, emotion recognition, and speech translation. Prompting in low-resource languages such as Burmese and Tamil lags behind English by up to 41 percentage points. This highlights limitations in current models and the need for more inclusive development.
Source: https://arxiv.org/abs/2609.09672