The TamilEOT dataset consists of 18,485 labelled turn boundaries extracted from 116 Tamil telephone conversations. Two audio-only detectors were fine-tuned from Smart Turn v3. Performance on a held-out split of 4,168 clips improved from 70.30% zero-shot accuracy to 83.71% (8.7 MB) and 86.13% (21 MB) as measured by ROC-AUC. The models run in under 150 ms on a laptop CPU.
Rule-derived labels, validated by human listening, had a 95.9% accuracy rate for positive labels and 44.4% for negative labels, below chance. Replacing these rules with an audio-LLM labeller achieved 97.5% human agreement. The cost of building the dataset was reported.
Only encoder capacity changes significantly impacted the results, with three identical runs yielding an accuracy floor of 0.87. Replaying the labelled boundaries through the production VAD and streaming adapter added a further 2.60 points to the accuracy. 7.8% of boundaries were never surfaced to the model. Data, weights, code and negative results are publicly available.
Source: https://arxiv.org/abs/2609.05631