The study investigated document-level machine translation (MT) evaluation, a method that presents entire documents to human annotators. The research hypothesis was that this approach would elicit document-level judgments, leading to distinct results compared to segment-level evaluations. A counterfactual condition, termed MIX, was employed to test this assumption. This condition combined segments from different MT systems within a single document, maintaining the document-level presentation but disrupting cross-segment consistency. Across 18,420 English-to-Korean annotations and 14 automatic metrics, no statistically significant differences were observed between coherent and incoherent documents. Raters identified the coherent passage as originating from a single translator in 87.3% of trials when presented with matched passages. The study concluded that the protocol itself, rather than the annotator, is the key factor driving the observed equivalence. This raises concerns about the effectiveness of current resources invested in document-level systems, metrics, and annotation.
Source: https://arxiv.org/abs/2609.05949