A study evaluated Google Gemini 3.0 Pro’s ability to detect planted errors in a corpus of 150 academic papers. The corpus contained 450 known contaminants, categorized as typographical corruption, semantic reversal, and absurd out-of-context insertion. The evaluation tested three prompting regimes: single document, small batch, and large batch. Detection performance initially held at small scales, reaching 50% and 60% recovery rates for single documents and small batches, respectively. However, at large batch sizes, recovery dropped to 2.8%. The failure mode was not abstention but confident fabrication, with the model generating invented contaminants like "telepathic squirrel" and "quantum-powered toaster". Detection varied by contamination type, with absurd insertions recovered at 75%, while semantic reversals and typographical corruptions were each recovered at 50%. The research suggests that LLMs degrade detection performance deceptively with increasing batch sizes. The findings underscore the importance of constrained batch sizes, direct content injection, and mechanical verification of reported findings.
Source: https://arxiv.org/abs/2609.09696