The research introduces a method for integrating language model forecasts into event forecasting systems. The core principle is to assess the marginal value of a model’s output relative to existing forecasts, rather than simply evaluating model accuracy. The system employs a competence gate that estimates domain-level source weights based on resolved outcomes. This gate then shrinks uncertain estimates toward a global weight and recalibrates the pooled forecast.
Experiments were conducted across 2,357 resolved binary questions using five language models. The competence gate improved the main external baseline from 0.0771 to 0.0732 Brier score. This improvement was significant even with leakage controls and against a leakage-safe time-series prior on the pooled structured set, including data from FRED.
Across the ForecastBench market subset, the gate provided no significant improvement, deferring largely to the market forecast. Analysis of four Qwen models revealed that verbal confidence did not reliably identify model outperformance compared to the external forecast. Outcome-estimated competence, however, supported better abstention decisions.
The research provides a practical approach for selective model use, focusing on measured marginal value. This allows for more efficient integration of language models into forecasting workflows. Source: https://arxiv.org/abs/2609.12101