The research introduces WinSyn, an automated pipeline designed to address limitations in current question-answering benchmarks. Existing benchmarks often lack the complexity and realism of enterprise environments, relying on short-form questions and unnatural queries. This work focuses on generating synthetic datasets that more closely mirror the challenges of enterprise settings, including distributed information across emails, chat messages, and documents. The pipeline simulates long-running projects involving up to 25 interacting employees across multiple roles, spanning several months. The generated data incorporates ambiguity and naturally occurring queries. Evaluations were conducted using standard agentic baselines and the latest frontier models on the generated datasets. Aggregate scores averaged across all queries remained below 80% for each dataset. This indicates a significant gap between current agent performance and the demands of realistic enterprise deployments. The findings emphasize the importance of high-complexity evaluation data for developing robust real-world enterprise DR systems.
Specifically, the pipeline generates long- and short-form questions with gold answers grounded in the simulated enterprise data. The simulations involve multiple interacting employees, reflecting the dynamic nature of workplace communication. The system’s output includes data representing evolving and potentially conflicting information. The research suggests that further development is needed to improve agent performance in enterprise question-answering tasks.
The evaluation utilized standard agentic baselines and the latest frontier models. The datasets generated by WinSyn represent a significant step towards more realistic and comprehensive evaluation of enterprise question-answering agents. The results provide a clear indication of the challenges that remain in deploying these systems effectively.
Source: https://arxiv.org/abs/2609.12171



