arXiv introduces $ au^ au$-bench, a benchmark focused on constructing customer-service agents from actual business records, client requirements, and production APIs. It provides a realistic environment for testing agent development processes.
The benchmark presents 53 tasks across four domains, with the best configuration passing 23.9% of evaluation simulations. An expert reference scores 82.2%, highlighting the gap between current models and human performance.
Failures observed include shallow queries, limited client communication, and minimal experimentation with architecture and costs. This benchmark aims to turn agent building into a measurable, standardized task for coding agents.
Source: https://arxiv.org/abs/2609.04611