Research1 min read
$ au^ au$-Bench: A Benchmark for End-to-End Agent Construction
The $ au^ au$-bench evaluates agent building from real business data, requirements, and APIs, measuring performance across multiple tasks to reflect real client engagement conditions.
From arXiv cs.AI
