LLMs1 min read
GAUGE: Evaluating LLM Judges in Agent Evaluation
Research reveals a significant disconnect between LLM-as-a-judge rankings and actual task success in user-simulated evaluations of task-oriented agents. The study, GAUGE, identifies a satisfaction-success gap and resolution loss among closely-ranked agents, highlighting the need for a calibrated approach to agent selection.
From arXiv cs.CL
