The paper introduces GAUGE, a protocol designed to assess the validity of using LLM-as-a-judge in evaluating task-oriented agents. The protocol evaluates ranking validity and construct validity across 25 agents from six providers using the $ au^2$-bench and SimulatorArena benchmarks. A satisfaction-success gap was observed, where conversations rated as ‘satisfied’ by a blind panel were often unsuccessful in completing the customer’s task. Specifically, 57.5% of these conversations resulted in task failure, a pattern consistent across five rater populations, both benchmarks, and all subjective dimensions rated.
The study found that while the LLM-as-a-judge ranking remained robust across a broad range of agent capabilities, it lost resolution when comparing agents with similar performance. The decision-disagreement rate increased from less than 1% when evaluating wide reward pairs to 31% when evaluating close reward pairs. This suggests that the LLM-as-a-judge is not consistently differentiating between high-performing agents.
Researchers validated the protocol with human raters, confirming the observed discrepancies. The findings demonstrate that the LLM-as-a-judge is human-validated but mis-anchored to actual task success. The research proposes a ‘calibrate-then-trust’ cadence, suggesting a judge-free completion bit as a zero-cost tripwire for detecting regressions in agent performance.
This work provides critical insights for engineers deploying task-oriented LLM agents. The observed limitations underscore the importance of supplementing LLM-as-a-judge evaluations with more robust, ground-truth metrics. The protocol offers a framework for understanding and mitigating the risks associated with relying solely on LLM-based scoring.
Source: https://arxiv.org/abs/2609.12191



