The study investigates improving LLM tool agents by altering their runtime harness, focusing on prompts and tool-boundary middleware. The protocol evaluates changes through mean held-out lift, worst-condition lift, repeatability, and RelLift95, providing a conservative estimate of the gain. PRISM, a clustering and routing optimizer, was instantiated on multiple benchmarks, achieving mean held-out lifts of 14.2, 14.9, and 10.1 percentage points on BFCL, tau2-Retail, and tau2-Telecom respectively. The results attribute the primary margin to failure-surface routing and the edit-pattern constraint. The research highlights the importance of reporting harness reliability alongside average held-out lift, as some search procedures can identify large gains but produce brittle updates.
Source: https://arxiv.org/abs/2609.05736