A study investigates enhancing LLM tool agents by adjusting the runtime harness surrounding a fixed model. This approach focuses on prompt and tool-boundary middleware modifications, avoiding retraining. The protocol measures performance using mean and worst-condition held-out lift, repeatability, and RelLift95(B). It instantiates this protocol with prompt-only and prompt-plus-middleware optimizers, including PRISM, which identifies failures and directs repairs to prompts, middleware, or joint edit surfaces. On BFCL, tau2-Retail, and tau2-Telecom benchmarks, PRISM achieved mean held-out lifts of 14.2, 14.9, and 10.1 percentage points, respectively. The study attributes the primary margin to failure-surface routing and the edit-pattern constraint. Results indicate that some search procedures find large gains but may select brittle updates, emphasizing the need to report harness reliability alongside average held-out lift.
Source: https://arxiv.org/abs/2609.05736