The study addresses the challenge of query-to-agent matching failing due to the confusion between semantic relevance and actual executable capability. The research proposes a framework of "capability-bound process supervision" to guide annotation. Debate-to-Skill utilizes reusable decision principles, structured deliberation, and a verifier for verdict extraction, incorporating disagreement-driven refinement. Comparisons were conducted on a Query2Agent benchmark against direct-label supervision, reasoning-SFT, and structural ablation experiments. The results indicate that gains are realized through supervising the critical capability-determining decision process, especially in ambiguous cases where semantic similarity and executable ability differ. This approach is relevant for engineers managing models and agents in production environments.
Source: https://arxiv.org/abs/2609.11176