AutoFyn introduces a novel approach to agent refinement. It operates by establishing a fresh model session each round, utilizing persistent state updated solely through interfaces like memory files and repository state. An orchestrator explores and plans using specialized agents, while a task-grounded verifier assesses the work and provides an objective reward. This reward is then incorporated into the persistent state, influencing the effective policy for subsequent rounds. The system demonstrates this process across three domains: olympiad mathematics, data science, and cybersecurity. Initial results indicate improved performance compared to provider agents on the 2026 International Mathematical Olympiad problems. Furthermore, AutoFyn achieved the top-ranked agent on the Spider 2.0 dbt benchmark. The system has also generated $16$ maintainer-confirmed vulnerability advisories across various projects including Next.js, MetaMask, pnpm, Warp, LiteLLM, Langflow, and Open WebUI.
Source: https://arxiv.org/abs/2609.05446