Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
NVIDIA AVO Agent Hits 100% on ARC-AGI-3 Benchmark - OpenSmartRoute
NVIDIA's AVO agent system scored 100% on ARC-AGI-3 using Claude Opus 5. It completed all levels with fewer actions than the VISTA system.
Key points
AVO achieved a perfect 100.00 RHAE score on ARC-AGI-3.
The agent finished all 183 levels using 6,624 environment actions.
VISTA required 7,542 actions to complete the same 183 levels.
AVO used approximately 12% fewer actions than VISTA with the same model.
Why it matters: System design enables frontier models to sustain long-horizon work better than raw capability alone.
By OpenSmartRoute editorial · written through the router by writer-small
From NVIDIA technical blog - “NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents”
NVIDIA AVO Agent Achieves Perfect Score on ARC-AGI-3
NVIDIA released a new agent system called Agentic Variation Operators, or AVO. This system scored 100% on the ARC-AGI-3 benchmark using Claude Opus 5. The score represents Relative Human Action Efficiency, known as RHAE. It measures how efficiently an agent completes tasks compared to humans.
The agent finished all levels in every environment without failure. There were 25 different game environments in the test set. The system solved a total of 183 individual levels across these games. This result shows that the full agent architecture achieved perfect performance.
The Role of System Design Over Raw Model Capability
A frontier language model is only one part of an AI agent. The surrounding system determines how the model works in practice. Engineers often focus too much on the model itself while ignoring the system.
The AVO project proves that system design can unlock high performance. It does not rely solely on the raw intelligence of the underlying model. The architecture handles memory, tools, feedback, and recovery automatically.
How AVO Handles Long-Horizon Autonomous Tasks
Long-horizon tasks require sustained operation over many steps. The AVO system uses persistent memory to keep track of past work. It stores prior implementations and evaluation results for later use.
A supervisor component monitors the overall search trajectory. It detects when progress stalls or cycles repeat unproductively. The supervisor can redirect the main agent toward new strategies if needed.
Key Performance Metrics and Efficiency Gains
The benchmark uses Relative Human Action Efficiency, abbreviated as RHAE. This metric combines task completion with action efficiency relative to human baselines. AVO achieved a score of 100.00 RHAE on all levels.
The system completed the full public set using approximately 12% fewer actions than VISTA. Both systems used the same underlying model, Claude Opus 5. This highlights significant gains in efficiency through better system design.
Reflection AI launched Beam, a 501-billion-parameter text-only model designed to match Chinese open models in reasoning while using significantly less compute.
NVIDIA released version 1.0 of its AI Cluster Runtime to standardize GPU cluster configurations. The update adds signed validation evidence and a live dashboard for operators.
Comparison with VISTA and Other Agent Architectures
VISTA is another agent architecture that competed on this benchmark. It also uses Claude Opus 5 to solve the same 183 levels. However, VISTA required more environment actions to reach completion.
The observation interface differs between the two systems. VISTA often renders images while AVO uses text-only grids. This difference in input modality affects how the model processes information.
Why System Architecture Matters for Real-World Agents
Real-world agents face complex environments with hidden rules and goals. They must infer objectives through interaction rather than receiving explicit instructions. The ARC-AGI-3 benchmark tests this exact capability.
System architecture determines how effectively a model converts its capabilities into progress. A good design preserves useful knowledge across long sessions. It prevents the agent from wasting time on repeated mistakes.
What Engineers Can Learn from the AVO Approach
Engineers should build trusted agent stacks with performance and reliability in mind. They must design these properties across the full system, not just the model. The AVO approach emphasizes autonomous evolution through supervised search.
Read the official paper for technical details on the Agentic Variation Operators architecture. Engineers can explore the ARC-AGI-3 benchmark to understand the evaluation framework better. Reviewing the scoring methodology helps clarify how RHAE is calculated.
The research demonstrates that model capability alone does not guarantee success. The surrounding harness determines how reliably a model works on extended tasks. Engineers should consider system-level mechanisms when building autonomous agents.
How it Compares to Previous Architectures
AVO replaces the fixed variation steps found in older evolutionary search systems. Traditional methods often require humans to define every possible change direction. AVO allows an agent to decide its own next move. This shift moves control from a predefined script to autonomous decision-making. The agent inspects code, forms hypotheses, and executes tests itself. It repeats this cycle without needing constant human intervention.
Before AVO, agents struggled with long-horizon tasks due to memory loss. Older systems frequently forgot context after many steps. They could not maintain state across days of work. This caused them to restart efforts or make redundant mistakes. The previous generation relied heavily on short-term memory buffers. These buffers filled up quickly during complex optimization runs.
AVO introduces persistent memory as a core architectural feature. It stores past implementations and evaluation results permanently. The agent retrieves this history when faced with new problems. This prevents the need to reconstruct the entire search from scratch. The system remembers what worked and what failed in previous iterations.
The supervisor component also differs significantly from earlier designs. Past agents lacked a dedicated monitoring layer for trajectory analysis. They often continued down unproductive paths without realizing it. AVO's supervisor detects stagnation patterns automatically. It can redirect the main agent toward better strategies when needed. This dual-layer design adds robustness to the overall operation.
Both systems use Claude Opus 5 as their underlying language model. The core reasoning engine remains identical across architectures. However, the wrapper around the model differs greatly in function. VISTA focuses heavily on visual rendering and image-based inputs. AVO prioritizes text-only grids and command-line execution. This input modality difference changes how information flows into the system.
VISTA often renders full images to present problems to the agent. The agent must process these complex visual signals to understand the task. AVO uses simplified text representations of the environment state. This abstraction allows for faster processing and clearer logical reasoning. The trade-off involves losing some visual detail in exchange for computational efficiency.
The ARC-AGI-3 benchmark evaluates both systems on the same 183 levels. Both agents attempt to solve identical problems within the same constraints. Yet, their performance metrics diverge significantly due to architectural choices. AVO achieves a Relative Human Action Efficiency score of exactly 100.00 RHAE. VISTA falls short of this perfect efficiency mark consistently.
AVO completed all public-set environments using fewer total actions than VISTA. The system saved approximately 12% in environment actions compared to its competitor. This reduction comes from smarter planning and execution strategies. AVO avoids redundant steps that VISTA might take by default. The agent learns which paths lead to progress more quickly.
The underlying model capability is not the deciding factor here. Both systems run on the same foundation of Claude Opus 5. The difference lies entirely in how they orchestrate the model's output. System design determines whether the model reaches its full potential. Architecture dictates the reliability and speed of long-running tasks.
AVO demonstrates that autonomous evolution requires specific system-level mechanisms. It combines persistent memory with active supervision to drive progress. Older architectures lacked one or both of these critical components. They could not sustain the same level of autonomy over extended periods. The result is a clear hierarchy in performance between generations.
Questions This Leaves Open
The source text does not specify the exact hardware configuration used for the ARC-AGI-3 run. It mentions NVIDIA DGX B200 systems for the kernel optimization phase only. Readers cannot confirm if the same GPU cluster handled the full benchmark suite. The number of GPUs involved in the 183-level completion remains unknown.
The paper does not disclose the total token count consumed during the ARC-AGI-3 evaluation. Engineers interested in cost efficiency need this data to calculate expenses. Without it, they cannot compare AVO's operational costs against other agents. Token usage is a primary metric for budgeting AI projects.
There is no information about the specific failure modes encountered by VISTA. We only know it required more actions, not why those extra steps were needed. Did VISTA struggle with visual parsing or logical reasoning? The text omits details on where its efficiency dropped off compared to AVO.
The source does not state how long a single task took for each agent to complete. It focuses on total action counts rather than elapsed time. Speed and latency figures are absent from the provided metrics. Readers cannot determine if AVO was faster per level despite using fewer actions.
AVO's ability to generalize beyond GPU kernels remains unverified in the text. The paper highlights success in software engineering and kernel optimization tasks. It does not list other domains where the architecture has been tested or validated. General-purpose claims need evidence from diverse application scenarios.
The exact threshold for supervisor intervention is not defined in the article. We do not know how many failed attempts trigger a strategy redirect. This parameter could significantly impact resource usage and task success rates. Engineers building similar systems must define their own intervention logic.
The paper does not mention any safety checks or guardrails implemented by AVO. Frontier agents often face risks related to code execution or system access. The text focuses on performance metrics without addressing potential hazards. Readers should verify the safety protocols used in production deployments.
Token budget limits for the agent are not specified in the source material. Knowing the context window size and token limits is crucial for model selection. Engineers need this data to ensure agents do not hit memory ceilings during long runs. The absence of these numbers creates uncertainty about scalability.
The cost per action or per level is not provided in the article. Calculating true economic efficiency requires knowing the price of every operation performed. Without pricing data, managers cannot perform a full cost-benefit analysis. They lack the inputs needed to justify the investment to stakeholders.
The source text does not detail how AVO handles non-public or private environments. The ARC-AGI-3 benchmark uses public sets for evaluation purposes. Real-world deployments often involve proprietary data and restricted access. Adapting the architecture for such scenarios requires further investigation.
Readers cannot determine if the 100 RHAE score is a new record or just a high baseline. Contextualizing this achievement against historical benchmarks helps assess its significance. The article does not provide comparative scores from prior years or other organizations. This limits the ability to measure progress over time accurately.
The paper does not explain how AVO handles tasks requiring external API access beyond standard tools. Many real-world agents must interact with third-party services and databases. The scope of supported tools in the current architecture remains unclear. Engineers need to know if custom integrations are feasible within the system design.
There is no information on the training data used to develop the Claude Opus 5 model. Understanding the model's knowledge base helps explain its performance on specific tasks. The source assumes prior knowledge of the model without providing details. This creates a gap in understanding the root capabilities being tested.
The article does not specify the version numbers for any software components involved. Software versions change frequently and can affect reproducibility of results. Engineers need exact versions to replicate the experiments or audit the codebase. Missing version data hinders verification and transparency efforts.
Readers cannot confirm if the supervisor agent operates in real-time or with delays. Latency between the main agent and the supervisor impacts overall system responsiveness. The text does not clarify the timing of supervisory interventions during execution. This affects how the system handles dynamic changes in the environment.
The source does not mention any human-in-the-loop requirements for specific tasks. Some safety-critical applications mandate human approval before executing code changes. AVO's fully autonomous claim needs qualification regarding acceptable risk levels. Engineers must define their own boundaries for autonomy based on use cases.
There is no data on how the agent handles conflicting instructions or ambiguous goals. Real-world problems often present contradictory requirements that confuse models. The text assumes clear objectives without addressing edge cases involving ambiguity. This limits the practical applicability of the architecture in messy environments.
The paper does not disclose the computational resources consumed per optimization iteration. Energy usage and carbon footprint are growing concerns for large-scale AI deployments. Without power consumption figures, it is impossible to assess the environmental impact of running AVO.
Readers cannot determine if the performance gains scale linearly with problem complexity. Simple tasks might benefit from AVO's design while complex ones do not. The article focuses on a specific set of 183 levels without broader scalability evidence. Engineers need data across a wider range of difficulty levels to judge suitability.
The source text does not mention any security audits or third-party reviews of the system. Independent validation adds credibility to claims about safety and reliability. The absence of such reports leaves questions about potential vulnerabilities unanswered. Managers should seek external verification before deploying critical systems.
There is no information on how long it takes to deploy AVO compared to VISTA. Deployment time affects operational readiness and time-to-market for new projects. Engineers need lead times to plan their infrastructure upgrades effectively. The article omits these logistical considerations entirely.
The paper does not specify the minimum hardware requirements to run AVO locally. Cloud vs. on-premise decisions depend heavily on resource availability and cost structures. Without minimum specs, engineers cannot size their clusters correctly for this architecture. This creates uncertainty in procurement planning processes.
Readers cannot confirm if the agent can operate without internet connectivity. Offline capabilities are essential for many industrial and remote applications. The text assumes networked environments without addressing disconnected scenarios. This limits the versatility of the solution for certain use cases.
The source does not mention any regulatory compliance standards met by the system. Industry regulations vary widely across different sectors and regions. AVO must adhere to local laws regarding AI usage and data privacy. Unaddressed compliance issues pose significant legal risks for adopters.
There is no data on how the system handles user feedback during operation. Human correction loops are vital for continuous improvement of agent behavior. The text does not explain mechanisms for incorporating external human input into the workflow. This affects adaptability to evolving user needs over time.
The article does not specify the cost of GPU instances used for the seven-day run. Operational expenses can quickly exceed initial development costs if not managed well. Managers need clear pricing models to forecast ongoing budget requirements accurately. Missing financial data makes ROI calculations impossible without assumptions.
Readers cannot determine if the agent's decisions are explainable or interpretable by humans. Transparency is a key requirement for trust in autonomous systems. Black-box decision making raises concerns about accountability and debugging difficulties. The source provides no insight into the reasoning processes used internally.
Simon Willison tested Claude Opus 5.5 on composing Monkey Island-style game music. The model produced surprisingly high-quality results in a text-based format.