NVIDIA AVO agent hits 100% on ARC-AGI-3, solving all 183 levels

NVIDIA's AVO reached a perfect 100% on ARC-AGI-3, an interactive reasoning benchmark, solving all 183 levels across the suite's 25 public games. The company's framing is as notable as the score: AVO is an agent architecture, and NVIDIA argues explicitly that a frontier language model is only one component — the surrounding agent harness (planning, memory, tool use, error recovery) determines real-world performance.
That thesis reframes the frontier race away from raw model IQ and toward engineering the scaffolding around the model. Hugging Face CEO Clement Delangue amplified the point, noting NVIDIA 'built its own coding harness to optimize CUDA GPU kernels' and predicting a shift from a world where optimizing and post-training your own models is hard to one where it's routine. NVIDIA released a 550B instruction-following 'teacher' model on Hugging Face in the same window, reinforcing the toolchain narrative.
The result lit up developer forums. On r/singularity the story drew 1,152 upvotes and 205 comments, with debate centering on exactly NVIDIA's point — is a 100% score a triumph of the harness or does it expose ARC-AGI-3 as gameable once you throw enough agent engineering at it? Skeptics argue a saturated benchmark tells you more about the benchmark than about general intelligence.
Strategically, AVO fits NVIDIA's broader 2026 pivot: alongside the Poolside deal and Nemotron push, it signals the company wants to own not just chips and models but the agent frameworks that determine how much useful work a watt of compute produces — a theme echoed in its DSX MaxLPS performance-per-watt messaging. What to watch is whether AVO ships as a product or remains a research demonstration.