Nvidia’s AVO Wrapper Just Took Claude from 30% to a Flawless 100% on ARC-AGI
It turns out today's flagship language models aren't fundamentally broken at solving logic puzzles; they just desperately needed a competent digital manager to stop them from running in circles.
Engineers at Nvidia initially developed an agentic architecture dubbed AVO (Agentic Variation Operators) purely to optimize their internal graphics chip layouts. After noticing how well it cracked stubborn engineering bottlenecks, the team decided to unleash it on the ARC-AGI-3 test — a puzzle benchmark famous for letting average humans breeze through while leaving giant neural networks completely baffled.
Instead of retraining the models from scratch, Nvidia slapped AVO on top as an autonomous strategic supervisor. Tested with heavyweights like Claude Opus 5 and GPT-5.6 Sol, the architecture acts as an executive steering wheel: it plans long-horizon moves, catches recursive loops before the AI melts down, and dynamically scraps broken plans in favor of fresh hypotheses.
The intervention took baseline models struggling at a mediocre 30% mark and rocketed their problem-solving success rate straight to a clean 100% on the benchmark.
Multi-billion dollar foundational models were essentially sitting on massive untapped reasoning power, waiting for an external scaffolding script to do the actual thinking for them.
Source: Nvidia Developer Blog
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.