Drop a large language model into an entirely unfamiliar 2D mini-game with no rules provided and no objective given. All the model can do is navigate using directional keys and click grid cells to probe the environment’s feedback, deducing on its own the game’s underlying mechanics and winning logic. This is the testing ground set up by the interactive reasoning benchmark ARC-AGI-3. In this benchmark, widely acknowledged as exceptionally challenging, frontier large models previously scored only around 40%. Yet an MIT team developed an external scaffolding system, VISTA, that scored 100% across all public test environments. The paper attributes this breakthrough to unlocking visual capabilities, arguing that models inherently possess strong reasoning abilities and that their previous mediocre scores were primarily due to the limitations of plain-text numerical grids. However, the benchmark’s lead designer offered a completely different assessment, stating unequivocally that achieving a perfect score by manually building a system around the benchmark tasks does not prove the model itself possesses general intelligence.
The leap from a 40% score straight to 100% sparked considerable discussion in the tech community regarding model reasoning capabilities, bringing the evaluation value of benchmarks back into focus. To see clearly how this 100% was actually produced, several layers of facts need to be untangled: how the test itself defines actions and scoring, what specific changes the external scaffolding system actually introduced, and what principles the benchmark creator originally established when drawing the evaluation lines.
A large language model is fundamentally a computational engine that takes in context and predicts subsequent tokens; on its own, it cannot see a screen or click a mouse. To enable a model to participate in such interactive games, engineers must construct helper software around it—commonly known as a harness. This external scaffold is responsible for receiving state signals emitted by the game environment, formatting them into representations the model can interpret, and translating the decision text output by the model into concrete movement or click actions within the game. The model never directly touches the game itself; all inputs and outputs are relayed entirely through this external transceiver system.
In the VISTA paper published by the MIT team, they made no changes to the model’s internal weights and added no additional fine-tuning; all modifications occurred entirely within the external harness layer. They introduced three changes to the model, corresponding to seeing, remembering, and inspecting.
Assisted by this suite of external tools, a frontier model cleared all 25 publicly available game environments, delivering a perfect scorecard. Not only did it achieve a clean sweep across all tasks, but the total number of environment interaction actions consumed to complete them was roughly half that of human players playing for the first time. The paper formulated an explanation based on this: frontier models already possessed spatial understanding and long-horizon reasoning abilities all along, and their prior mediocre benchmark performance was entirely due to clumsy, lossy text representations and frequent context truncations tying their hands. In the authors’ framing, VISTA simply removed obstacles in visual perception, unleashing capabilities the model inherently possessed.
However, the body of the paper analyzes the impact of each enhancement on the score in detail, observing score changes by introducing one modification at a time. Rather than supporting it, these analyses weaken the abstract’s claim attributing all credit to visual contributions, for two reasons.
The first reason is that replacing the original sequence of numbers with generated color images yielded a substantially higher score. The paper itself hypothesizes that this score gain stemmed primarily from saving context space: an image consumes only a few hundred tokens, whereas a board written as numbers requires several thousand tokens. A flood of numbers quickly fills up the context, leaving less room for reasoning. In other words, the credit for this step lies in compacting information, having little to do with letting the model see the visual scene itself.
The second point is even more critical. After putting lossless memory and inspection tools in place, they swapped the images back to plain-text numerical grids; the score barely dropped, requiring only a few more actions. If the cause of failure had truly been that the model could not see, the text version should not have performed on par with the image version. What truly underpinned the perfect score was the external engineering of memory and tools; vision was merely a more space-efficient representation. In the limitations section, the paper also conceded a caveat: these 25 public games had been online for a long time, and they could not rule out the possibility that the tasks had already entered the model’s training data.
What this experiment actually demonstrates subtly diverges from the immediate impression conveyed by the abstract. The core driving force behind the system’s perfect score was not merely enabling the model to understand images; the far greater impetus came from the lossless archiving mechanism and inspection toolchain built by the engineering team for the interactive environment. This perfect score, pushed upward layer by layer through external scaffolding, quickly prompts a deeper question: the more human scaffolding added around the exam tasks, whose capability is the resulting high score actually measuring?
To understand why the benchmark creator remains skeptical of this perfect score, one must return to the original design intent of the test. François Chollet, who led the design of this benchmark, has long been dedicated to building technical yardsticks that can genuinely measure general intelligence. In the conceptual framework of the official methodology, the core focus of the test is to evaluate the on-the-fly reasoning and adaptive efficiency of artificial intelligence when entering an entirely unfamiliar new environment without prior preparation. It deliberately steers clear of testing memorized existing knowledge, locking its evaluative focus squarely on the ability to explore in real time and induce underlying rules.
The leaderboard metric is called Relative Human Action Efficiency (RHAE), and its algorithm is straightforward. For a given game, the number of actions taken by multiple first-time players to clear each level is recorded, with the upper median taken as the human baseline for each level. Then, the number of steps taken by the model on each level is counted. If the model takes fewer steps, its score is higher; if it takes more, its score is lower; failing to clear a level yields a score of zero. Within each game, levels are combined via a weighted average by level index; the total score is then the average across all games, capped at 100.
Crucially, however, what is counted here are only actions that actually alter the game state—such as pressing an arrow key to move a character one cell, or clicking a tile to trigger a mechanism. How much thinking the model conducts in the background, how many tool calls it makes, how many pixels it reads, or how much scratchpad trial-and-error it performs do not count as long as they do not change the game itself. The original intention behind this design was to provide a fair baseline for models with different compute budgets by looking solely at the efficiency of final moves. Yet it also left a loophole: an external system can run extensive search, reasoning, and local troubleshooting in the background, only executing the single critical move precisely on the game board, thereby keeping action counts extremely low.
To ensure the long-term validity of tasks, the organizers split the test suite into two parts: a public test set and a private test set. The former is open source, allowing anyone to download it, inspect the source code, and run programs; the latter is held securely by the organizers for official blind evaluations and competitions. The officially published evaluation policy states that scores on the public test set will never enter the official leaderboard. The official technical report even explicitly notes that the performance gap between public and private tasks serves as an important indicator of overfitting. In the view of the organizers, handcrafting agents with knowledge specific to the public environments or customizing harnesses for the benchmark’s idiosyncratic interactions constitutes task overfitting in a broad sense.
This has been Chollet’s consistent position. As early as March 2026, he wrote on social media that if human engineers are allowed to look at specific test tasks and craft bespoke programs to solve them, solving all public tasks becomes engineering-wise quite mundane; the benchmark team itself had previously released a harness that achieved a perfect score by replaying human actions. However, the core metric the official leaderboard aims to measure has always been a model’s adaptability when confronted with unfamiliar and unprepared situations; the ability of human engineers to fine-tune external scaffolding for specific task types is not what the benchmark seeks to evaluate. After explaining this rationale, he concluded: If it’s AGI it shouldn’t need a human in the loop handcrafting a system for every new task.
This commentary cuts directly to the vulnerability of engineering hacks on public benchmarks. When scoring 100% on an interactive benchmark, the party truly being tested is actually the group of software engineers sitting behind the screen: it is they who figured out the task rules, meticulously orchestrated memory structures, and iteratively adapted the interaction interface; as for the foundation model encapsulated at the core of the system, it acts more like a computational unit answering the calls of the external scaffolding. This reality hardly needs outsiders to point out; the benchmark creator drew this line unmistakably clear from day one of setting the rules.
If the experimental results from the VISTA team still left the impression that multimodal vision was key to conquering the benchmark, another industry achievement pushed that explanation into a contradictory dilemma. NVIDIA unveiled its general-purpose long-horizon agent architecture, AVO, on its official technical blog. Derived from NVIDIA’s AVO architecture research, this system was originally designed to explore high-dimensional search spaces for low-level chip compute kernel code, achieving impressive speedups in challenging systems engineering optimizations.
The NVIDIA team transferred this system, originally designed for code compilation and engineering search, directly onto the testing grounds of the very same interactive game benchmark. In system design, AVO forms a striking contrast with VISTA. In its technical blog post, NVIDIA stated that the team abandoned building explicit programmatic world models, adopted the direct interaction design principles outlined by VISTA, and independently reimplemented the task interaction interface. However, in its choice of perceptual modality, the model operates entirely within pure text: every environment observation is provided strictly as a 64×64 plain-text character grid, with no images involved and not a single visual token fed to the model throughout the entire run. The two systems share the same philosophy in task interface design, but diverge to polar opposite extremes in the modality fed into the model.
Yet it was precisely this text-only system, devoid of any visual input, that likewise achieved a 100% clean sweep across all 25 public game environments in the same benchmark. What is more, the total number of environment interaction actions AVO accumulated during its runs was roughly 12% lower than that of VISTA, which relied on high-resolution visual rendering and pixel-level inspection tools. Reaching the exact same 100% finish line, one path led toward high-resolution visual rendering and local pixel zooming, while the other reverted to seemingly primitive plain-text matrix formatting. Two systems taking polar opposite routes in perceptual information ended up delivering virtually identical results on the same benchmark.
Faced with this comparison, NVIDIA’s research team included an explicit disclaimer in their blog post, cautioning readers not to treat this comparison as a strictly controlled, single-variable ablation experiment. Significant differences existed between the two systems in their underlying agent backends, text formatting of environment states, long-term memory retrieval mechanisms, and context engineering details. Furthermore, like VISTA, the record reported by AVO occurred strictly within those 25 open-source public tasks without undergoing blind verification on the official private test set—making them essentially self-reported scores by their respective teams. To date, neither of these perfect scores has been independently reproduced by any third party.
Behind this divergence in modalities lies a particularly vivid human connection. Yeyin Zhu, one of the co-authors of NVIDIA’s AVO technical blog post, was also a co-author of the pioneering MIT paper ARC Is a Vision Problem!—a paper studying ARC-1’s static grid tasks that argued they should be approached as visual problems. The very same researcher who previously advocated for the visual route went on, while leading the ARC-AGI-3 effort at NVIDIA, to achieve a more action-efficient 100% score on the same benchmark using an all-text engineering system that completely abandoned image input.
This fact clearly demonstrates that whether a system can achieve a high score in such interactive tasks does not depend on whether the model inspects images or reads characters from a table. What consistently plays the decisive role is the external system’s engineering organization of search paths, its pacing constraints on long-horizon task execution, and its systematic management of trial-and-error processes. As long as the external scaffolding is built sturdily and structured completely enough, different perceptual modalities are more like interchangeable components plugged into the same mechanical exoskeleton—either can execute the workflow smoothly.
In today’s AI news, headlines proclaiming that a model has conquered a difficult benchmark, broken human records, or hit an all-time high on general intelligence evaluations are commonplace. Such narratives often give the impression that the model’s own general thinking abilities have taken another qualitative leap. Yet by dissecting how such external systems reached perfect scores on public tasks, one can cultivate a cooler, more grounded perspective when encountering such claims.
The next time you see similar high-score reports, consider asking yourself three reflective questions that get closer to the technical reality:
The first question concerns the attribution of capability: To what extent does this score actually reflect the foundation model’s own adaptive reasoning, and how much of it stems from the state machines, planning trees, and toolchains handcrafted by external engineers specifically tailored to the task characteristics?
The second question concerns the nature of the evaluation ground: Are the numbers cited in the report self-reported scores repeatedly iterated on open-source, known-rule public debugging tasks, or certified blind evaluations conducted inside the benchmark organizers’ strictly closed testing environment?
The third question concerns hidden costs: Beyond the operational actions officially tracked on the leaderboard, how many rounds of self-correction did the external scaffolding run behind the scenes, how many auxiliary tool calls were made, and how many off-the-record retries was the model permitted in the dark?
The core requirement of general intelligence is that when faced with an unfamiliar scenario it has never encountered before, a model can demonstrate on-the-fly generalized understanding and problem-solving ability under finite resource constraints. When an evaluation task originally designed to test general adaptability has its pathway mapped out crystal clear by engineers through elaborate external workflows, allowing a model to score 100% by following the pipeline, what is truly on display is the power of modern software engineering in system integration and task decomposition. As the benchmark creator has repeatedly stated, this perfect score records what humans did to get the model across the finish line. Whether the model truly understands the world is a question the scorecard leaves unanswered.