Most evaluations of autonomous AI research agents focus on final benchmark scores or whether the agent can pick the optimal configuration. On September 23, 2026 (2026-09-23), a team from Carnegie Mellon University and Tsinghua University introduced a different approach in their preprint WhatWorkedBench. Rather than just observing which configurations the agent ultimately chooses, the benchmark asks whether the agent truly understands what would happen in the experiments it never ran. In this benchmark, the environment presents dozens of candidate experiment combinations, but the agent is permitted to run only a limited few. After running these tests, the agent must predict the scores for every single unrun combination in the space. Surprisingly, the agents’ experimental intuition fell short of a classic statistical method from sixty years ago. Picking the optimal configuration and understanding what actually works turn out to be two entirely different problems. Where agents lose points is in the last mile of translating data and rules into judgment.
The researchers set up a standardized exam scenario: a typical machine learning pipeline written in NumPy, SciPy, and scikit-learn, with fixed input data and random seeds. Embedded in the code are 4 or 6 binary switches (called factors in the paper), each representing a concrete engineering decision—such as whether to binarize term frequencies or filter stopwords. 4 switches yield 16 configurations, while 6 switches yield 64. Each pipeline run produces a score normalized to [0, 1]. The environment is deterministic and noise-free; running the exact same configuration repeatedly yields identical results.
The exam rules are straightforward: the agent starts with 2 free anchors—one with all switches off, and one with all switches on. It then receives a budget of additional real runs — 4-switch tasks typically get 8. Together with the 2 free anchors, only 10 total observations across the entire combinatorial space have real ground-truth measurements. During decision-making, the agent can inspect the source code and run Python computations (capped at 30 CPU seconds, with 20 calls per task). Once the budget is exhausted, the agent must submit a complete prediction table, estimating scores for every one of the 16 or 64 configurations, including all the configurations it never tested.
What makes this difficult is that the effect of a switch is not constant. The exact same switch can increase scores in one combination and decrease them in another. For instance, in the real-world scientific literature retrieval task SciFact (5183 documents, 64 queries), when high-frequency term compression is disabled, turning on the term-frequency binarization switch boosts the retrieval score by +0.160. But when high-frequency compression is enabled, turning on the very same binarization switch actually reduces the retrieval score by -0.084. A single switch produces two diametrically opposed effects, resulting in an interaction of -0.245 Paper §1. The underlying mechanism in the code is that binary counts neutralize the compression transform, leaving no difference in ranking order between the two branches.
Such switch effect reversals are far from rare. Across the benchmark’s 36 task conditions, 35 exhibit sign reversals depending on background context, with 24 having reversal magnitudes exceeding 5% of the task’s total score spread Paper Appendix L.4. To provide objective ground truth for inference, WhatWorkedBench spans 30 data sources and 8 workflow categories (classification, regression, clustering, time-series forecasting, image inpainting, retrieval, heartbeat detection, and graph link prediction), comprising 1248 configurations in total. The authors exhaustively ran every single configuration in advance; single-threaded replay takes just 128 s (127.87 s) Paper §4.
Evaluation relies on an effect recovery metric, which measures the gap between predicted and true effects relative to the magnitude of the true effects themselves. More accurate inference scores closer to 1; if the mean error exceeds the magnitude of the true effects, the score drops straight to zero. The metric is formulated as , where is the mean absolute error across all 32 / 192 conditional effects (32 for 4 switches, 192 for 6 switches) and is the average magnitude of the ground-truth effects. During scoring, configurations actually tested by the agent are populated directly with authoritative ground truth, isolating the test strictly to the quality of unmeasured inferences. Compared to MLE-bench, which evaluates only end-to-end outputs; AblationBench, which grades whether ablation plans look like those written by human experts; and CAFE, where effects are estimated via statistical models, WhatWorkedBench relies on exhaustive ground truth to provide an objective answer key for evaluating experimental understanding.
In algorithm engineering, we often joke about graduate student descent, poking fun at how dependent hyperparameter tuning is on human researchers. But it is not really a joke. When facing an unfamiliar problem, an experienced human engineer rarely needs to blindly brute-force hundreds of parameter sets. They glance at the pipeline code, run a few ablation trials, skim the terminal logs, and already form a decent mental picture of unrun combinations before deciding what to try next based on experimental intuition.
WhatWorkedBench deconstructs this intuition into three independently scoreable components: selecting experiments, making inferences, and reading code Paper §1:
The fundamental flaw of past end-to-end evaluations was confounded attribution: when an agent ultimately failed, there was no way to tell whether it had picked the wrong experiments or drawn flawed inferences. WhatWorkedBench’s core methodological contribution is a controlled evaluation protocol that swaps estimators on identical observations. Once an agent finishes its exploration, the evaluation framework feeds the exact observations it collected into three standard estimators (main-effects ridge regression, interaction ridge regression, and Gaussian processes) for refitting. Here, each estimator’s own sampling strategy is turned off, and its inputs match what the agent received down to the last digit. This factors out differences in upfront exploration, leaving a pure showdown of inferential capability: does an LLM’s own experimental intuition beat mature, traditional statistical methods?
Picking the best configuration and articulating the effect of every switch are two independent capabilities. Pairwise ridge regression is an old-school statistical method that follows a textbook routine: it selects combinations to test following a fixed pattern, fits the observed results to a mathematical equation, and projects untested combinations. Across 22 questions, it selected the overall optimal configuration 15 times (15/22). But when subjected to a different test requiring its entire prediction table (predicted scores for every configuration) to pass, it succeeded only 3 times (3/22) Paper §6.1. How closely a prediction table matches the ground truth is captured by a recovery score—scaled up to 1, where higher means more accurate. It is like how pointing out the best dish on a menu is easy, but describing the exact flavor profile of every recipe is hard. Picking the right configuration does not mean understanding the system. This separation leads to three surprises.
Surprise 1: The LLM’s experimental intuition loses to a sixty-year-old Gaussian process. Intuitively, a capable LLM should be able to reason effectively once shown experimental data. In reality, the recovery score of the model’s own prediction table on that very data was only 0.632. By contrast, a Gaussian process is a sixty-year-old statistical method designed to extrapolate across an entire space given a handful of observed points. Fed the exact same data, the Gaussian process scored 0.698 Paper §6.2. Even with standard statistical tools readily available, the agent handed in a worse scorecard. The evidence was solid; the judgment collapsed.
Surprise 2: Reading the code didn’t improve inference—and sometimes made it worse. One would naturally expect that letting an agent read the full source code would deepen its understanding. In practice, performance was virtually identical whether the agent saw the code or not; the code-reading group sometimes even scored lower. Yet the code contains free structural equivalence clues: when an upstream switch is turned off, adjusting any downstream parameter under it has zero effect, meaning two seemingly distinct configurations behave identically in production. When this clue was supplied as a rule to statistical methods, the inference recovery score jumped from 0.248 to 0.462 Paper §6.3, without running a single extra experiment. Agents simply failed to translate the code they read into usable knowledge.
Surprise 3: Able to recite the rules, but unable to submit a compliant table. Common intuition suggests that if an agent understands rules in the code, it should abide by them when submitting its answers. The researchers first asked eight agents to restate the constraint rules found in the code; in text, every single one answered correctly. But when examining the prediction tables they actually submitted, every agent violated the exact rules it had just articulated. It is like a student who can recite a theorem verbatim yet fumbles the moment it appears on an exam. In a post-hoc batch fix, the researchers forcibly aligned conflicting cells in the prediction tables to equal values. Without adding a single experiment, the average recovery score climbed from 0.338 to 0.507 Paper Appendix K.3. The rules were recited at the linguistic level, but entirely disregarded in the output data.
I think this study is essential because as AI grows more capable, the tasks we assign it become increasingly difficult.
Computations previously handed to machines were largely traceable. Every step followed explicit rules, and errors could be audited step by step. Today, the tasks assigned to research agents are mostly non-traceable. Agents explore massive combinatorial spaces without fixed trajectories. We cannot audit their step-by-step reasoning like balancing an accounting ledger; we can only judge whether the final proposal is sound. That intuition for making trade-offs under uncertainty is judgment. Because judgment cannot be audited step by step, we need a benchmark that measures it in isolation. That is precisely where the value of this exam lies: turning the elusive quality of judgment into a score with a clear ground-truth key. The gap at the execution layer is narrowing, but the gap at the judgment layer is widening. That is our next bottleneck.
Picking the best configuration is not the same as understanding an experiment. For today’s AI research agents, the critical bottleneck remains the last mile of inference and grounding rules into predictions.
A few practical boundaries should be kept in mind when interpreting these findings: this work is an unreviewed preprint, the evaluation focused on a single model family stratified by date due to API routing changes, and the code has been open-sourced on GitHub (EthanNing/WhatWorkedBench, Apache-2.0 license).
Setting aside those caveats, the study leaves us with two actionable assets.
For agent developers and researchers: There is no need to demand that large language models perform high-dimensional mental math in-context. A much sounder path is equipping agents with dedicated statistical tooling while integrating static code analysis to translate existing structural rules into explicit prediction constraints. Let the LLM steer high-level strategy, while statistical estimators and code constraints anchor micro-level inference. The paper’s own repair experiment proves this path: enforcing the agent’s violated rules back onto its prediction table boosted scores from 0.338 to 0.507 without running a single extra experiment.
For benchmark designers and system evaluators: The protocol of fixing the exact inputs an agent collected and replaying them across different estimators is reusable across domains. When evaluating agent systems, separating exploratory sampling from subsequent judgment isolates the true source of failure.