Industry & CompetitionTrust & Governance

Who Is Moving the Ruler: The Dual Tensions of AI Leaderboards

When building applications or calling APIs and choosing models, many developers refer to public evaluation leaderboards. These leaderboards condense various coding, math, and Q&A tests into a single composite score and rank models accordingly. I used to make technical selections the same way: check who ranked first, and prioritize picking that model for the architecture. Picking from an off-the-shelf ruler saves time and effort, and people naturally assume the highest-scoring model is the strongest.

However, a recent event broke this convenient habit. OpenAI released its next-generation flagship model, GPT-6 Astra, and a benchmarking organization promptly published its evaluation score. Within a week of the new model’s release, this organization changed its scoring rules twice, turning Astra’s score from lagging behind into a tie for first place, without any change to the model itself. If you regularly rely on leaderboards for technical decision-making, you need to understand how these rankings actually come about. When rankings change, it is not necessarily because the models have changed; it may simply be that the ruler measuring them has moved.

Two Rule Changes in a Week: After Astra’s Release, the Ruler Moved Faster Than the Model

On September 3, 2026, OpenAI officially released GPT-6 Astra. At the end of the launch event, company president Greg Brockman declared that this marked the beginning of the AGI era and proposed discussing whether Astra already constitutes artificial general intelligence (Trending Topics). Expectations were widespread that this flagship model would pull ahead on public benchmarks. On the day of the release, the Intelligence Index, a composite index run by commercial evaluation firm Artificial Analysis, was at version v4.1.1. The published results immediately drew attention: Astra scored only 61 points. This score not only tied the previous-generation GPT-5.6 Sol, but also visibly trailed Anthropic’s Claude Fable 5.1 (66 points) and even fell below Meta’s Muse Spark 1.3 (max) (TechTimes).

The next day, September 4, 2026, Artificial Analysis launched Intelligence Index v4.2. This update introduced AA-Briefcase, focusing on agentic knowledge work, as well as GDP.pdf, a document reasoning benchmark sourced from data provider Surge that spans 4,592 pages of long text; it also removed GPQA Diamond, on which frontier models were universally approaching perfect scores. A more fundamental change lay in the reweighting: the weight of private held-out test sets undisclosed to model vendors was raised from 20% to 40%. Under the new rules, Claude Fable 5.1, which had held a 5-point lead, dropped to 57 points, while Astra shifted to 55 points, narrowing the gap between them to 2 points (AA v4.2 更新公告).

The revisions did not stop there. Three days later, on September 7, 2026, Artificial Analysis released v4.3. In this revision, the terminal coding benchmark Terminal-Bench was upgraded from v2.1 to v4.0; the previous banking customer service test τ³-Banking was replaced by AutomationBench-AA, built in collaboration with workflow automation provider Zapier, featuring 657 tasks across finance, HR, marketing, operations, sales, and technical support, all using private hidden data; and the composite weight of private held-out test sets was increased further to 45%. The scores were updated accordingly: both Astra (max) and Claude Fable 5.1 (max with fallback) landed at 53 points, tied for the global top spot.

Within four days, the scores of the two models went from 66:61 to 57:55, finally settling at 53:53. Over these seven days, neither lab released a new model version. In response to questions, Artificial Analysis officially cited three reasons: bringing forward evaluation elements originally planned for v5, frontier models evolving too fast to wait for the original schedule, and introducing a higher proportion of private test items to curb leaderboard gaming.

Each of these reasons has logical merit, but in terms of practical effect, the ruler shifted noticeably faster than the models under evaluation evolved. Changing the rules twice in a single week allowed a new model that started at a disadvantage to erase the score gap and tie for first place with zero updates to the model itself. For teams relying on leaderboards to select models, fluctuations in composite scores often arrive before any actual change in model capabilities.

The same model saw its scoring rules adjusted twice in one week, narrowing the score gap from 66:61 to 53:53, with all changes occurring on the ruler

Leaderboards Are a Business: Who Measures, Who Pays, and Who Judges

Understanding the violent swings of the ruler requires confronting the commercial drivers behind benchmarking. Public AI leaderboards involve clear cash flows, customer relationships, and growth objectives; they are commercial businesses in their own right. The ecosystem consists of three core players: the model vendors being tested, the independent evaluation firms acting as measurers, and the audience of technical decision-makers and engineers.

This business model imposes institutional constraints on any notion of absolute neutrality. Artificial Analysis, for example, prices its Pro subscription at $417 per seat per month for professional users; for large enterprises, it sells Enterprise plans that include dedicated APIs, full data exports, custom reports, and advisory services (Artificial Analysis 定价页面). In its public-facing presentations, endorsements and citations from Amazon, Google, Meta, and Microsoft are displayed on the homepage as core commercial assets.

Precedents for this dynamic exist in the financial sector. In the aftermath of the 2008 subprime mortgage crisis, scholars and regulators reviewing Moody’s and Standard & Poor’s pointed out an inherent flaw in the issuer-pays model: when revenue depends on the entities being evaluated, objective measurement comes under pressure from the threat of customer churn. A similar tension exists in large model evaluations. Evaluation organizations need endorsements and citations from major tech companies to sell high-priced seats, while tech giants need the authority of leaderboards to showcase their scorecards. If rankings persistently diverge from major vendors’ expectations by relegating new flagships to second-tier status, those vendors will turn to competitors. The degree of adoption by evaluated vendors directly determines the survival space of the benchmarking firm.

The tied score of 53 points between both sides in v4.3 clearly reflects this balance. On the surface, Astra and Fable 5.1 share the top spot, but the detailed sub-scores point to completely different engineering profiles. Fable 5.1 maintains an advantage on the knowledge work benchmark AA-Briefcase and the scientific coding test SciCode; Astra leads in the terminal environment Terminal-Bench v4.0 with 59.1% versus 52.0%, and takes the top score across the field on AutomationBench-AA, the enterprise workflow automation benchmark built with Zapier, at 68.5%.

The difference in financial cost for engineering deployment is even starker. Running the same evaluation suite, Astra costs approximately $3.26 per task on average, consuming a total of around 60 million output tokens; Fable 5.1 costs an average of $7.63 per task, consuming roughly 190 million output tokens across the full evaluation suite, more than double the API expense and token consumption. Commercial composite indices fold these critical dimensions for practical deployment directly into a tied 53 points. A weighted composite score records the outcome of balancing various demands, and a single number cannot reflect genuine engineering characteristics.

Tension One: It Must Match Intuition, but Cannot Just Be Intuition

External commercial interests are only one contributing factor; a more intractable conflict within evaluation products is rooted in the logic of measurement itself. The core value of a leaderboard lies in providing a measurement residual that deviates from the public’s everyday experience; yet the credibility of the leaderboard relies entirely on that very same everyday experience for endorsement.

Developer intuition from everyday use of models is what people call subjective feel. If leaderboard rankings perfectly match everyday intuition, the leaderboard is merely an expensive reiteration of consensus, offering no incremental value. The value of evaluation lies precisely in valid counterintuitive residuals: using systematic testing to reveal hidden shortcomings or extreme capabilities that casual trial runs cannot detect.

The evaluator’s paradox arises here: the public’s only anchor for judging whether a ruler is objective happens to be subjective intuition. Once the residual exceeds an acceptable psychological threshold (such as ranking an acknowledged flagship below an older model), people instinctively conclude that the ruler is distorted. The criterion for distinguishing valid residuals from measurement noise lies in predictive power over time: high-quality residuals are corroborated in engineering practice over subsequent months and settle into new consensus; poor residuals merely expose defects in the measurement itself during deployment.

Public discussions surrounding Artificial Analysis’s index methodology exposed specific sources of measurement noise. In the weighting scheme of v4.1.1, agent capabilities accounted for 34% (GDPval-AA v2 at 20%, τ³-Banking at 14%), coding capabilities accounted for 24% (Terminal-Bench v2.1 at 16%, SciCode at 8%), and the remaining share was distributed among AA-LCR, AA-Omniscience, Humanity’s Last Exam, GPQA Diamond, and CritPt (Artificial Analysis 方法论). The evaluation organization claimed that the overall index had rigorous statistical precision, with a 95% confidence interval of less than ±1% (Artificial Analysis 评测说明).

Zhihu contributor 某科学的仙人仉 analyzed the leaderboard and noted that approximately 76% of its weight suffered from measurement flaws (知乎讨论). GDPval-AA v2, carrying a 20% weight, was based on an original benchmark that relied on blind evaluations by industry experts, whereas AA substituted pairwise LLM-as-judge decisions and Elo ratings, introducing scoring drift; Terminal-Bench v2.1, with a 16% weight, was an outdated version that only had its validators fixed and faulty problems removed in the subsequent 4.0 update; τ³-Banking, carrying a 14% weight, relied on the lightweight model GPT-5.4-mini to act as both user and judge, and according to this contributor, an official report stated that simply fixing script misjudgments caused scores on identical trajectories to rise by about 9 percentage points; AA-Omniscience, carrying a 12% weight, contained an exploit where models could achieve a perfect zero-hallucination score by universally refusing to answer unfamiliar difficult questions, taking advantage of rules that spared refusals from penalties; roughly 54% of the scored sub-problems in SciCode, with an 8% weight, were affected by flaws that depressed scores of correct models; and AA-LCR, with a 6% weight, had distorted reference answers while the automated judge failed to read the full 100k-token source documents.

Zhihu contributor imnoob142536 also listed specific ranking inversions: Claude 4.6 Sonnet scored higher than the flagship Opus of the same generation, and the top-tier version of Meta’s Muse Spark 1.3 ranked above Astra. Furthermore, approximately half of the reference answers in the high-difficulty test Humanity’s Last Exam were disputed, yet it held a 12% weight; roughly 50% of the entire index’s scoring was directly tied to agent tool execution, making overall results vulnerable to amplified volatility from sporadic environment interaction errors.

Regarding questions about scoring points through refusals, Hacker News user jascha_eng expressed a different view in community discussions: based on everyday experience, AA-Omniscience correlated strongly with real-world usefulness, and the scoring rules were designed to reward correct answers while heavily penalizing fabricated hallucinations, thus exempting refusals to answer from penalties (Hacker News 讨论). In practice, this mechanism intended to encourage cautious responses turned into a route for leaderboard gaming, allowing models to avoid losing points through passive refusal.

Beyond scoring rules, the execution harness driving a model also sways results. An execution harness refers to the runtime environment that ingests inputs, orchestrates tools, maintains state, and parses outputs. On the abstract reasoning benchmark ARC-AGI-3, Astra achieved a 62.7% pass rate at a cost of $26,098 under the official neutral harness; under OpenAI’s proprietary adapter harness, because opaque reasoning state was allowed to persist across requests, the identical Astra weights facing the same set of questions surged to a 99.9% score, with costs dropping to $18,817 (Orbilontech; Finance Biggo).

At the same time, non-profit organization Epoch AI, in an evaluation aggregating over 50 benchmarks across 267 models, ranked Astra number one globally with a composite score of 169 points (The Decoder; Epoch AI 方法论). The same model produced 62.7% and 99.9% under different harnesses, and ranked starkly differently across organizations. This gap clearly demonstrates that when a leaderboard diverges from intuition, it may reflect deep-testing insights, or it may merely reflect biases in execution harnesses and validation scripts.

Tension Two: It Must Serve as a Target, but Cannot Be Gamed to Death

Beyond internal complexity, a measurement system must also contend with targeted test-taking by evaluated subjects. This gives rise to the second tension: a leaderboard must command industry influence to attract top vendors to tune and adapt for it, but once vendors pivot their R&D focus toward targeted leaderboard gaming, the entire metric system rapidly loses its measurement utility. This decay conforms to Goodhart’s Law: when a measure becomes a target, it ceases to be a good measure.

In large model R&D, metric failure follows a three-stage progression. It begins with data contamination, where public test questions leak into pretraining corpora, and correct answers merely reflect memorization rather than generalized reasoning; it is followed by targeted test-taking, where teams perform reinforcement learning alignment tailored to leaderboard frameworks, templates, and judge preferences, enabling models to score high on specific question types without gaining general capability; and it culminates in metric saturation, where frontier models approach test set ceilings, clustering tightly near perfect scores, and the benchmark loses discriminatory variance. Artificial Analysis removed the doctoral-level Q&A benchmark GPQA Diamond in v4.2 precisely because the frontier tier universally approached perfect marks, leaving the question set without measurement resolution (AA v4.2 更新公告).

The release of deepseek-v4-flash on July 31, 2026, became a classic case of targeted benchmark climbing. In terms of hardware and architecture, the model entirely retained its original design of 284B total MoE parameters and 13B active parameters, with the team solely rerunning post-training (Kie.ai). Official technical charts claimed that the realigned V4 Flash outperformed the larger-parameter flagship V4-Pro-Preview across nine mainstream agent benchmarks, covering DeepSWE, Terminal Bench, Cybergym, Toolathlon, and DSBench. The comparison provided in the official model card was even more direct: on the same architecture with only post-training rerun, V4 Flash rose from 7.3 points in the preview edition to 54.4 points on DeepSWE, and from 61.8 points to 82.7 points on Terminal Bench 2.1, surpassing the larger V4-Pro-Preview on both counts.

Yet on the model page of independent evaluation organization Artificial Analysis, the composite index for V4 Flash 0731 (max) stood at 52 points, tied with the much smaller Qwen3.8 27B (xhigh) and GPT-5.6 Luna (max) (at release, AA had scored it at 50 points). While this represented roughly a 10-point improvement over the previous-generation Flash at release, it remained distant from true flagships. Furthermore, these gains could not be reproduced externally: an outside technical analysis delving into officially disclosed data found that among the nine highlighted benchmarks, two belonged to unreleased private test suites; none of the execution harnesses used in the testing were open-sourced; and the official report table included only a single column of configuration notes lacking detail (Medium 专栏分析).

These paper gains stood in stark contrast to everyday engineering experience: in routine coding and long-instruction following, the model’s performance showed no substantive gap from the older version. Targeted tuning against benchmark metrics can inflate paper scores in the short term, but under targeted post-training adaptations, fixed static benchmark suites often degrade within months of release into showcase numbers of little reference value.

Public benchmarks reach metric saturation under targeted test-taking by model vendors, with measurement capability decaying even as scores rise

How to Build a Good Ruler

Having seen through the commercial compromises and gaming tactics behind leaderboards, if I were to design an evaluation system for fellow practitioners, the overall logic would have to be derived directly from the specific pitfalls outlined above. The first issue to address is vendors memorizing test suites. The fact that high-difficulty suites like GPQA Diamond had to be outright removed after frontier models universally saturated them shows that any public test questions posted online will eventually be scraped into training corpora. To eliminate such rote memorization, test suites must use private held-out test sets that vendors cannot inspect or prepare for in advance. Artificial Analysis raised the weight of private sets to 40% in v4.2 and further to 45% in v4.3 precisely to plug this loophole. Yet secrecy alone is insufficient; as long as questions are repeatedly called during evaluations, answers will inevitably leak over time. Once a set of questions suddenly shows uniform convergence across different models’ outputs, it should be decisively retired, with fresh original questions continuously added on a fixed schedule.

Hidden test items alone, if only judged by comparing the final line of generated text, still cannot prevent targeted guessing and output-format fine-tuning. Even more problematic is the immense variance introduced by execution harnesses: swapping in OpenAI’s proprietary harness on ARC-AGI-3 propelled scores from 62.7% on the neutral harness to 99.9%, a difference of 37.2 percentage points; DeepSeek similarly posted inflated scores using proprietary harnesses that outside parties could not reproduce. If the scripts and environments used in evaluations are not open, one cannot tell whether higher scores stem from a smarter model or an assist from the harness. Therefore, evaluations must place models in fully open-source, transparent interactive environments to test complete multi-step execution trajectories. The evaluation focus must center on the sequence of tool calls, the accuracy of provided parameters, the ability to self-correct upon execution errors, and whether goals are genuinely achieved within the system. The 657 cross-departmental task flows built by AutomationBench-AA in collaboration with Zapier reflect this thinking: rather than dwelling on whether the model outputs polished phrases, they verify in a sandbox whether it can process real finance reimbursements or customer support tickets all the way to completion.

Even when rich data is gathered, many leaderboards compress performance across all dimensions into a single overall score to cater to public ranking habits and commercial subscriptions. This practice flattens a model’s most crucial engineering traits. Astra and Claude Fable 5.1, which tied for first place at 53 points in Intelligence Index v4.3, possess diametrically opposed strengths in practical applications. Astra excels at terminal operations and process execution, posting a 59.1% to 52.0% lead on Terminal-Bench v4.0 and achieving a field-leading 68.5% on the automation benchmark AutomationBench-AA; more critically, completing the test cost an average of only $3.26 per task, consuming roughly 60 million output tokens across the full suite. Fable 5.1’s advantages lie in long-horizon knowledge work on AA-Briefcase and scientific coding on SciCode, but its average per-task cost reaches $7.63, consuming around 190 million output tokens for the entire evaluation, representing more than double the cost and resource consumption. Engineers designing systems must select the right tool for specific scenarios, which means evaluations must faithfully output multidimensional profiles detailing granular capabilities and cost expenditures, abandoning composite scalars that force every trait into a single number.

Beyond scoring methodology, the timing of rule changes equally determines a ruler’s credibility. Artificial Analysis modified its rules twice in just one week after Astra’s release, turning an initially lagging flagship into a co-leader without any update to the model itself. Ad hoc adjustments keyed to vendor release schedules inevitably invite suspicion regarding commercial motives. A sound ruler must follow a preregistration approach: fix test suite admission criteria, scoring rubrics for each sub-item, and version update schedules in advance, publish them openly, and update strictly according to a fixed cadence. Regardless of how significant a new model release is, even if a new flagship stumbles unexpectedly on existing tests, the evaluation organization should not retroactively inflate weights of selected sub-items, much less swap questions midstream. Completely decoupling rule update cycles from major vendor launch events is essential to cutting off legitimate concerns about biased interests.

To achieve this, the evaluation organization’s funding sources must be sufficiently independent, and its statistical treatments must remain grounded in reality. If a commercial evaluation firm derives its primary revenue from high-priced enterprise subscriptions and displays major vendor endorsements and citations across its homepage, maintaining neutrality under the threat of client churn becomes exceedingly difficult. Non-profit research organization Epoch AI demonstrates an alternative path. Their funding relies primarily on grants and donations rather than charging evaluated vendors; in statistical methodology, they employ item response theory, modeling the difficulty of different benchmarks alongside model performance to normalize more than 50 benchmarks of varying difficulty spanning multiple years onto a single unified capability scale, complete with error bars showing confidence intervals for each score (Epoch AI 方法论). This explicitly reminds readers that even with a two-point gap between two models on a leaderboard, one cannot simply conclude that one has surpassed the other without checking whether the error bars indicate statistical significance.

Even when individual organizations maintain rigorous standards, the industry should not rely on a single source of measurement. Commercial proprietary benchmarks, academic open-source test suites, and internal business evaluations built by development teams must collectively form an ecosystem of checks and balances. The fact that Astra initially received only 61 points on a commercial composite leaderboard while securing the top spot with 169 points in Epoch AI’s academic evaluation illustrates that any single evaluation organization will inevitably have perspective biases and blind spots. Only when multiple evaluation entities with independent funding and segregated test suites coexist can respective strengths and weaknesses surface through cross-comparison, calibrating one another so that no single vendor can dominate the market through PR maneuvering or localized benchmark gaming.

Finally, in addition to scoring models, evaluation organizations must establish a meta-scorecard for themselves. As community analyses have highlighted, leaderboards frequently harbor substantial measurement defects: some test suites have over half of their sub-questions affected by scoring flaws, some benchmarks see scores increase by 9 percentage points merely by fixing validation script misjudgments, and other benchmarks permit unearned perfect scores through universal refusal. A qualified ruler must regularly disclose quality audit reports to the public, transparently reporting the proportion of broken test questions, misjudgment rates of validation scripts, and the drift variance of judge models themselves, reflecting score fluctuations driven by leniency or strictness. If an evaluation firm shies away from disclosing these self-examination metrics and instead dynamically patches composite rankings whenever major labs drop new releases, engineers who treat those leaderboards as absolute truth when committing to production technical stacks will ultimately face recurring crashes under live traffic, paying the steep price of spending months tearing down and re-architecting their systems.

As a Builder, How to Use a Heavily Gamed Ruler

Having recognized the commercial bargaining and gaming tactics behind leaderboards, when making architectural selections today, my first step is still to open the leaderboards, though I use them very differently than before. With dozens of new models emerging every month, nobody has the bandwidth to test each one individually. The greatest utility of public leaderboards is serving as a coarse filtering funnel, helping teams eliminate clear laggards in minutes and narrowing hundreds of candidates down to two or three targets. After this initial filtering, however, the composite score offers no practical guidance for finalizing architecture. That overall number is designed for marketing promotion, artificially fusing divergent model strengths together; real-world engineering depends solely on individual characteristics that solve specific operational bottlenecks. If your application involves automated DevOps scripting, look directly at granular rankings on Terminal-Bench v4.0; if you need to read lengthy contracts or perform compliance audits, focus on error rates on long-context benchmarks like GDP.pdf; if you are building an enterprise cross-tool agent, evaluate workflow completion rates on benchmarks like AutomationBench-AA. Debating a two- or three-point composite score gap divorced from concrete operational scenarios contributes nothing to writing software.

Once candidates are selected from granular sub-items, the next step is examining the evaluation harness and runtime configuration in the official documentation. The evaluation harness represents the complete execution environment and invocation scripts configured for the model during testing. Different harnesses yield vastly different results. In the logical reasoning test ARC-AGI-3, Astra scored 62.7% under the official neutral harness, but jumped to 99.9% under OpenAI’s proprietary harness that retained private state. If your production system cannot provide those customized prompts, multi-round reflection retries, or proprietary state interfaces, the high scores seen on the leaderboard will not materialize in your own codebase. When reviewing these sub-benchmarks, one should place greater value on procedural metrics assessing multi-round dynamic interactions. Having a model repeatedly invoke tools, handle system errors, and adjust parameters step by step inside a sandbox tests comprehensive interactive capability, which vendors cannot easily game through simple fine-tuning. Conversely, single-turn multiple-choice questions, standard Q&A, or short-text completions can often have their scores inflated through a few rounds of targeted data training, offering limited reference value.

When assessing models that make the candidate pool, rather than focusing on an absolute score at a fixed slice in time, one should pay closer attention to the velocity of progress across version updates. In mathematics, this rate of improvement is known as the first-order rate of change; in practical selection, it measures whether a model can continuously climb when facing novel tasks. A model architecture with a brisk release cadence and solid gains with each iteration often carries far greater evolution potential, even if it currently trails by two points on an older benchmark, than a mature model relying on repeated fine-tuning against stale question suites to preserve its numbers.

After choosing candidates, the real line of defense resides within your own engineering architecture. Production systems should avoid binding directly to any single vendor’s proprietary API; instead, you can place a unified adapter layer, commonly known as an abstraction layer, between application code and underlying models. Keep prompt construction, context management, and tool orchestration logic entirely within your own system, preserving only the most generic invocation contract for the underlying models. This decoupling of business logic from specific models reduces the model to a hot-swappable compute component. With a decoupled architecture, online validation becomes straightforward: the team can route 2% of live traffic in the production cluster to the new model as a small-scale canary test. Running under live production traffic for a few hours provides empirical data on conversion rates, response latencies, and billing costs; if metrics fall short of expectations, changing a single line of configuration in the backend immediately reverts to the previous setup.

University rankings correspond to one-off, irreversible, hard-to-verify decisions, whereas model selection corresponds to high-frequency, measurable, replaceable engineering choices for which teams should build their own evaluations

The analogy to university rankings illustrates this cognitive disconnect well. Calling a model API is entirely different from choosing a university for four years. Applying to college involves paying substantial tuition, moving into dorms for years, and bearing immense costs if one drops out or transfers; moreover, prospective students can hardly verify actual teaching quality beforehand, forcing everyone to rely on external rankings for certainty. University administrations understand this mindset thoroughly, optimizing metrics specifically around international student ratios or citation counts.

Model selection in engineering exhibits the exact opposite characteristics. Every API call comes with millisecond-level latency, itemized billing down to fractions of a cent, and explicit error stack traces. Engineers hold complete measurement instrumentation, fully capable of running offline evaluations against historical business data or routing 2% of traffic in a cluster to observe real-world performance. If a new model times out, violates format constraints, or exceeds cost limits, modifying a single configuration file rolls back to the previous version in seconds.

Transferring a university-ranking mindset, forged around irreversible and high-stakes decisions, into software systems that can be measured and replaced at any time overlooks the inherent flexibility of engineering. By building domain-specific test suites in your local environment and treating models as swappable compute commodities, external leaderboards can change their scoring rules every few days without disrupting your online development cadence.