In July 2025, AI achieved gold-medal-level performance at the International Mathematical Olympiad (IMO). Many viewed this as a major breakthrough in mathematical reasoning capabilities, but in fact, as early as three years ago, someone had already written down a definitive date for this milestone.
Between June and October 2022, a forecasting research organization called FRI (Forecasting Research Institute) conducted a months-long forecasting experiment. They brought together two groups—one composed of top machine learning experts, the other of superforecasters with long track records of consistent win rates in probabilistic forecasting—to make quantitative year and probability judgments on several technical milestones. One of the questions was: In what year will an artificial intelligence system win a gold medal at the International Mathematical Olympiad?
The median year predicted by the expert group landed on 2030; the superforecasters were a bit more conservative, predicting 2035. In reality, the gold medal was secured in July 2025. These forecasts were recorded in black and white during the experiment back then, and can now be verified line by line.
Beyond guessing the year, the organizers also evaluated implied probabilities—that is, how much probability forecasters assigned in advance to the outcome that actually materialized, effectively the average score assigned to reality. For the actual achievement of a gold medal, the implied probability given by experts beforehand was 8.6%; for superforecasters, it was even lower at just 2.3%. Across the four benchmarks covered in that experiment, the average implied probability experts assigned to the actual outcome was 24.6%, compared to 9.7% for superforecasters. These raw data points are publicly available in full in FRI’s analysis report.
This is where the narrative that experts underestimate AI came from, a claim heard everywhere over the past year. Confronted with this comparison, many people’s immediate reaction is that experts saw technical progress moving too slowly, concluding that the entire expert community is excessively conservative in its judgment.
However, if you look further, that turns out to be only part of the story.
Starting from the International Mathematical Olympiad question, similar biases cropped up again and again elsewhere.
On FrontierMath, another frontier benchmark testing mathematical reasoning, Wave 1 of FRI’s LEAP expert forecasting project (forecasted 2025-06~08) asked participants to forecast the solve rate for FrontierMath (Tiers 1–3) by the end of 2025. Experts gave a median of 31%, and superforecasters gave 30%. By 2026-01-01, according to the initial resolution records (v1 criteria) in the Epoch FrontierMath report, the actual solve rate of frontier models reached 40.7%. On 2026-06-12, Epoch corrected errors present in 42% of the questions; after the questions were fixed, benchmark scores shifted upward across the board, widening the underestimation gap even further.
Commercial revenue followed the exact same lagging trajectory. In an economic forecasting study conducted by FRI (forecasted 2025-07~08), the question asked what the annualized recurring revenue (ARR) of a single leading AI company would be by the end of 2026. AI experts gave a median prediction of $20B, economists gave $16B, and superforecasters gave $25B. Yet by 2026-09-22, according to reporting by Axios, the annualized revenue of the front-running individual company had already reached roughly $100B.
In Wave 11 of the LEAP project (surveyed 2026-07), the organizers asked for the combined annualized revenue of OpenAI and Anthropic by the end of 2026. Experts gave a median of $70B, superforecasters gave $90B, and the public group gave $41B. Yet according to the LEAP Wave 11 report, barely two months after the survey closed, the reported combined annualized revenue of the two companies had already reached $105B ($40B for OpenAI and $65B for Anthropic). The median forecasts set by all three groups for the end of 2026 were completely blown through within two months.
If we look over a longer timescale, we find that expert timeline forecasts for core technical milestones have also been steadily shifting forward, at a remarkably consistent pace.
The Expert Survey on Progress in AI (ESPAI), led by AI Impacts, systematically reviewed how people’s judgments on the arrival timeline of High-Level Machine Intelligence (HLMI) have evolved in its latest large-sample survey (ESPAI 2024, surveyed in 2024-12, report published in 2026-09, n=1,580). From the perspective of mean aggregation, the year with a 50% probability of achieving High-Level Machine Intelligence was forecast as 2061 in 2016, 2059 in 2022, shifted to 2047 in 2023, and moved forward to 2042 in 2024. Combining the mean-aggregated results across all four rounds of surveys, one can observe that for every year that passes in reality, the remaining time to the experts’ expected arrival date shrinks by roughly 3.4 years (see the full ESPAI 2024 report).
On specific individual tasks, this lag was already routine long before the boom in large language models.
For example, if we look at the task timeline listed in the ESPAI 2023 survey, the median year experts gave for a 50% chance of AI defeating human professional players in the real-time strategy game StarCraft II was 2027. But in reality, DeepMind’s AlphaStar had already achieved this in 2019—a full 8 years earlier.
And in the World Series of Poker (WSOP No-Limit Texas Hold’em), the median year predicted by experts was 2026, while the Pluribus algorithm beat top human professionals in 2019—likewise 7 years earlier.
Across different years, different tasks, and different institutions, these comparisons all point to the same conclusion: underestimation is not an isolated case, but a phenomenon that recurs repeatedly across diverse independent quantitative evaluations.
If we stopped here, it would be easy to draw a very natural conclusion: human experts simply have limited imagination, and algorithmic progress always outpaces everyone’s expectations.
Yet within the very same time window, people from the very same institutions made forecasts on another set of questions that ran in the exact opposite direction.
This reversal first surfaced in the real-world production environments of software engineering.
In July 2025, METR, a non-profit AI safety evaluation organization, conducted a randomized controlled trial on software developers, documented in an arXiv paper.
The basic setup of the trial was that before running it, participants and observers gave quantitative expectations. The software developers participating in the trial predicted that with AI coding tools, they could reduce their task completion time by 24%. The median expected time reduction among invited economists was 39%, and among machine learning scholars, the median prediction was 38%.
From frontline engineers writing code to scholars conducting macroeconomic research, everyone anticipated before the experiment that these tools would substantially cut down time-consuming and tedious steps.
Yet the empirical findings were the exact opposite: when engineers were equipped with AI tools, task completion time did not decrease; it increased by 19%. As the paper stated, granting access to AI tools actually made completion times longer.
The speed gains calculated by experts in an office setting were completely eaten away in real-world engineering environments by the friction of debugging, boundary alignment, patching subtle defects, and deciphering context.
A similar bout of overoptimism appeared shortly thereafter in biosecurity.
In the Active Site biorisk randomized controlled trial supported by FRI (forecasted in mid-2025), organizers tested whether large language models could help undergraduate students complete hazardous biological tasks.
During the forecasting phase, biosecurity experts believed that model assistance would increase experimental task success rates by 22.5%, professional virologists predicted a 40% increase, and superforecasters gave a median estimate of 16.2%.
Yet after running the controlled trial, the actual data recorded on FRI’s analysis page showed that model assistance improved the success rate by merely 5.2%. The hazard uplift that experts intuited was several times higher than actual hands-on capability in a real environment.
Even looking solely at code benchmarks, overestimation was impossible to shake off.
In their annual review of technical progress in 2025, Epoch and AI Digest noted that for the software engineering benchmark SWE-bench Verified, forecasters between 2024-11 and 2025-01 believed the median solve rate by the end of 2025 would be 88%. Yet when empirical results were announced at year-end, frontier models achieved a solve rate of 80.9%, far below previous collective expectations (see Epoch’s summary report).
On one side stood mathematical benchmarks and commercial revenues repeatedly shattering expectations; on the other stood engineering productivity gains, biological wet-lab experiments, and coding benchmarks visibly falling short of forecasts.
Why did the very same forecasters stumble in both directions during the exact same period?
Observing how experts assess AI progress is like watching two entirely different tracks.
One track has a stopwatch: math competitions, programming tests, company business revenue. Once achieved, the time freezes right there for everyone to see. Another track is fitted with instruments that measure real-world friction—such as deploying models into actual development environments to accelerate programmers, where coordination drag and toolchain integration come into play. Then there is a third track, where instruments have not even been built yet—such as robots doing household chores or the tangible impact of AI on macroeconomic aggregates—where success or failure lacks any objective, quantitative yardstick.
On the track with the stopwatch, expert assessments almost always arrived too late: competition gold medals were won ahead of schedule, revenues multiplied exponentially, and reality consistently outran the forecasts.
Yet on the track measuring real-world friction, expert judgments were almost universally too optimistic: workflows once assumed to yield dramatic speedups turned out in practice to be slower due to debugging and context switching.
The direction of expert forecasting error does not depend on what models are good at; it depends on whether there is an instrument in front of our eyes, and what that instrument is measuring. A stopwatch measures how fast pure capability runs—on these questions, experts appear pessimistic. An instrument measures how much real-world friction resists deployment—on these tests, experts appear overly optimistic.
As for the track where no instruments exist, the situation is even more delicate. Whether experts predict a breakthrough next quarter or in the distant future, the results cannot be graded on the ground. This is not because people are being lazy; the issue is that without clear acceptance criteria, it is impossible to define whether the experts were right or wrong. In domains where no yardstick has yet been built, the only honest answer is that nobody knows.
FRI, which hosts these large-scale forecasting initiatives, acknowledged this asymmetry in its interim report. Under their accounting rules, the moment an actual value surpasses the median forecast, underestimation is immediately established and booked on the spot; however, proving that a forecast was set too high requires waiting until the resolution deadline has passed. Consequently, looking only at questions resolved ahead of schedule will naturally surface far more instances of underestimation, as detailed on FRI’s analysis page.
Even more troublesome, our yardsticks sometimes change as well. In FRI’s forecasting questions on labor-hour task assistance, the official historical baseline used by forecasters was revised from 2% to 3.35%, while the actual record for 2025 ultimately landed at 5.7%. Measured by the old ruler, forecasters severely underestimated; but once the revision history of the yardstick is taken into account, the entire win-loss picture completely reverses.
Ultimately, biases in evaluating AI progress consistently track the measurement instruments: where stopwatches are present, experts are perpetually late; where friction gauges are installed, experts prove overly optimistic; and where instruments have yet to be built, no one knows the true answer.
Having weathered four consecutive years of reality checks and data resolutions, where do these frontier experts and forecasters stand today?
The answer is: they have indeed updated their expectations, but only halfway.
Wherever leaderboards can be tabulated, benchmark test suites exist, and progress is compute-driven in the software domain, everyone is moving timelines aggressively forward.
According to Wave 8 of the LEAP project (surveyed 2026-04~05, see the LEAP Wave 8 report), regarding the year Artificial General Intelligence (AGI—explicitly defined in the survey as commercial AI systems outperforming the 90th percentile of full-time human workers on the vast majority of non-physical tasks) will be achieved, the experts’ median estimate has moved forward to 2050, and superforecasters have moved it forward even earlier to 2047.
When asked about the probability of achieving AGI before the year 2100, the median assessment across the entire respondent pool reached 80%.
On the Technological Revolution Scale (TRS), which gauges the depth of transformation, 35% of experts and 34% of superforecasters classified AI as Level 8, equivalent to a century-defining technology on the scale of the Industrial Revolution; 24% of experts and 23% of superforecasters viewed AI as Level 9, a once-in-a-millennium technological transformation.
Comparing with Wave 1 and tracking the attitudes of the same cohort over time, the average score among experts rose by 0.20, while the average score among superforecasters rose by 0.39. Faced with technical breakthroughs that consistently materialized, superforecasters adjusted their judgments by a margin that even exceeded that of the experts themselves.
Behind these numbers, the ecosystem of forecasting tools is also shifting.
On 2026-07-16, FRI released evaluation results for ForecastBench, showing that frontier large language models on a comprehensive suite of forecasting questions had approached the level of top human superforecasters.
In its official explanation, FRI specifically noted that current data shows parity with superforecasters rather than a significant lead, and that the human superforecaster benchmark in the evaluation remained frozen at July 2024 (see FRI’s evaluation announcement).
This creates an interesting situation: the very scoreboard that once measured our human forecasting performance is now being rapidly flattened by technology.
Yet the moment one shifts attention from leaderboards to the real world, it becomes apparent that the other half of everyone’s cognition remains firmly pinned to historical empirical baselines.
When it comes to macroeconomic output growth, total factor productivity gains, and real-world deployment into complex physical systems, long-term forecasts by most experts and economists have barely budged. While they acknowledge that models on screens can rapidly pass many exams, when thinking about how to apply these capabilities to the real economy, they still tend to cling to their prior assumptions of slow diffusion.
They have not turned radical across the board; they have simply reserved radicalism for where benchmark scores exist, while reserving conservatism for an uninstrumented reality.
In the previous section, we understood why misjudgments have a direction. In engineering decision-making, the rule is straightforward: when facing any judgment, first ask which track it sits on.
On tracks equipped with a stopwatch, where metrics have explicit scoreboards, do not place blind faith in expert intuition medians; look directly at trend extrapolations. Expert year estimates often carry psychological defense mechanisms of the herd, whereas trendlines extrapolated from empirical measurement data offer far more reliable readings.
METR has long observed the ability of models to handle long tasks autonomously, finding that the task duration models can sustain at a reliable success rate follows a stable exponential growth pattern: effective duration doubles roughly every 7 months, with 212 days for a 50% success rate and 213 days for an 80% success rate. Extrapolating along this pattern, systems capable of independently handling one month of a human engineer’s working hours (167 hours) have an arrival window between late 2028 and early 2031, as detailed in METR’s official blog and arXiv paper. The dates of 2047 or 2050 that experts guessed by intuition simply do not belong to the same level of informational depth as this extrapolation from continuous readings.
On tracks measuring real-world friction, keep a close eye on calibration signals from the first batch of instruments rather than listening to sideline debates. Debating whether AI can save programmers time or will proliferate biosecurity risks is far less useful than waiting for instrument readings. The METR developer randomized controlled trial and the Active Site biosecurity trial are both prime examples of instruments providing friction readings as soon as they were set up, using control groups to measure productivity discounts and the gulf with expert expectations.
First-generation instruments also have engineering noise. FrontierMath corrected errors present in 42% of its questions on 2026-06-12, after which scores shifted upward across the board (see the FrontierMath evaluation report). But a noisy reading is better than no reading at all. Run the instrument once against real-world friction, and the numbers are right there.
On tracks where instruments have not even been built, treat both optimism and pessimism as weak priors. If a domain has no clear win-loss line, no reproducible experiments, and no feedback mechanism, whether someone asserts it will take fifty years or claims it will change everything next quarter, neither should serve as a basis for decision-making. In the fog without instruments, many confident pronouncements are nothing more than uncalibrated sentiment.
Look back again at the 2022 IMO gold medal forecast. Four years ago, the superforecasters sitting in that room did not make a single arithmetic mistake in their probability calculations, nor could one find logical flaws in their reasoning process. Their only mistake was placing far too much weight on humanity’s millennia-old priors about mathematics.
Yet reality moves forward, and it never accommodates anyone’s priors.
Engineers, typing out code, are erecting one new scoreboard after another. As new measurement instruments pierce into reality, the needles begin to move. What comes next will be guided by the numbers on the dashboard, no longer by the silhouette of authority.
| Forecast Item | Forecaster & Date | Forecast Value | Latest Progress & Comparison | Nature of Error & Notes | Source Link |
|---|---|---|---|---|---|
| Training run compute / cost scale | FRI XPT (2022) | Largest training run by end of 2030: Experts $180M, Superforecasters $100M | Grok 4 official median $490M (Epoch assessment 2025-09-12) | ~5 years early, underestimated by 2.6–5x | Epoch Data Insights |
| Human-level speech transcription | ESPAI 2023 survey | Median year for 50% milestone probability: 2026–27 | Microsoft achieved 5.1% WER human parity on Switchboard in 2017-08 | Task definition involved noise and accent nuances; category milestone, actual achievement preceded forecast | Microsoft Research Report and ESPAI Paper |
| Directional check of date forecasts | Rethink Priorities | Analyzed 41 AI date questions | Of 7 resolved questions, 6 resolved earlier than community prediction, 5 resolved before 25th percentile of forecast | Date-type forecasts systematically resolved earlier than predicted | Rethink Priorities Analysis |
| IMO gold medal achievement | Metaculus Community | Community median in 2022-07 was 2029 (per historical screenshot accounts) | Gold-medal-level performance achieved 2025-07; formal resolution of Q6728 remains contested | Forecast was late, but closer to reality than XPT expert panel | Metaculus Q6728 Page |
| Extinction risk from high-level machine intelligence | ESPAI 2024 expert survey | Extinction risk probability median 10%, mean 18% | Subjective belief assessment, pooled across 3 question framings | Expert risk perception slowly trending upward, median reaching double digits for the first time | Full ESPAI 2024 Report |
| Rideshare autonomous vehicle penetration | FRI LEAP survey (2025-06~08) | AV share of rideshare by end of 2027: Experts 7.3%, Superforecasters 2% | Projected at 2.5% based on current trend extrapolation (unresolved) | Model extrapolation projection, not resolved fact; experts skewed optimistic on physical deployment | FRI Analysis Report |
| Framing effects on extinction risk questions | FRI 2023 XPT experiment | Percentage framing median: 5% | 1-in-X framing median: 1 in 15,000,000 | Massive disparity on same risk topic across framings, illustrating how subjective probabilities are shaped by question framing | EA Forum Discussion Archive |
| AGI arrival year | Metaculus Community | Current community median 2031-03 | 80% confidence interval 2028-05~2037-09 | Community prediction timeline substantially contracted compared to pre-2022 | Metaculus Q5121 Page |
| Cybersecurity benchmark Cybench | Epoch × AI Digest 2025 retrospective | Median predicted solve rate by end of 2025: 62% | Actual reached 82% | Forecast was 20 percentage points too low | Epoch Retrospective Report |
| Combined revenue of leading labs | Epoch × AI Digest 2025 retrospective | Median forecast for three labs’ combined revenue by end of 2025: $16B | Actual reached $30.4B (almost double) | Commercial revenue scale exceeded forecaster estimates by nearly double | Epoch Retrospective Report |