American hospitals and commercial health insurers are locked in a war of attrition using the exact same algorithmic toolkit. Hospitals run charts through algorithms to recast routine inpatient stays as high-complexity cases and bill insurers more; insurers deploy algorithms in reverse across those very same records, hunting for pretexts to deny claims. With the marginal cost of each volley plummeting to near zero, the two sides can trade blows over a single chart through multiple rounds of escalation. The Blue Cross Blue Shield Association tallied the damage: in just two years, cases where the association determined patient care had undergone no substantive change cost its commercial health plans an extra $942 million.
What fuels this escalating contest is cheapness itself. In the past, reviewing a chart and drafting a denial notice consumed expensive human labor, leading both sides to call a truce after two or three rounds; today, with the marginal cost of each exchange rendered negligible, the volume of skirmishes has soared. The cheaper intelligence gets, the more expensive the bill becomes. Healthcare’s algorithmic slugfest is merely the most visible case study; the same dynamic is operating elsewhere.
The plunge in model pricing is no illusion; it is anchored in hard empirical data. Epoch, an organization that tracks model compute costs over time, systematically measured inference pricing across five benchmarks for several leading models. In their September 2026 report, they documented an unambiguous trend: holding performance constant, token query prices for frontier models fall by an average of roughly 47% per quarter. In the immediate wake of technological breakthroughs, single-quarter drops reached as high as 66%; even two years into technology diffusion, the quarterly decline for equivalent capability remained steady at around 32%.
Yet while prices halving nearly every quarter is undeniable, the relentless climb of the overall bill is equally indisputable. The API list price developers see in their developer consoles is, at best, the thinnest surface slice of the system’s true ledger. Peer beneath that list price: consumers enjoy massive hidden subsidies, enterprises contend with explosive query volume, public-sector testing runs up enormous bills, and industries wage algorithmic warfare against one another—each force driving aggregate expenditures higher. Around the outer perimeter of the entire apparatus hangs a risk exposure so volatile that actuarial models cannot even attach a price tag to it.
Public list prices are open for anyone to inspect, creating a seductive illusion of cost control. To trace where the money actually goes, one must begin with the layer closest to the individual developer: subsidized consumer subscriptions.
Many engineering teams treat the $200-a-month individual subscription as their psychological cost anchor, assuming a flat membership fee can cap their total system expenditures. In June 2026, semiconductor and frontier technology research firm SemiAnalysis conducted an empirical test. They purchased top-tier individual subscriptions and ran complex, long-horizon coding tasks in controlled environments until exhausting weekly quotas, then converted the consumed tokens into theoretical invoices based on published API list prices.
The resulting disparity was staggering. According to SemiAnalysis’s estimates (unconfirmed by vendors), at peak utilization, a $200-a-month ChatGPT Pro account yielded compute equivalent to up to $14,000 at API list prices; an identical $200-a-month Claude Max 20x subscription translated into roughly $8,000 of theoretical value. For power users, vendor subsidies on API list prices reached between 40 and 70 times, dwarfing the industry’s prior rule-of-thumb estimate of roughly 10x.
Business model analyst Gennaro Cuofano argues that the deficit vendors absorb on subscriptions is fundamentally a strategic procurement budget for the future. For frontier labs, the scarcest input today is the interaction trajectory of senior engineers guiding agents through complex, real-world tasks—invaluable training corpora for the next generation of models. Meanwhile, flat-rate subscriptions serve as a call option betting on rapid compute deflation, locking in user workflows ahead of time. So long as underlying inference costs continue their downward march, real-world delivery costs will steadily contract, eventually nudging the business toward profitability.
Yet such lavish subsidies, designed to harvest high-quality signal, carry an unmistakable expiration date. When OpenAI unveiled GPT-6 Sol and Luna in late September 2026, official API pricing was roughly half that of the prior generation, with Sol falling to $2 per million input tokens and $10 per million output tokens while delivering broadly comparable benchmark performance. Frontier capabilities from leading labs are increasingly being sequestered into gated safety programs requiring rigorous clearance rather than packaged into cheap consumer tiers. The window for harvesting forty- to seventy-fold subsidies through personal subscriptions is closing. Labs can afford to buy elite data on the consumer side, but once models cross into enterprise production, they enter a commercial settlement structure that tolerates no losses.
In enterprise procurement and private deployments, there is no such thing as an unlimited free lunch. According to industry research estimates from SemiAnalysis, roughly 75% to 85% of Anthropic’s revenue originates from usage-based commercial contracts. Every single API call made by an enterprise customer is an API call paid for.
Why, then, are corporate technology budgets spiraling out of control even as per-million-token list prices fall off a cliff every quarter? Peter Walker, an analyst at model routing platform OpenRouter, shared a telling dataset from the compute supply chain. In the legacy single-turn Q&A paradigm, a prompt consumed only a few hundred to a thousand tokens. But once implementations transition to autonomous agents that plan, invoke tools, and repeatedly self-correct, executing a single workflow can devour a thousand times the tokens of a traditional chat. OpenRouter’s data reveals that between February 2026 and the present, agent-driven token consumption on the platform surged 14-fold, with roughly 70% of that volume consisting of prompt cache hits.
Why hasn’t falling pricing driven down overall spending? The arithmetic is straightforward: even if underlying prices halve every quarter, when enterprises expand query volume a thousandfold to automate workflows, the volume multiplier completely swamps the dividend of cheaper compute. To support complex multi-agent orchestration, the framework layer imposes further overhead for state management and context assembly, effectively doubling consumption once more.
For enterprises facing volatile usage and runaway budgets, this unpredictability has accelerated a migration toward new billing paradigms. According to exclusive reporting from The Information, OpenAI has begun piloting charging based on completed tasks with select enterprise customers, though OpenAI declined to comment. In vertical enterprise software, automation provider Sierra and customer service platform Fin—acquired by Salesforce for $3.6 billion—have already shifted to charging exclusively for tasks resolved autonomously without human intervention. AI software development startup Cognition has even pledged up to $10 million in credit refunds to enterprise clients if its system fails to deliver agreed-upon engineering outcomes.
This pricing evolution signals that buyers and sellers of AI services are fundamentally redefining value. As the era of token-metered billing by character count recedes, enterprises are no longer paying for cheap synthetic text; they are paying for tangible, deterministic delivery.
When enterprises pay for agents, expenditures can at least be calibrated against business outcomes. But when model capabilities intersect with public safety and state regulation, the price tag for safety evaluation quietly explodes at an even more astonishing velocity.
Evaluating whether frontier models introduce catastrophic risks has become an unavoidable mandate for regulatory agencies worldwide. In the United States, the primary entity tasked with national security risk evaluation is the National Security Agency’s Artificial Intelligence Security Center (AISC).
According to reporting by Jeff Stein in The Washington Sun on September 24, 2026, two anonymous sources familiar with classified intelligence estimates revealed that the center spent billions of dollars evaluating and testing frontier AI models in 2026 alone. The largest expenditure went toward procuring cutting-edge compute, followed by compensation for industry technical specialists. Following the revelation, a Pentagon spokesperson declined to comment on security and confidentiality grounds, stating that the Department of Defense does not publicly discuss the technical architecture or resource allocation of such capabilities.
Some U.S. lawmakers calculate that establishing a permanent, nationwide regulatory testing infrastructure for frontier models could cost the federal government tens of billions of dollars annually in testing and verification alone. That dwarfs the official fiscal projections produced by the Congressional Budget Office (CBO) when reviewing the AI Safety and Innovation Act (H.R. 9363), which proposed establishing a dedicated AI risk center. At the time, the CBO estimated average annual operating costs at just roughly $20 million; a separate proposal for a nationwide tracking system carried a five-year budget of only $36 million. The intensive testing expenditures documented in classified estimates exceed policymakers’ early desktop calculations by two to three orders of magnitude.
Civilian safety testing bodies face a vastly more constrained reality. The primary entity testing civilian frontier models in the United States is the Center for AI Standards and Innovation (CAISI, formerly the U.S. AI Safety Institute), and its actual operating budget for fiscal year 2026 stood at only roughly $15 million. Across the Atlantic, by contrast, the UK Artificial Intelligence Safety Institute commands £66 million in steady annual appropriations, priority access to over £1.5 billion in public compute infrastructure, and a roster of more than a hundred full-time technical specialists.
The spiraling cost of frontier model safety testing stems primarily from regulators and commercial labs competing for the same strained compute market. According to The Washington Sun, citing third-party compilations of data from The Information, Anthropic signed roughly $517 billion in compute agreements and letters of intent over the past 11 months, underscoring the sheer expense of procuring frontier compute. Frontier models exhibit pronounced emergent behaviors; stress-testing a highly autonomous system to verify it will not cross catastrophic thresholds in cyber warfare or biosecurity requires evaluators to marshal compute clusters of equivalent scale for adversarial penetration. Each time model capability scales to a new tier, verifying that it will not inflict systemic harm costs almost as much as building a peer-scale computing platform from the ground up.
Against this backdrop of soaring evaluation expenses, a central question emerges: who should foot the bill? Nat Purser, a researcher at the AI Verification and Evaluation Research Institute, has proposed establishing a statutory testing levy requiring model developers to contribute proportionally toward the compute and personnel costs of independent evaluations. The Washington Sun also reported that Anthropic and OpenAI signaled in private discussions that they may be willing to absorb a share of regulatory testing expenses. Leading frontier developers are likewise preparing a self-regulatory body called SAFA—the Standards Agency for Frontier AI—hoping to curb testing expenses by harmonizing evaluation benchmarks.
The government’s massive testing expenditures represent, in essence, a public defense expenditure to prevent systemic catastrophe. But in real-world commercial sectors where financial interests directly collide, arming opposing parties with dirt-cheap intelligence tools does not compress aggregate costs; instead, it drags both sides into a compute war of attrition from which neither can disengage.
Unpacking the $942 million bill introduced at the beginning of this article brings the full anatomy of this confrontation into sharp relief. When discussing why his hospital network procured algorithmic billing software, Dave Mazurkiewicz, chief financial officer of Michigan health system McLaren Health Care, put it bluntly: “All the insurers are using A.I. to scan our charts to look for claims to deny. For the same reason, we’re looking at the same charts today.” Almost simultaneously, Caroline Pearson, executive director of the Peterson Health Technology Institute, offered a near-identical assessment from the vantage point of commercial insurers. Asked by a reporter why hospitals and insurance carriers dispute single contested claims across so many iterative rounds, her answer was just three words: “because it’s cheap”.
Precisely because the cost per negotiation cycle has fallen so precipitously, the confrontation has spiraled into a self-perpetuating loop. On September 24, 2026, the Blue Cross Blue Shield Association (BCBSA) published a data analysis that revealed the first detailed balance sheet of this algorithmic battle. The figures show that between 2023 and 2025, BCBSA commercial health plans paid out an extra $942 million in reimbursements solely because hospitals broadly adopted algorithm-assisted coding software to inflate billing complexity.
Luke Chalker, an executive at the association, explained that this $942 million in excess spending counts only cases where billing codes were escalated by algorithms despite the association concluding that clinical care remained unchanged. Of that $942 million surplus, $653 million stemmed directly from secondary diagnoses extracted and appended from medical chart narratives by algorithms. Each case classified as high-complexity by these tools cost insurers an average of roughly $11,000 more.
As The New York Times reported, the mechanics are systematic: hospitals run algorithms over patient discharge summaries to automatically append secondary diagnosis codes such as anemia or hyponatremia, repackaging routine hospitalizations as complex admissions to collect nearly $12,000 more per case without administering additional medical care. In major bowel surgical procedures, for example, algorithmic intervention caused the proportion of highest-complexity cases to leap from 10.2% to 22.7%, driving an estimated $61 million cost increase in that single surgical category alone. Luke Chalker highlighted a telling discrepancy: while hospitals reported surging numbers of anemia diagnoses, comprehensive prescription and procedure audits revealed no corresponding increase in actual blood transfusions.
A mature monetization flywheel has coalesced around these chart review algorithms. McLaren Health Care’s CFO stated that after deploying chart-auditing software from SmarterDx, the hospital system recaptures roughly $1 million in additional revenue each month from insurers, with the software vendor collecting a percentage cut of the newly recovered sums.
The total bill for this contest is governed by simple arithmetic: the number of skirmishes multiplied by the cost of each round. When reviewing paper charts and drafting denial appeals relied entirely on human labor, every volley was slow and costly, prompting both sides to compromise after two or three rounds. Today, generative tools reduce the cost of drafting bulletproof appeals and identifying denial loopholes to near zero; as the cost per round hits bottom, the volume of exchanges climbs indefinitely. As Shiv Rao, cardiologist and founder of healthcare AI platform Abridge, observed, what is underway is algorithms warring against algorithms and agents dueling agents. As the marginal cost of interaction collapses to zero, the aggregate adversarial bill stacked up across the system swells dramatically.
Technology-induced inflation of this kind is hardly unprecedented. Dr. David Brailer, a former senior U.S. health official, noted that while the widespread adoption of electronic health records (EHRs) improved chart legibility, it also lowered the barrier to upcoding. Today, the influx of generative technology is similarly projected to accelerate healthcare inflation. Algorithms first master the billing arenas where structured data is abundant and collections are immediate; yet in the core therapeutic domains that actually treat illness and restrain costs over long horizons, technological diffusion remains sluggish. Actuaries echo this reality: when PwC surveyed chief actuaries across 27 US health insurance plans, nearly 70% of respondents ranked intelligent documentation and automated coding tools among the top three drivers pushing up healthcare costs for the coming year, with medical cost trends running at a nearly two-decade high. (writing in Health Affairs)
Runaway healthcare spending is not an isolated anomaly; it is simply the domain where precise dollar figures have surfaced first. In cyber warfare, financial fraud detection, and content moderation—any arena structured around bilateral adversarial dynamics—cheap intelligence is dragging both sides into the same quagmire of compute attrition.
Yet expensive as healthcare invoices may be, they at least appear as explicit line items on a financial balance sheet. In more complex commercial arenas, the most dangerous potential liability has no price tag at all.
Around the outer perimeter of the total intelligence ledger hangs a vastly larger, unpriced systemic liability that looms over every enterprise deploying frontier models.
According to an assessment by international reinsurance intermediary Lockton Re cited by CSIS in February 2026: “It is clear that CGL insurers do not currently model, underwrite, or price AI risks, so there is likely a growing gap between what insurers intend to cover and what they actually cover based on the policy language.” In other words, underwriters of standard commercial general liability (CGL) insurance have not yet developed actuarial models or pricing structures tailored to autonomous algorithmic failures and business disruptions. (dedicated industry analysis)
In a subsequent policy analysis, the Center for Strategic and International Studies (CSIS), a Washington think tank, offered an industry forecast: because the potential blast radius of autonomous agents is so expansive and assigning causality for accidents is extraordinarily difficult, mainstream commercial insurers may, during the next annual renewal cycle, exclude liabilities arising from autonomous algorithms either partially or entirely from standard coverage. While this remains an industry projection rather than a settled outcome, its gravity is illustrated by historical comparison: cyber insurance required nearly two decades to establish even a rudimentary underwriting framework, and underlying verification capabilities remain fragile today. If standard commercial policies exclude AI liabilities outright, enterprises stripped of insurance backstops will face unmanageable balance-sheet exposure as they roll out autonomous systems.
Tufts University scholar Josephine Wolff has emphasized that historical loss data is fundamentally inadequate for modeling the catastrophic forward-looking risks of frontier systems. In a June 2026 research report, reinsurance broker Gallagher Re noted that when leading labs like Anthropic release frontier models like Mythos under strictly closed access regimes, external insurers are completely barred from conducting penetrative safety evaluations. Confronted with systems they cannot independently audit, underwriters retreat into defensive postures, choosing to “price uncertainty rather than risk,” as the report observed. Unquantifiable uncertainty inevitably translates into astronomical premiums or outright denials of coverage. Although the industry has begun exploring tying policy pricing to standardized model evaluations, until robust external auditing frameworks mature, many business units deploying autonomous agents are operating functionally uninsured.
From the hidden subsidies embedded in consumer subscriptions, to explosive enterprise volume, to sovereign testing expenditures, internecine algorithmic warfare, and unpriced systemic tail risk, the true ledger of artificial intelligence is vastly more entangled than a few lines of lightweight API calls in application code.
Returning to the architectural perspective of engineering practice: compressing prompt lengths and routing calls to cheaper models can trim immediate costs in specific scenarios. But relying solely on localized technical optimizations to evaluate an organization’s AI investment obscures the systemic whole.
To reckon with the true total cost of intelligent systems, leaders must move beyond spreadsheets that track only per-token API prices. Unpacking each ledger examined above reveals five critical questions that technology executives must audit during project scoping and budget reviews:
Controlling API list prices solves only the most superficial puzzle on the surface. Only by rigorously tracing the flows across all five ledgers—subsidy dynamics, volume compounding, verification overhead, adversarial spirals, and unpriced tail risk—can organizations keep their comprehensive intelligence costs firmly under control as unit token prices continue their downward plunge.