Over the past two months, two diametrically opposed views have been circulating across the industry. Some observed the rapid revenue growth among compute and model providers and argued that enterprise AI spending would continue to balloon uncontrollably. Their theoretical weapon is the Jevons paradox: in 1865, English economist William Stanley Jevons observed that as steam engines became more efficient at burning coal, England’s coal consumption did not decrease—it increased, because coal became cheaper, prompting factories and use cases that previously could not afford coal to adopt it at scale. Applied to AI, the logic goes: the cheaper models become, the more new use cases can be spun up, and overall spending actually increases. Others argue that due to macroeconomic headwinds and challenging project ROI, AI demand has hit a bottleneck, PoCs are failing to reach production, and enterprise AI spending is declining.
Which of these two arguments is correct? As it happens, there is a dataset available to test them. Ramp is a financial operations platform for US businesses—companies use it to issue corporate cards and pay invoices, leaving an audit trail for every expense. Its Chief Economist, Ara Kharazian, publishes a monthly AI Index tracking enterprise spending on AI; see the official Ramp methodology and data sources explanation. The three figures that matter most to us—how many tokens models consumed, how much money was spent, and the effective price per token—come from a cohort of businesses connected to Token Spend Management. Ramp integrates directly into companies’ model provider backends, tracking strictly API usage and billing—excluding subscriptions and out-of-payroll employee expense reimbursements (see the author’s clarification on methodology)—to pull real invocation volumes and billing data, rather than relying on self-reported surveys.
Looking across the entire timeline, enterprise AI demand in 2026 has experienced neither explosive growth nor a recessionary collapse. Instead, it is undergoing a transition: total token volume is rising, effective unit prices are falling, and usage patterns are shifting toward cheaper, standardized tiers. This trend of stabilizing or slightly declining bills is driven collectively by engineering teams tightening cost controls, enterprises instituting default model tiers, and the adoption of routing strategies.
Let’s first rule out one possibility using the initial data point: enterprises haven’t stopped using AI. Ramp tracks the share of companies that paid a model provider in a given month, counting both subscriptions and API tokens. The data shows this share continues to climb: in August, more than half of US businesses paid for AI, up roughly 11 percentage points from a year earlier (underlying data table). Anthropic saw the sharpest rise: a year ago, only one in six companies was paying Anthropic (July report); now, it is nearly half. The fundamental demand base is expanding, which invalidates the “AI ebb tide” narrative right out of the gate (Ramp August report).
Why is adoption still climbing while talk of an AI retreat persists on the outside? Invocation volume is rising, unit prices are sinking, and usage is shifting down to lower tiers—two opposing forces neutralizing the bills. To understand this transmission mechanism, look first at pricing.
The most straightforward reason bills and invocation volumes are diverging is that token unit prices have been steadily falling. Note that unit price here does not refer to official list prices on websites, but rather the effective average price enterprises actually pay on their invoices: factoring in input, output, and caching discounts, weighted by real usage.
According to figures from Ramp’s September report: the effective unit price per million tokens peaked at $1.15 in March of this year, followed by a continuous six-month decline to just $0.68 in early September, a drop of 40%. Recalculating from the chart’s underlying daily CSV series, the decline was still accelerating in August, falling faster and faster.
The impact of price cuts hit the two providers with different depth. Across the same cohort of enterprises and the same timeframe, OpenAI’s actual transaction price dropped nearly 60% from its peak, whereas Anthropic’s fell by only 40%. One is trading price for market share; the other is defending its premium.
This discrepancy in price declines is not solely driven by the frequency of each provider’s official price cuts; it also reflects differences in the composition of calls made by their respective customer bases. Blended unit price is an endogenous metric: even if the list price for individual models remains unchanged, as long as engineers shift the bulk of their workloads from flagship models to lightweight tiers, the weighted average price will naturally drop. The two providers have different task distributions across their user bases, so even if list price adjustments were identical, the slopes reflected on customer invoices would differ. This serves as a compositional explanation; the current dataset does not break down a quantitative attribution breakdown.
Even though unit prices dropped 40%, overall enterprise invocation volumes did not shrink. Reporting on October 2, 2026, based on Ramp data, The Decoder noted that relative to the July peak in spending, total token consumption by enterprises was actually up roughly 50% by late September, setting an all-time record in the final week of September.
A dose of cold water first: don’t use these two figures to calculate “price elasticity of demand.” The 50% increase and 40% decline come from different time windows, and the drop in unit price was largely driven by enterprises shifting tiers, not because the market offered cheaper prices and everyone rushed in.
Where the money actually went is clearest in Ramp’s weekly tier series. The most expensive flagship tier followed an inverted-V trajectory: climbing through July, peaking in early August, and falling for four consecutive weeks thereafter. Ramp’s September report notes that it accounted for roughly 45% of total usage, noticeably below its early-August peak.
The ground ceded by the flagship tier was entirely absorbed by the middle tier: over the same nine weeks, the standard tier’s share increased almost every week, climbing from negligible single digits to over a third. Meanwhile, lightweight tiers and miscellaneous models steadily shrank. Usage did not migrate to open source or self-hosted infrastructure; it simply stepped down one rung within the hierarchy of commercial models.
This migration in usage is, in reality, active intervention by engineering teams. Ramp’s report quotes feedback from multiple enterprise customers: they have standardized default model tiers company-wide to reduce reliance on frontier models; they find standard-tier models sufficiently powerful while delivering significantly higher cost efficiency. During initial feature exploration or PoC prototyping, engineers instinctively reach for the most expensive flagship model to guarantee performance; once the feature rolls out to production, doubling bills compel architects to revise routing. For pipeline tasks such as text classification, content summarization, format validation, and intranet retrieval, models in the GPT-5.6 Terra or Claude Sonnet class can achieve high accuracy without letting flagship models spin their wheels on routine work.
This is not unique to Ramp’s books. Independent benchmarking firm Artificial Analysis evaluated new and previous model generations across a fixed set of tasks and found the same result: running the exact same workloads, newer models cost only half as much as the previous generation while actually generating more tokens (September 22 report). While per-task bills appear only half as expensive on the surface, the actual generation volume was more than 50% larger—amounting to “spending less to do more.”
Another frequently misread trend: this tier shift did not flow toward open source or self-hosting. Ramp data shows open source accounts for less than 5% of enterprise AI spending, a point Kharazian specifically emphasized on X on October 1. Workloads have stepped down within the commercial model hierarchy, without leaving that hierarchy.
Prices dropped, enterprises proactively downgraded tiers, yet invocation volume kept growing. In August, bills began to gradually decline.
Ramp tracks the top 1% of enterprises ranked by “AI spend per employee” each month—the cohort that spends most aggressively (CSV series). The intuitive summary: their monthly AI spend per employee was around $3,000 in January and surged to roughly $8,000 in July, more than doubling in six months; August marked the first reversal, dropping to $7,200—a decline of about 10%. Even after that drop, it remains nearly triple the level of the same period last year. This dip was a pullback inside a steep six-month climb.
“Everyone went on vacation in August” was Ramp’s initial explanation. But its own three-year dataset punctures that theory: during the same summer holiday period, top-tier enterprise bills grew in August last year; only this year did they fall. The more plausible drivers are the two forces discussed in the preceding sections: falling unit prices and enterprises shifting tiers.
Looking across 36 months of history, single-month pullbacks are nothing new; the longest streak of four consecutive monthly declines was still followed by fresh all-time highs. An honest reading of August’s drop, therefore, is that while a topping out and retreat in bills may mark a new trend, there is currently only one month of evidence; the September report will deliver the verdict.
The latest weekly trajectory in September (disclosed by Kharazian on September 25): the August downturn had already reversed course and ticked back up by mid-September, though it has yet to reclaim its July peak. Bills are trending downward with volatility, not diving in a straight line.
Enterprises are spending more thriftily, yet news reports show model providers posting record-breaking revenues. Where is the money coming from? Take OpenAI: Axios reported in late September, citing people familiar with the matter, that its annualized revenue was nearing $70 billion (subsequently confirmed by Reuters), with enterprise revenue doubling within a single quarter. Providers’ growth does not rely on existing customers spending more; it is driven by a massive influx of new customers and new use cases. Ramp observes existing customer bills flattening, while providers capture revenue from an expanding incremental market—the two realities do not contradict each other.
Anthropic’s figures are even more divergent: self-reported annualized revenue from fundraising press releases (self-reported $47 billion in May), annualized numbers relayed by media outlets, and historical revenue reported in IPO prospectuses differ by an order of magnitude, none of which can be used to construct a reliable income statement. The overall picture, however, is clear: in the most recent week, Anthropic captured roughly half of every dollar enterprises spent on models (The Decoder citing Ramp data).
With this persistent tier shifting, what is truly being squeezed is the providers’ profit model. According to historical estimates from Epoch AI and Exponential View, it took four months of gross profits from the previous flagship generation (the GPT-5 family) just to break even on the R&D costs burned prior to release. Flagships carry hefty R&D amortizations; cheaper tiers do not. Now that customers are shifting en masse to cheaper tiers, top-line revenue may keep growing, but the quality of profit has fundamentally shifted. That said, this depends on conditions: if tier shifting halts, or if new customer inflows are sufficiently large, this pressure ceases to hold. Tech analyst Azeem Azhar directly asked Ramp whether this sample represents average enterprises, and received no data in response; when reading this index, remember that it reflects early adopters, not the industry-wide average.
For technical decision-makers, this curve holds direct engineering implications. The trajectory mapped out in Ramp’s charts is simply the aggregate of model selection decisions made daily by countless engineers. Teams are decomposing pipelines that previously routed everything indiscriminately to flagship models—handing data cleaning, format normalization, and intent routing to standard tiers, while reserving frontier models strictly for critical judgment calls. Every such tweak casts a vote on this curve.
Evaluation criteria for engineering selection are also shifting. Fixating on leaderboard rankings or list prices per million tokens offers limited value in real-world production. A far more meaningful metric is the total cost per successful delivery. If an expensive frontier model grasps complex intent on the first turn without requiring correction, its composite cost may well be lower than that of a standard-tier model requiring multiple retries and validations; conversely, deploying a flagship model on fault-tolerant tasks like text extraction or semantic summarization is sheer waste. The essence of cost discipline lies in aligning task complexity with model tier, rather than blindly choosing the cheapest model.
On August 19, 2026, Stripe announced the acquisition of model routing platform OpenRouter for approximately $7.5 billion. The New York Times reported the price citing people familiar with the matter (NYT report), and Bloomberg had also tracked details of the transaction (Bloomberg report). The official announcement press release on Stripe’s website quoted Patrick Collison: “Tokens are becoming the core currency of building products with AI.”
This acquisition gives Stripe visibility into the revenue side of businesses, knowing how much each customer charges end users; OpenRouter holds the gateway on the inference side, logging how many tokens each application routes to which model provider. When these two systems integrate, an infrastructure player can, for the first time at industry-wide scale, see both the top-line inflows and unit-economic outflows of software products simultaneously.
Usage is growing, unit prices are falling, and enterprise bills have begun a gradual decline. One final question remains: will we first uncover the true profitability of enterprise AI through the unified ledger pieced together by Stripe and OpenRouter, or will we see this tier shift recalibrate the compute boom in model providers’ next round of audited financial statements?