Carrier is a historic American industrial enterprise with a 110-year track record in HVAC and roughly 50,000 employees. This September, during a panel discussion (video) at an industry infrastructure summit bringing three major enterprises together, Carrier’s Chief Data & AI Officer Arun Nandi shared a striking figure: between January and July 2026, the company’s token bill surged eightfold. The spend primarily went toward serving edge infrastructure, software, and data teams for Carrier’s own products. Nandi stressed that this spike was entirely deliberate—it was the price of exploration: “that is the bill for discovery… for innovation.”
What Carrier buys is the model itself; pre-training foundation models is something none of the three companies do. Slightly larger is U.S. Bank, the fifth-largest bank in the United States, with 70,000 employees. When asked whether the bank would pre-train its own foundation models, Chief AI Officer Prashant Mehrotra was pragmatic: “never say never. It’s not there today.” At JLL (Jones Lang LaSalle), a commercial real estate services firm with 110,000 employees, foundation models are also fully procured from outside.
This is by no means an idiosyncrasy unique to these three companies. In November 2025, Menlo Ventures surveyed 495 enterprise decision-makers across the U.S. The data revealed that while 47% of enterprise AI projects were built in-house in 2024, that figure has now shifted to 76% directly buying off-the-shelf. Across enterprise LLM API usage, Anthropic, OpenAI, and Google command an 88% market share. These leaders treat underlying models as everyday consumable commodities, concentrating their engineering efforts on the layer directly above the models. Looking closely at the playbooks they forged in production, four key lessons stand out as directly applicable.
Traditional enterprise software procurement defaults to an evaluation horizon of three to five years. Today, however, leading AI labs ship a new model every two to four weeks, forcing U.S. Bank to overhaul its procurement mindset. As Prashant Mehrotra explained, the prerequisite for adopting anything today is knowing that “we can have it replaced in a 12 to 18 month window.”
This modularity is not theoretical: several internal business initiatives at U.S. Bank have already swapped out entire model families mid-stream. Mehrotra argues that infrastructure should be as ubiquitously present as roads, water, and electricity; enterprises should focus on building houses on top, with models merely being the latest utility hookup. Carrier’s Arun Nandi highlighted an even starker operational risk: new models can disappear for reasons entirely beyond your purview—sometimes even beyond the control of your country. As Nandi put it succinctly: they “can be pulled for reasons completely out of your control.” Given that external supply can shift at any moment, the ability to swap models becomes a tangible risk hedge. Prompts and evaluation suites should be architected under the working assumption that you will migrate to another provider next quarter.
Once models become consumable commodities, enterprises immediately confront the question of how to allocate internal costs. U.S. Bank’s approach is to charge token costs back to individual employees so they figure out which model makes sense: “everybody is paying for the token internally… so they know which ones to buy.” Chargebacks alone aren’t enough; they simultaneously run internal upskilling to teach employees which model tier fits which task. As Mehrotra cautioned, answering a simple question for a customer or an employee doesn’t require a trillion-token behemoth.
JLL offers the clearest demonstration of this playbook in action. Across an organization of 110,000 employees, technical staff account for only about 3,000. CTO Yao Morin’s strategy was to initially grant employees generous token quotas and then scale them back. While employees initially complained that tightened quotas hindered their work, they soon figured out cost-saving techniques on their own—such as avoiding re-uploading identical data over and over. Morin’s takeaway: “Human ingenuity is amazing when you give them constraints.” Carrier’s Nandi holds a similar view on costs: expenses shouldn’t be measured purely by cost per token, but rather by the business value delivered per task. That requires attributing costs on a per-task basis right from day one, rather than letting the accounting blur into overhead.
When selecting models, the easiest reference point is public benchmark leaderboards. At Carrier, however, Arun Nandi observed over the past eight to twelve months that employees in certain roles developed distinct loyalties to particular model families—preferences that didn’t align neatly with public benchmarks and had nothing to do with token pricing: “slightly different from what we’ve seen from evals… devoid of all of the cost per token.” Benchmarks cannot substitute for how users actually experience a model in real-world workflows.
Carrier’s strategy is to route workloads by role: multi-step reasoning goes to frontier models, while routine queries route to cheaper workhorses or open weights. U.S. Bank bifurcates latency and cost tolerances on the very same voice channel depending on whether the user is a customer or an employee. JLL built a unified gateway where Yao Morin treated response latency as an independent variable: background tasks that don’t need real-time answers can be deferred to off-peak compute windows, shaving costs even further. Morin pointed out that many frontier vendors offer no enterprise SLA guarantees—asking them for five nines of availability will get you laughed out of the room. At the routing layer carrying actual production traffic, routing decisions are anchored precisely in these accumulated user intuitions and operational realities.
As enterprises pour capital into AI, they inevitably weather plenty of failure. Arun Nandi noted that his mindset when managing these innovation bets resembles running an investment fund: two or three high-performing projects in the portfolio outweigh 95 unsuccessful ones, with a few winners offsetting all the losers.
A January 2026 update from Gartner indicated that through the end of 2025, at least 50% of GenAI projects had been abandoned after proof-of-concept, failing to reach production. A survey of 1,006 European and American enterprise professionals by 451 Research (part of S&P Global) revealed that the proportion of organizations explicitly abandoning the majority of their AI initiatives jumped from 17% to 42%. Faced with such high failure rates, Carrier skips lengthy project approval dossiers in favor of reviewing prototype code directly for rapid triage. Mid-sized and smaller teams can adopt the same playbook: set upfront termination conditions, and ruthlessly cut projects when metrics fall short.
The provenance of these insights also warrants clarity. The panel discussion took place on September 16, 2026, at the AI Infra Summit, hosted by third-party organizer Kisaco Research. Compute provider Lambda served as the platinum sponsor of the conference, and the panel was moderated by its Chief Commercial Officer. The publicly available recording is a 41-minute edited cut published on Lambda’s channel on September 25; the unedited full session remains unavailable.
It is also worth noting that Lambda draws 80% of its revenue from hyperscalers and frontier labs (as he noted in an interview). Its official blog post frames in-house builds as hardware hell, depicts an API-only route as vendor lock-in, and conveniently points to its own reserved compute as the ultimate answer. Even the buzzwords have a backstory: Mehrotra uttered “valuemaxxing” on stage (identifying where value lives before deciding where dollars go), after which Lambda’s official post claimed the term was coined at their panel. Yet an IBM tech column reveals that three months prior to the summit, the term had already been documented in black and white as coined by Marc Boroditsky, Chief Revenue Officer of Nebius—Lambda’s direct competitor.
Commercial motivations aside, looking strictly at what the three executives are implementing inside their respective organizations, their engineering choices hold up. Most of what was shared on stage is corroborated by public sources (outside of self-reported figures like the eightfold token surge). For instance, Mehrotra explicitly stated in another interview that the bank’s scale has firmly tipped toward renting models. Carrier CTO Markus Klausner confirmed in an interview with Fortune that the company’s model strategy embraces an all-of-the-above stance, pairing frontier commercial APIs with fine-tuned open-source weights. JLL’s model and public cloud architecture are detailed in an official press release and a Microsoft customer story. What can be adopted with confidence are the four engineering takeaways outlined above. What deserves a discount is merely the narrative framing and vendor pitch packaged around them by Lambda.