Industry & CompetitionChina Tech Ecosystem

The Third Path for Domain Models: A $40M Answer from a 175-Year-Old Company

On August 24, 2026, Thomson Reuters released its in-house large model, Thomson 1.0. The 175-year-old professional information publisher did not train a general foundation model from scratch to compete head-on with OpenAI. Their core assets are concentrated in proprietary corpora such as the Westlaw case law database, Practical Law, and the Checkpoint tax database. As one layer in a multi-model stack, Thomson 1.0 was integrated into the existing legal product CoCounsel Legal, outperforming several mainstream closed-source models in self-reported domain evaluations.

The ledger behind this investment reveals a clear structure. Thomson Reuters reported a cumulative two-year investment of $40 million, covering all expenses for talent acquisition and research team M&A. The base model chosen was Alibaba’s open-weights model, Qwen3.6-35B-A3B. According to Thomson Reuters, the team injected proprietary corpora for continual pre-training, utilizing less than 10% of their total content assets; hundreds of legal and tax experts subsequently participated in fine-tuning, preference alignment, and evaluation. The entire training consumed 35,207 B200 GPU hours, which translates to a pure compute cost between roughly $100,000 and $200,000.

Self-reported evaluation data shows that Thomson 1.0 achieved an average score of 75.2 in the legal domain, surpassing the 72.7 of the base Qwen3.6-35B-A3B, as well as Gemma 4-31B’s 70.5 and Anthropic Haiku 4.5’s 67.7. Thomson Reuters released the weights of the smaller version for academic and non-commercial use, while keeping the full version fully proprietary in-house. Titled “Continual Learning of Frontier Models for SovereignAI,” the technical report emphasizes that what the team built is a continual learning pipeline, with training the model being merely the pipeline’s first step.

Taken together, these developments bring an industry-wide question straight to the forefront. Can a legacy company holding massive industry-proprietary data genuinely establish its own model path by combining an open base with proprietary engineering? Is this an isolated, hard-to-replicate case, or does it represent a general solution for vertical domains? To understand the logic of this path, one must look back to an attempt three years ago that moved in the exact opposite direction.

Three Years Ago, the Same Type of Company, the Opposite Result

In March 2023, financial information giant Bloomberg published the BloombergGPT paper. At the time, the team used a 50B-parameter scale, feeding 363B tokens of proprietary financial data from the Bloomberg Terminal to train a base model from scratch. At that stage, building an in-house base model was the prevailing approach among vertical giants: proprietary data could not be fed directly into external commercial APIs, open-source base models were still weak, and pre-training from scratch appeared to be the only viable route.

This enormously expensive project eventually met an unexpected outcome. Wayne Barlow, Global Head of Terminal Product at Bloomberg, gave a candid response in an interview with Market & Alt Data Insight:

“BloombergGPT was a research model and is actually not used at all in any of our products. Our products are built on a combination of models. We have some commercial models, some open-weight LLMs, some smaller, custom language models that we build ourselves.”

Bloomberg’s current AI product, ASKB, entered beta testing in early 2026, reaching roughly one-third of its 375,000 terminal users. Its underlying architecture has evolved into a hybrid routing solution supported collaboratively by general commercial APIs, open-weight models, and in-house lightweight small models. The 50B large model that once carried the entire company’s expectations remained in the research phase and never entered actual production lines.

Two industry giants with deep data assets reached opposing business outcomes across a three-year span. Bloomberg committed a substantial budget and consumed 363B tokens of proprietary data starting from scratch, ultimately failing to land the model in any real product. Thomson Reuters, according to its own accounts, invested $40 million, used less than 10% of its proprietary content for post-injection on an open base model, and surpassed top general models in self-reported domain benchmarks.

The same type of data-rich company: in 2023, training a 50B model from scratch resulted in zero product usage; in 2026, post-training on an open base model for $40 million beat frontier models in-domain

I originally thought this divergence stemmed from differences in data cleaning difficulty or team execution. Only after clarifying the division of labor between RAG and model weights did I realize the root cause lay in architectural choice. In 2023, the industry broadly viewed general models plus RAG versus proprietary model training as mutually exclusive paths. When faced with this choice, Bloomberg bet entirely on the weight layer, attempting to force massive financial facts directly into a 50B model. In 2023, that approach had its rationale: general models lacked sufficient financial understanding at the time, and RAG technology was still in its early exploratory stages.

However, staking one’s moat entirely on the weight layer quickly invites the impact of open models. Once the capability of open base models catches up, the first-mover advantage of merely baking domain data into weights quickly evaporates. Without a private data flywheel, an evaluation loop, and industry distribution, such models naturally cannot sustain a foothold in production. The approach validated by Thomson Reuters in 2026 offers a new answer: open base models handle general reasoning, RAG carries dynamic facts, and a private data flywheel specializes in evaluation and behavioral alignment. The same problem found an entirely different solution three years later, with the weight layer relegated to a commoditized, easily swappable component.

Why Now

From 2023 to 2026, three core dimensions across the technology stack underwent noticeable shifts. Movements in base model capabilities, post-training toolchains, and capital thresholds collectively reshaped the industry’s baseline.

The first shift is that open base model capabilities crossed the threshold of practical usability. In 2023, the market lacked high-quality open-weight foundation models, leaving vertical companies with no choice but to pre-train from scratch to build custom intelligence. By 2026, open models such as Qwen 3.8 Flash Next and GLM 5.3 Flash were tracking frontier performance closely in domain understanding. According to evaluations by Artificial Analysis, GLM-5.2 scored 51, while the 2.8T-parameter Kimi K3 reached 57, trailing top closed-source foundation models by only 3 points.

At the same time, the costs of compute rental and API calls plummeted. In March 2023, around GPT-4’s initial release, the cost per million input tokens was approximately $30; by late 2025, API pricing for models of comparable capability dropped to around $0.30, an overall reduction of nearly 100x. With the marginal cost of renting general capabilities falling drastically, spending huge sums to pre-train a general foundation model from scratch lost all financial rationality.

The second shift is the maturation and standardization of the post-training toolchain. Full-parameter asynchronous reinforcement learning, preference alignment algorithms based on enterprise value rubrics, and methods for converting domain expert evaluation systems into continuous optimization signals were largely experimental patchworks in 2023. By 2026, this methodology had matured into highly complete infrastructure, allowing vertical teams to rapidly layer high-density professional judgment on top of open base models.

The third shift stems from a fundamental transformation in the economics of frontier general foundation models. Pre-training general large models has evolved into a heavy-industry arms race that ordinary product companies cannot participate in. According to confidential financial data obtained by The Wall Street Journal in April 2026 and disclosed in a SaaS industry analysis report, the combined 2026 spending on model training and operations by OpenAI and Anthropic alone approached $65 billion.

OpenAI closed a $122 billion funding round in April 2026 at a post-money valuation of $852 billion, generating $2 billion in monthly revenue. Even so, its compute budget for 2028 alone is projected to reach $121 billion with an $85 billion loss, pushing the break-even horizon past 2030. While training GPT-4 in 2024 cost roughly $78 million to $100 million, by 2026 the budget for a single pre-training run of a frontier foundation model ballooned to $10 billion to $50 billion. Currently, not a single frontier lab can sustain a healthy standalone P&L.

The key players in this frontier arms race are also shifting toward trillion-dollar conglomerates. In February 2026, SpaceX acquired xAI (the creators of Grok) in an all-stock transaction, valuing the combined entity at $1.25 trillion; SpaceX subsequently completed its IPO on Nasdaq in June, raising approximately $75 billion at an offering price of $135 per share, with its first-day market cap surpassing $2 trillion. NVIDIA simultaneously stands as the largest shareholder across the funding rounds of OpenAI, Anthropic, and xAI. Terms in OpenAI’s latest funding round even stipulate that the disbursement of $35 billion is contingent upon the company achieving Artificial General Intelligence or completing an IPO.

In contrast, targeted post-training in vertical domains requires a total budget of only tens of millions of dollars, with pure compute consumption compressed to hundreds of thousands of dollars. The capital scales and risk profiles of the two approaches are no longer in the same order of magnitude.

Harvey’s Three Stages

If Thomson Reuters offers a static corporate blueprint, the legal AI unicorn Harvey presents a highly representative dynamic evolutionary trajectory. This company, focused on specialized attorney workflows, underwent two pivotal architectural overhauls over the past three years.

The first phase spanned 2023 to 2024. In its early days, Harvey was deeply tied to OpenAI, fine-tuning closed-source models on US case law. Third-party case studies showed that practicing attorneys participating in tests preferred the customized model in 97% of tested scenarios. As subsequent general frontier base models upgraded, general model capabilities quickly caught up with fine-tuning gains; Harvey decisively deprecated its proprietary fine-tuned model and shifted to a multi-model routing strategy centered on general APIs.

The second-phase breakthrough came in June 2026. Harvey turned to a partnership with Applied Compute, implementing full-parameter, fully asynchronous reinforcement learning on the open-weight GLM 5.3 Flash on the AC2 compute platform. In a retrospective on its official blog, Harvey noted that the core reason for choosing GLM 5.3 Flash was its exceptionally solid initial baseline among candidate open models.

The experiment yielded self-reported breakthrough evaluation results: in professional legal benchmarks, baseline pass rates increased from 0.853 to 0.913, and all-criteria pass rates rose from 0.059 to 0.126. Its professional score set a new industry record, directly outperforming OpenAI’s GPT-5.5 xhigh and Anthropic’s flagship Opus 4.8 Max. This marked the industry’s first time using vertical post-training to beat contemporary top closed-source models on professional benchmarks. In its official technical retrospective, Harvey revealed the core laws of post-training across two statements. The first pointed out:

“post-training, harness optimization, and grader design cannot be treated as separate problems”

They followed with a further summary:

“the model only learns as well as the environment allows”

These two statements pull the optimization focus squarely back onto the evaluation environment itself. The quality of the grader and test suite directly sets the ceiling on the model’s eventual capability. If the evaluation environment cannot provide highly discriminative scoring signals, no amount of compute investment will widen the gap.

Broadening the perspective across 12 representative enterprises in legal, financial, and healthcare sectors, industry exploration displays a clearly stratified landscape. Two to three companies have clearly proven this path, including Thomson Reuters, Harvey’s second-generation approach, and healthcare AI company Abridge, whose valuation soared 12x within 12 months. Three companies are in a hybrid transitional state: Bloomberg (shifting toward small models collaborating with commercial APIs), C3 AI’s Narwhal, and Salesforce’s xLAM framework. Two companies stick to purely renting external APIs: financial intelligence platform AlphaSense and Thomson Reuters’ longtime rival LexisNexis. Meanwhile, three companies encountered setbacks or contracted their strategies: BloombergGPT (which abandoned production deployment entirely), Harvey’s first-generation fine-tuned model, and the retired Databricks DBRX.

Analyzing the gains and losses across these 12 enterprises clarifies a key watershed. Struggling companies often treated model training as a technical identity and an ultimate goal, either sinking into the quagmire of pre-training from scratch or relying on shallow fine-tuning over public corpora. Successful teams, by contrast, treated in-house models as an engineering mechanism to optimize product cost, latency, and operational control, consistently embedding them within composite systems backed by general frontier models.

What Is Truly Scarce

Out of the $40 million total budget reported by Thomson Reuters, pure compute expenses accounted for only $100,000 to $200,000. This stark contrast clearly delineates the real center of gravity in vertical post-training: foundational compute has become a highly standardized public commodity readily acquirable with capital. What truly consumed tens of millions of dollars and cannot be bought off the shelf are 175 years of proprietary content assets (spanning case law, practical guidance, and tax data), the meticulous annotation and deep involvement of hundreds of veteran domain experts, and a highly discriminative professional evaluation system.

When examining the RAG framework and post-training within the same system, the boundary dividing their roles becomes crystal clear. General models with RAG and proprietary model post-training are actually complementary planes within the same hybrid architecture, each responsible for different layers. The facts layer is dynamic, citable, deletable, and roll-backable; it is best kept outside the model, carried by RAG and external knowledge bases. The behavior and professional judgment layer is stable, high-frequency, and scoreable; this is what needs to be baked into weights through mid- and post-training.

This division of labor points directly to where scarce resources actually accumulate. The weight layer itself has been commoditized by open-source base models; simply stuffing facts into weights is both expensive and difficult to maintain. The truly scarce assets are, on the one hand, dynamic facts and knowledge indexes that the weight layer cannot reliably support, and on the other hand, the private data flywheels and evaluation systems capable of continuously scoring behaviors. What dictates the ceiling of the system is precisely the ability to manage facts externally while refining behaviors internally.

This shift is likewise corroborated by the evolution at the application layer. In 2023, mainstream teams typically prepared hundreds of Q&A pairs to fine-tune a small model for a specific task. Today, ordinary developers have pivoted wholesale to in-context learning and agentic workflows, rendering fine-tuning no longer the default choice in day-to-day development. Conversely, among platform enterprises holding deep data assets, continual mid- and post-training on open base models is generating massive momentum, propelling Thomson Reuters and Harvey to outperform general closed-source models in their respective self-reported professional benchmarks.

Transmission chain of the post-training stack: open base models can be purchased, while private data and evaluation systems cannot be bought, ultimately compounding into cost efficiency and operational control for in-domain products

While recognizing the potential of this path, negative signals disclosed in the technical report also warrant careful consideration. Data reported by Thomson Reuters shows that following intensive continual injection in the legal domain, the model’s general capabilities suffered measurable regressions. On the mathematical reasoning benchmark AIME, its score dropped from the base model’s 93.3 to 90.0; on the command-line environment benchmark Terminal-Bench, it fell from 45.2 to 40.5; and its coding test score slipped from 39.8 to 37.4. The assumption that domain injection can completely avoid catastrophic forgetting does not hold up in real-world testing.

This capability trade-off precisely explains why Thomson Reuters must maintain a multi-model co-existence architecture in its CoCounsel product. The specialized in-house model tackles high-value professional logic, while general frontier models provide fallback assurance for general reasoning and open-ended tasks. Specialized post-training comes with a clear cost to general performance, making multi-layered architectural hedging an essential engineering safeguard.

The Narrowing Prerequisite

The ability of vertical companies to forge a third path rests on a critical assumption: that the open-source community will continue to deliver top-tier base models of high quality without added commercial restrictions. In 2026, however, this prerequisite is undergoing a substantial narrowing.

In July 2026, Moonshot AI released Kimi K3 under a Modified MIT license, explicitly stipulating that commercial products with more than 100 million monthly active users or over $20 million in annual revenue must negotiate a separate licensing agreement. Shortly after, as documented by industry media in reports on Alibaba shutting down the free tier for Qwen Code, MiniMax adjusted its licensing model to include a community commercial license with a $20 million annual revenue ceiling. At the same time, Alibaba began rolling out revenue-sharing mechanisms for large-scale commercial users of the Qwen series.

The commercial strategies of Chinese open-source vendors are collectively pivoting. The era of offering permissive open-source terms to cultivate developer ecosystems is gradually drawing to a close, replaced by a mainstream push toward monetization through controlled open weights. The window for purely unrestricted, permissive licensing is accelerating toward closure. DeepSeek released its V4-Flash model weights under a standard MIT license in late July, representing one of the few remaining pure open-source channels, while contradictory industry reports persist regarding the eventual distribution model for its flagship V4-Pro weights.

Looking back at Thomson Reuters choosing Qwen and Harvey betting on GLM, third-party teams must position open-source license review as a mandatory prerequisite before initiating engineering projects. If top-tier open models in the future universally carry commercial thresholds or revenue-sharing requirements, the financial model for the entire vertical post-training path will have to be recalculated.

Three Paths, One Decision Rule

Synthesizing the technical and commercial realities of 2026, enterprises face three clear evolutionary paths in model strategy. Companies of different scales and data reserves need to align with completely different combinations of factors.

The first path is developing frontier foundation models from scratch. This path demands massive capital scale between $10 billion and $50 billion and the capacity to bear losses unconstrained by a standalone P&L. Its barriers to entry include hundreds of billions in external funding commitments, multi-year high-intensity compute procurement contracts, and gigawatt-scale power supply—restricting players to frontier labs deeply tied to capital titans. The second path is purely renting general commercial APIs. Under this logic, models are relegated to plug-and-play components at runtime, with product value residing entirely in application orchestration, interaction design, and workflow integration; this is the rational choice for the vast majority of standard software companies. The third path is investing tens of millions of dollars to build proprietary mid- and post-training capabilities on top of high-quality open base models, serving exclusively industry incumbents equipped with high-value data assets.

Hierarchical positioning of the three paths: training frontier from scratch operates at the $10B–$50B scale without a standalone P&L, the third path operates at the $40M scale serving data-rich domain incumbents, and pure renting retains value in the orchestration layer

Faced with these three paths, the evaluation framework for enterprises is straightforward. For vertical enterprises sitting on massive proprietary data, possessing industry-leading evaluation capabilities, and operating in high-value, low-tolerance scenarios, the third path provides powerful strategic defensive leverage. Before embarking, however, teams must evaluate themselves against three rigid prerequisites: first, whether they have an automatically iterable continual learning pipeline ensuring rapid retraining whenever upstream base models update; second, whether their internal evaluation standards are rigorous enough to serve directly as optimization signals for reinforcement learning; third, whether they have established a hybrid multi-model architecture where general frontier models provide a safety net for complex, long-tail tasks. For conventional software teams lacking accumulated proprietary data, blindly pursuing post-training will only repeat the missteps of early fine-tuning.

The pace of future evolution hinges on two key bellwethers. One is the rate at which open-source vendors tighten base model weight licensing, which will dictate whether the foundation of the third path remains stable over the long haul. The other is the penetration speed of general frontier models into specialized vertical domains; once the reasoning capabilities of general base models make another generational leap in legal or financial domains, the relative advantage window afforded by vertical post-training will face renewed compression.