AI AgentModel Architecture

Distilling Frontier Research Capabilities into a Small Model in Just Two Days

Distilling a Capability with Hard-to-Define Boundaries

The most common approach to distilling knowledge from large language models targets tasks with well-defined inputs and outputs—such as writing a few lines of compliant code, extracting entities, or translating a passage of text. Because these tasks inherently have clear criteria for right or wrong, we can provide unambiguous instructions and ask the model to respond in a specific format. However, for time-consuming and complex research tasks, single-step alignment methods no longer work.

The difficulty with research capability lies in its fuzzy boundaries. Faced with an ambiguous question, humans need to organize their thoughts, search diverse sources, read long reports, compare details, backtrack when encountering contradictory evidence, change search keywords, and ultimately synthesize various pieces of evidence into a compelling conclusion. In engineering terms, acquiring thousands of hours of high-quality human demonstration recordings is infeasible. Without these off-the-shelf recordings of human exploration, traditional imitation learning lacks its most critical data foundation right from the start.

The problem then boils down to a single question: distillation requires demonstration data, but where does that data come from? Real human recordings are out of the question. Letting an LLM connect directly to the live web while solving problems and recording its actions is the next most natural path, but it comes with its own hidden costs, which we will tally in the next section. As for synthetic data, while it sounds like it offers the most freedom, in practice it hits a much harder wall: research is not a math problem—you can neither generate questions programmatically nor judge correctness mechanically. There isn’t even a metric you can write into code to evaluate whether a data sample is good or bad. Two and a half of the three paths are dead ends. The OpenResearcher paper, presented at the EMNLP 2026 main conference by a joint team from Texas A&M University, the University of Waterloo, and other institutions, sets out to answer this very question.

How It Learns: Creating an Exam Arena for a Capable Model

If you actually let an LLM connect directly to the internet to solve difficult problems and record its trajectories, the engineering budget fails right out of the gate. To assemble hundreds of thousands of usable training samples from a few thousand seed questions, you would need to make millions of web searches back and forth. If you use commercial search APIs directly, search costs alone would run into thousands of dollars; switch to a more expensive API and the bill multiplies. Even worse, public web pages change constantly: redesigns today, timeouts and takedowns tomorrow. Search results recorded yesterday no longer match today, making the entire training data pipeline impossible to reproduce reliably.

The team’s solution was to bring the entire information environment completely local. They selected thousands of seed questions from public benchmarks, fetched tens of thousands of core evidence documents once based on questions and reference answers, and mixed them into tens of millions of public web pages, constructing a high-concurrency local retrieval system that requires no internet connection. This closed exam arena not only completely erased hefty search bills, but also ensured that results returned by each retrieval are fully deterministic and reproducible at any time.

With the exam arena ready, it was time to bring in the teacher model. The team brought in the open-weight model GPT-OSS-120B, equipped it with three basic actions—search, open web page, and find in page—and let it navigate and solve problems on its own within the closed arena. With 64 H100 GPUs running continuously for about two days, the teacher model conducted nearly 100,000 exploration runs, fully recording how to follow clues, how to read long documents, and how to reformulate keywords and search again when clues run cold.

alt: The offline closed-book repository and teacher rollouts replace expensive web search with a local pipeline, ultimately distilling into a small model with ~3B active parameters

Armed with these comprehensive problem-solving recordings, the subsequent student-training phase became much more straightforward. For the student model, the team chose an open-source lightweight architecture from NVIDIA, with only around 3B parameters actively engaged per inference step. The team avoided complex online trial-and-error; training here was remarkably straightforward: a few GPUs running overnight, with the lightweight base imitating the teacher’s problem-solving steps from start to finish, acquiring this systematic suite of research actions in just a few hours.

Why the Simplest Method Wins: Environment and Actions, Not Answers

First, let’s examine the experimental results. The benchmark used here is the open-source BrowseComp-Plus. BrowseComp-Plus places 830 deep question-answering problems and approximately 100,000 human-verified documents locally for offline search without internet access, testing long-horizon search and multi-hop reasoning capabilities. Before any specialized training, a lightweight base model scored only in the 20s on this benchmark. But after observing the teacher model’s exploration recordings, this small model—with very few active parameters per step—achieved a score of 54.8. Meanwhile, commercial flagships like GPT-4.1 and Claude-4-Opus scored only around 36 under the exact same controlled retrieval environment.

The small model’s high score is not the result of rote memorization of problem answers. Ablation experiments reveal an interesting detail: even when fine-tuned deliberately on failure paths where the teacher model ultimately answered incorrectly, the small model still achieved a score of 55.06, virtually identical to the results from training on correct paths. What the small model learned was the behavioral strategy: when a search comes up empty, try different keywords, and locate evidence within long documents; the specific answers to individual questions were not memorized.

Conversely, this also explains why the smartest general LLMs perform poorly on long-horizon tests. Benchmark data shows that without any search, directly placing the gold document containing the answer into the context yields an accuracy of 93.5% for GPT-4.1. Flagship models do not lack reasoning ability; they score low on these tests because they have not developed the behavioral habit of continuously using tools and iteratively verifying evidence in unfamiliar environments.

However, pure behavior cloning has clear limitations in its applicable scenarios. The paper’s ablation experiments show that once the original gold documents are removed from the document corpus of this exam arena, the model’s score plunges to 6.35. In the AgentFlow paper published by the same author, when facing the real dynamic internet, merely imitating trajectories via supervised fine-tuning drops performance to 19.5, leaving online reinforcement learning with continuous trial-and-error exploration as the only viable path forward. These two controlled experiments demonstrate that pure behavior cloning is truly effortless only when the information environment and tool interfaces are completely fixed.

The Three Prerequisites for Distillation and Where the Real Costs Lie

Tracing through this engineering practice leads to a bold takeaway: this is fundamentally a low-cost distillation of advanced behavioral capabilities from frontier LLMs. Deep research capability has long been the proprietary asset of closed-source frontier models, leaving ordinary developers to pay per token via APIs. OpenResearcher demonstrates that with the right methodology, top-tier capabilities can be distilled from large models at low cost. To make this distillation pipeline work, three hard prerequisites must be met: first, the ability to self-host an open-source LLM as the teacher to sample freely and locally at any time; second, the research behavior to be learned can be decomposed into a sequence of tool-call actions; third, the capability to construct a closed exam arena and implement automated evaluation.

These three prerequisites completely flip the traditional economics of model development. The conventional wisdom has always been that model training is the primary bottleneck consuming compute and burning money. Yet in this pipeline, supervised fine-tuning of the small model becomes remarkably cheap—deliverable after just a few GPUs running overnight. What is truly valuable and forms the real moat across the entire system is the environmental engineering effort spent upfront curating literature and designing tool primitives, along with the massive pool of high-quality behavioral trajectories accumulated within that environment.

alt: Compute weight lies in teacher rollouts and environmental engineering; student distillation only takes a few GPUs running overnight

Subsequent moves by major industry players confirm that trajectories are indeed the hard currency. NVIDIA adopted these synthetic trajectories as part of the fine-tuning data for its new flagship Nemotron 3 Ultra, and its official synthetic data platform, NeMo Data Designer, has also integrated this generation methodology. This move makes things clear: weights themselves become obsolete, but the pipeline capable of continuously distilling behavior from top-tier models, along with the accumulated trajectories, constitutes the truly reusable foundational asset.

Boundaries and What They Mean for Us

While the efficiency of this distillation playbook is truly impressive, a realistic assessment requires recognizing its objective technical boundaries. The ceiling of the small model remains capped by the teacher model, with no substantive leap occurring in the student. On the official leaderboard of the authoritative benchmark, the new-generation general flagship GPT-5 has already broken past 70 points, showing that the generational gap in general intelligence remains firmly in place.

Furthermore, the model’s high score is heavily dependent on the specific exam environment. The current closed-book repository primarily handles factual information retrieval. Once this lightweight model steps outside its familiar controlled arena and is evaluated directly on the live open web, its score immediately drops to 26.3, trailing noticeably behind specialized models honed by peers using large-scale reinforcement learning. Moreover, as of now, no third party in the public tech community has independently reproduced these results, and the open-source data layer remains restricted by non-commercial licensing terms.

Putting these boundary conditions together, deciding whether to build such a pipeline comes down to four strict criteria: the business scenario must involve long-term, high-frequency reuse; the task must have an objective, clearly defined automated scoring mechanism; typical failure bottlenecks must stem from specific decision-making actions such as searching and browsing, rather than infrastructure issues like retrieval recall misses or tool failures; and the information environment relied upon by the business can be packaged and reset at relatively low cost. Only when all four criteria are met does investing compute to construct a local exam arena and distill LLM research experience into a private small model offer a favorable ROI. When facing open-ended tasks filled with unknown variables and demanding extensive impromptu reasoning, relying honestly on commercial LLM APIs remains the path with the highest certainty.