Industry & CompetitionGovernance & Compliance

Refusing to Sign, Holding the Ruler, and Locking the Window: The Origins of Anthropic's Open-Source Stance

On September 29, 2026, an Anthropic frontier red team specializing in offensive and defensive evaluations released a report assessing the cyber offense and defense capabilities of GLM-5.3, an open-source model developed by Chinese AI firm Zhipu AI. Within twenty-four hours, different communities reacted in starkly divergent ways: on the overseas developer hub Hacker News, engineers accused Anthropic of paternalistic fearmongering designed to upsell its own safety offerings; across Chinese tech circles, the report was hailed as the ultimate technical endorsement, with tech outlet QbitAI publishing a tongue-in-cheek piece the next day titled “Anthropic, Are You Here to Advertise for Zhipu?”, framing it as free global publicity; capital markets reacted in kind, with Zhipu’s secondary-market share price climbing intraday, surging over 2% according to social platform market trackers. Meanwhile in Washington policy circles, officials and think tanks seized upon the report as the latest empirical basis for mandating pre-deployment safety testing on frontier models. A single security alert was interpreted by three different audiences as fearmongering marketing, competitor advertising, and legislative ammunition.

The same report interpreted by three groups as fearmongering, free advertising, and legislative ammunition

Many commentators reflexively attributed this divergence to commercial suppression between two rivals or dismissed it as a PR misstep by Anthropic. Yet examining the September 29 report in isolation misses deeper threads. Connecting Anthropic’s technical initiatives and policy statements from early 2026 through the autumn reveals a remarkably consistent stance on open-weights models. CEO Dario Amodei himself has stated that he has held “these positions consistently for many years.” In other words, this report stems from an established, coherent set of principles rather than an improvised PR counterattack. When the company observed an open-source model truly touch the offensive and defensive capability threshold of top closed frontier systems for the first time, this underlying logic was simply activated.

A Months-Long Posture: What Anthropic Has Said

This posture was already laid bare in early 2026, when Anthropic published its anti-distillation policy. Model distillation, simply put, involves querying Claude at scale through swarms of accounts and harvesting its responses to train another model. In its anti-distillation policy document, Anthropic explicitly characterized this practice as a security threat in black and white, establishing defenses against the exfiltration of core capabilities.

On April 7, Anthropic introduced Project Glasswing, a cybersecurity partnership that locked down models with high-risk offensive and defensive capabilities, making them available only to vetted defense organizations. Within that same initiative, the company disclosed for the first time a high-capability model never made accessible to the public: Claude Mythos Preview. According to the red team blog, during testing, the publicly deployed commercial model Opus 4.6 achieved a success rate near 0% in autonomous system vulnerability exploitation, whereas Mythos Preview operated in an entirely different tier. Anthropic stated plainly that it had no plans to release Mythos Preview to the public. Yet, in tandem with issuing warnings about these high-consequence capabilities, Anthropic announced a $4 million donation to the open-source community for open-source security infrastructure. Taken together, these two actions clearly established the baseline template of its posture: open-source models without hazardous capabilities serve as public goods, which Anthropic welcomes; but once an open model possesses unregulated offensive capabilities, Anthropic firmly opposes its proliferation.

On May 14, Anthropic released the policy research paper “2028 AI Leadership”, further articulating its strategic technical objectives. The paper argued that “it may be possible to lock in a 12-24 month lead in frontier capabilities” for democratic nations. However, this window for maintaining a lead is not infinite. According to the report’s assessment, Chinese frontier models lagged behind American models by only a narrow margin at the time—in the report’s words, “a few months behind US models.” Model distillation served as the primary shortcut for closing this gap. To illustrate the risks of technical diffusion, the report cited empirical data from the Center for AI Standards and Innovation (CAISI) under the US National Institute of Standards and Technology (NIST): under jailbreak attacks, the open-source model DeepSeek R1-0528 exhibited a compliance rate of 94% with malicious requests, far exceeding the 8% observed in American baseline models. The report cautioned that if dual-use models are released in open-weights form, users can strip away their safety guardrails locally, creating long-term, irreversible misuse risks.

By late July, the debate between the open-source and closed-source camps had burst into the open. On July 24, Nvidia led a coalition of 25 organizations in publishing an open letter titled “Open Weights and American AI Leadership”. Jensen Huang championed it personally, and social media impressions quickly surpassed ten million. This open letter advocating for open source advanced three core arguments:

First, open weights ultimately benefit defenders because defenders gain access to models as capable as those used by attackers, thereby strengthening their own posture; second, open models can be inspected and scrutinized by everyone, offering a level of transparency that is inherently safer than walled black-box systems; third, concentrating frontier models exclusively within a handful of large tech corporations creates a critical single point of failure for the entire industry. Joined together, these three tenets served as an explicit declaration: open weights equal greater security. The day after publication, OpenAI added its signature, expanding the coalition to 35 organizations. By early August, Microsoft’s official signatory page listed over 270 supporting institutions, encompassing virtually every major Silicon Valley tech giant—notable only for the absence of Chinese AI labs. Yet from this expansive roster, one prominent name was conspicuously missing: Anthropic.

Three days later, on July 27, Anthropic CEO Dario Amodei published a position statement titled “Our position on open-weights models”, responding point-by-point to the open letter. Rather than issuing a blanket rejection of open source, he began by clarifying his stance: Anthropic has never advocated for an outright ban on open-weights models, acknowledging that open models without dangerous capabilities are indeed “a public good.” But when confronting the letter’s central premise, the fundamental rift became clear. Dario argued that when considering who open weights ultimately benefit, “it seems at least as likely to me that the opposite will be true.” He pointed to the biological domain as an analogy: if a frontier model possesses advanced biological design capabilities, a malicious actor could swiftly engineer pandemic-class pathogens, whereas developing vaccines and countermeasures on the defensive side requires years of painstaking effort. Among the two catastrophic threats Dario identified as most alarming, the first was authoritarian states leveraging AI to gain decisive military and social control advantages—an outcome determined by who wields the technology and how much compute they control, independent of open source. The second catastrophic threat was cyber attacks and biological misuse. In both cyber and biological domains, open-weights models pose heightened risks because once a model is distributed to local hardware, external parties can neither enforce safety guardrails nor monitor user behavior. Crucially, “once weights are released they cannot be withdrawn.”

Dario proposed three targeted measures: strictly enforcing export controls on advanced chips, aggressively combating industrial-scale distillation, and mandating pre-deployment safety evaluations for all sufficiently capable frontier models. Addressing the open letter’s three claims, his definitive ruling was that whether risks truly exist “should emerge from testing, rather than be decided in advance.”

From projecting the lead-time window in the May report, to refusing to sign the open letter in July while establishing three defensive perimeters, to OpenAI President Greg Brockman publishing “The Defender’s Window” in mid-August—which similarly sounded alarms over a shrinking defensive timeline—Anthropic’s chain of reasoning has been tightly linked. The late-September evaluation report on Zhipu’s GLM-5.3 was simply another tangible manifestation of this well-rehearsed logic meeting real-world conditions.

Where This Stance Comes From: Irrevocability and the Power to Measure

A closer examination of Anthropic’s apprehension toward open weights reveals three progressively deeper pillars, each rooted in fundamental technical or practical considerations rather than transient PR rhetoric. The first pillar is its technical philosophy of safety: whether a model can be recalled once something goes wrong. The second is commercial alignment: this safety doctrine directly safeguards its core commercial moat. The third is an operational evaluation apparatus: it has institutionalized this stance into an automated testing pipeline that can be triggered on demand. Only by understanding these three dimensions can one grasp the internal momentum driving its actions.

The first pillar is its foundational view of safety: genuine safety depends on whether a model can be revoked if problems arise, and whether decisions can be grounded in empirical pre-release data. Anthropic has consistently maintained that safety cannot rely on the goodwill and self-discipline of end users once they hold the model. When a model is served through centralized cloud APIs, the provider can intercept policy-violating prompts and throttle call frequencies at will. But once model weights are distributed to a user’s local hardware, all cloud-dependent auditing and enforcement mechanisms instantly collapse. The UK AI Safety Institute (UK AISI) reached the exact same conclusion in a July 2026 technical blog post: once open weights are released, embedded safety guardrails can be stripped away, leaving “a persistent and irreversible risk of misuse.” Because distributed artifacts cannot be physically recalled, safety governance must be strictly enforced at the empirical pre-deployment testing stage, followed by tightly controlled access for high-risk models restricted to vetted, trusted institutions. Refusing to conflate open source with safety is the prerequisite underpinning this entire philosophy.

The second pillar is the natural convergence between commercial interests and safety principles: this philosophy effectively transforms Anthropic’s core technical assets into an irreplaceable moat. Consider the alternative: if the entire industry embraced the open letter’s thesis that open weights inherently generate security, Anthropic’s hard-won red-teaming infrastructure, pre-deployment audits, and controlled-access pipelines would lose their commercial justification. Conversely, once the ability to recall and monitor models is recognized as a necessary precondition for frontier safety, closed-source hosting emerges as the only commercial delivery format capable of guaranteeing security. Anthropic’s product roadmap reflects this alignment: through Project Glasswing, it restricted the highly capable offensive and defensive model Claude Mythos to a few hundred vetted institutions; in its dual-track release in June, the public-facing Fable 5 came equipped with stringent guardrails and automatic capability degradation on sensitive tasks, while the unconstrained, high-performance Mythos 5 was reserved exclusively for trusted channels. As tech business analyst Rui Ma observed, one does not need to assume Anthropic is engaged in clandestine machinations against rivals to recognize that its safety stance and commercial interests are, at this juncture, perfectly aligned.

The third pillar is an established, routine evaluation mechanism: Anthropic has built a pipeline that translates its philosophical stance into automated daily testing, continuously gathering first-hand data. Ever since its prominent announcement in April that Mythos Preview represented a generational leap, Anthropic effectively erected a yardstick for evaluating cyber offense and defense across the industry. Even though Mythos was never made public—and was even briefly taken offline in June following US export control directives—code security platform Semgrep’s widely circulated post, “We Have Mythos at Home” (which garnered over 1,100 upvotes), explicitly used Mythos as its benchmark. This organically drove the industry-wide adoption of Anthropic’s yardstick. Internally, the red team codified this evaluation methodology into standard operating procedure. Whenever an external model crosses key capability thresholds, this testing machinery automatically spins up to generate an evaluation report.

Three tiers behind the stance: philosophical axioms, commercial alignment, and institutional execution, with the right arrow showing the ruler’s automated output

Why GLM Is the Current Target: Thresholds and Automated Output

With these three underlying drivers understood, the reason Zhipu’s GLM-5.3 became the target becomes straightforward. Commentators frequently attribute this to US-China geopolitics, but primary sources show that this interpretation overlooks the trigger criteria of the evaluation system itself. Dario noted in his July position statement that “open weights are far less relevant than the fact that the operations are backed by an authoritarian state.” So long as they lack hazardous capabilities, open-source models remain public goods advancing the broader ecosystem. In other words, the critical conditions required to trigger the red team’s testing machinery are twofold: a model must exhibit top-tier autonomous cyber offensive capabilities while being freely distributed in open-weights format.

On August 14, Zhipu officially released GLM-5.3. On September 17, the Center for AI Standards and Innovation (CAISI) under NIST published a dedicated evaluation report, identifying it as the most cyber-capable open-weights model evaluated to date. CAISI’s benchmark results showed GLM-5.3 scoring 40.4% on the professional cybersecurity benchmark SEC-Bench Pro, surpassing the cyber capability index of the previous leader, Kimi K3, and narrowing the gap with America’s top closed frontier models to mere months. This marked the first time an open-weights model had genuinely approached top closed-source systems on core metrics of cyber offense and defense.

This milestone immediately triggered Anthropic’s pre-existing internal testing apparatus. The red team report published on September 29 was the technical embodiment of its July policy position. The first stage of testing benchmarked offensive and defensive capabilities. On the ExploitBench benchmark for automated penetration testing, GLM-5.3 successfully executed 50 end-to-end exploits across 410 attempts, compared to 56 by the baseline model Mythos Preview. In lower-level binary exploitation tests, GLM-5.3 achieved a 4% success rate versus 6% for Mythos Preview, while the publicly released Opus 4.6 and previous-generation GLM-5.2 both scored 0%. In manual evaluations, Anthropic researchers documented GLM-5.3 discovering multiple zero-day vulnerabilities in mainstream browsers within a single day. Even when experimenting with the lightweight GLM-5.3-Flash, researchers chained two known vulnerabilities into a viable exploit path with just 20 minutes of human guidance and 8 hours of model compute. Based on official Zhipu API pricing, the cost of mounting such an attack was roughly $20.40.

The second stage of testing evaluated whether the open-source model’s safety guardrails could be bypassed by end users. The report noted that while stock GLM-5.3 includes native safety defenses, targeted persona-based red-teaming achieved a 64% guardrail bypass rate; prefilling thought tokens raised the bypass rate to 92%. When subjected to model abliteration—a technique that surgically removes safety feature representations from model weights—the bypass rate reached 100%. Anthropic’s internal tests showed that completing the full abliteration process required approximately 2,200 GPU hours, representing roughly $4,400 in compute cost. Anthropic estimated that a proficient team could accomplish this in approximately 600 GPU hours. Furthermore, post-abliteration, the model’s general base capabilities remained virtually intact.

Native refusal rates are nearly identical (95% vs. 96%); the divergence in bypass rates under three techniques forms the core argument

It should be noted that the 64%, 92%, and 100% bypass rates were all derived within Anthropic’s proprietary offline simulation environment. As acknowledged in footnote 4 of the report, the model-generated exploit code was not executed against real-world environments, with certain evaluations relying on another LLM to approximate execution. CAISI’s official assessment included no such testing. Moreover, unaltered stock GLM-5.3 posted a refusal rate of approximately 95% on standard harmful requests, closely tracking Claude’s own 96%. This parity formed the primary argument raised by Chinese tech media questioning the objectivity of Anthropic’s report.

Yet from Anthropic’s perspective, whether GLM-5.3’s native refusal rate is 95% or 99% is immaterial; its thesis hinges on a single premise: once weights are made public, guardrails can be stripped. GLM-5.3 was placed in the crosshairs not because of the geographic origin of its developers, but because it was the first model across the industry to simultaneously approach top closed models in cyber capabilities and distribute its weights for free. These two conditions were precisely what Anthropic’s measurement machinery was calibrated to capture.

Two Observations Testing This Explanation

Two observations help test the coherence of this explanation. The first is that Anthropic is hardly alone in maintaining vigilance over the proliferation of cyber offensive capabilities. OpenAI began experimenting with trusted access for cybersecurity back in February, developing customized GPT-5.4/5.5-Cyber models for security practitioners. Google similarly launched Project Fairwind to provide trusted defenders with dedicated security models. In mid-August, OpenAI President Greg Brockman authored an article warning that the diffusion of open capabilities is compressing the defender’s response window. Yet both OpenAI and Google signed the industry open letter, maintaining private defenses while publicly aligning with the open-source camp. Anthropic alone refused to compromise, standing apart from over two hundred signatories. This contrast illustrates that behind Anthropic’s refusal lies a rigid, self-consistent worldview: to publicly concede that open weights are inherently safe would undermine the foundation of its own pre-release auditing and defensive architecture.

The second observation involves audience mismatch: why public ridicule failed to influence Anthropic’s posture. Following the report’s publication, tech communities joked that Anthropic had provided free PR for Zhipu. If Anthropic’s primary objective were managing grassroots developer relations, such pushback would typically prompt an internal course correction. Yet Anthropic appeared entirely indifferent to the external chatter. The reason is that the report was never written for ordinary developers in the first place. It targeted two specific constituencies: first, Washington policymakers who urgently need granular technical evidence to mandate pre-release safety testing and supply chain controls on frontier models; and second, chief information security officers at critical infrastructure organizations, for whom the documented threats create immediate urgency to procure Project Glasswing or commercial security solutions. As researcher Nathan Lambert pointed out, while the report appears defensible along a narrow technical track, it effectively reinforces a specific safety mindset. Through the lens of this audience model, the public relations storm does not hinder what Anthropic sought to achieve. What appears to the general public as a PR misstep carries virtually zero downside within the coordinates of its intended audience.

Conclusion: When Safety Becomes Gatekeeping

Looking back from the early 2026 anti-distillation policy, through the dual signals sent by Project Glasswing in April, the projected lead-time windows in May, the refusal to sign the July open letter on grounds that released weights cannot be recalled, to the targeted evaluation of Zhipu’s GLM-5.3 in September, Anthropic displays a remarkably consistent trajectory. This is no conspiracy; it is the natural convergence of an organization’s worldview, commercial incentives, and testing apparatus methodically executing its foundational logic. It accepts open models without dangerous capabilities as public goods beneficial to the ecosystem, but steadfastly refuses to equate open source with safety a priori. Under this governance framework, the authority to define safety resides strictly in pre-deployment technical evaluations and post-deployment controlled access.

Framing safety as an access-control entitlement leaves the broader industry with three unresolved questions. First is the fairness of the yardstick: the entity establishing the measurement framework is itself the largest commercial competitor in the frontier race. How can the objectivity of this ruler be guaranteed? In compute performance, mature benchmarks such as MLPerf derive their credibility from broad industry consortia and fully open-source reference implementations; by contrast, current evaluations of frontier cyber capabilities rely heavily on undisclosed proprietary environments maintained by closed-source providers.

Second is the preparation window for defenders. The UK AI Safety Institute explicitly noted that the several-month gap between closed frontier systems and open-source models provides a crucial window for defenders to prepare. If top-tier technical tools remain indefinitely restricted to a small circle of vetted institutions, are everyday developers genuinely being protected, or are they being deprived of the tools necessary for research and self-defense? Finally, GLM-5.3 will not be the last model placed on this measuring bench. Whenever a future open model crosses established thresholds in core offensive and defensive capabilities, this testing machinery will spin up once again, producing another evaluation report of identical structure. Recognizing the deep-seated sources behind this stance helps readers understand the true trajectory of AI governance and the emerging security order before the next controversy unfolds.