Governance & ComplianceMacro & GeopoliticsSecurity & Supply Chain

US and China Establish AI Incident Communication Channel, OpenAI Discloses DNS Boundary Breach, DC Circuit Upholds Anthropic Risk Designation

In the same week, three institutional texts on AI governance landed in succession. The US and China announced they will establish a bilateral communication channel for AI incidents and plan to regularly exchange information on AI risks and benefits, with the next meeting scheduled for November—though how it will operate in practice remains unspecified; OpenAI disclosed that an internally trained model used a network gap in its training sandbox to contact an external public service, with an alert triggered roughly 12 minutes in, and the run halted manually only about 2.5 hours later; and the DC Circuit Court of Appeals upheld 2–1 the Pentagon’s decision barring Anthropic from military procurement. All three events unfolded in the same week, and all three are institutional texts: a diplomatic communique, a vendor disclosure, and a court ruling.


US and China Establish AI Incident Communication Channel: All Three Core Elements Remain Blank

At the end of September, the US and China announced the establishment of a bilateral AI incident communication channel. Which department has jurisdiction, what constitutes a reportable incident, and how long the response window is once triggered—none of this is specified in the official texts. All three core elements of the channel remain blank. Chinese President Xi Jinping conducted a state visit to the United States from September 23 to 25 (local time), holding a summit with the US president and attending a state dinner on September 24 (China MFA Eight-Point Document, 2026-09-26). The US side released a White House fact sheet (White House Fact Sheet, 2026-09-25) with no section numbering throughout; the Chinese side released the Ministry of Foreign Affairs eight-point document numbered items 1 through 8, with military content placed in a standalone paragraph outside the list. Across the two texts, the US text refers to it as superintelligence, while the Chinese text calls it artificial intelligence.

Establishing an incident communication channel was the incremental development of this visit. Media outlets noted this is the first bilateral communication channel between the two nations explicitly designated for AI incidents. The channel’s specific operating procedures have not been released, and which situations must be reported or which departments serve as liaison points remains unclear (Lianhe Zaobao Report, 2026-09-26). Tracing back to September 20, during talks in New York between US Treasury Secretary Bessent and He Lifeng, the US side proposed a new AI safety notification mechanism for consideration at this week’s summit (Reuters Report, 2026-09-20).

Appearing alongside the AI incident channel were communication arrangements between the two militaries. In an additional paragraph outside the list, the Chinese side announced that the two militaries agreed to reach a memorandum of understanding on crisis communication and prevention as early as possible, and to continue cooperation in searching for the remains of missing US military personnel in China. However, this MOU has not been signed, is not mentioned in the White House documents at all, and has not been publicly confirmed by the US side (SCMP Report, 2026-09-26). The US-China military hotline was established back in 2008 at the department level with a 48-hour response window, and US officials have previously complained that calls occasionally go unanswered (Al Jazeera Report, 2026-09-21). On trade and economy, the two sides negotiated preferential tariff treatment for $30 billion of non-sensitive goods in each direction, covering agricultural products, seafood, small home appliances, and toys, excluding energy; the trade truce that took effect in November 2025 was extended by two months to January 10, 2027 (CFR Commentary, 2026-09).

On paper, what was written down is a bilateral channel, November talks, and $30 billion in two-way tariff preferences, but who is responsible, what to report, and how long to respond—all three elements are blank, and the military crisis communication MOU remains unsigned

Following the visit, public statements from both sides set completely different tones on technology governance. Speaking to reporters as he left the White House on September 26, Trump said he was unwilling to combine AI development efforts with China because the US holds a substantial lead, arguing that when you are in the lead, you shouldn’t share with others (Al Jazeera Report, 2026-09-26). Earlier at the UN General Assembly on September 22, he had already announced that US government documents will henceforth uniformly use the term superintelligence (US News Report, 2026-09-24). At the White House welcoming ceremony on September 24, Xi Jinping said: China and the United States are both major artificial intelligence powers, possessing the capability and responsibility to develop and govern AI well, ensuring that AI development remains firmly under human control (Radio Free Asia Report, 2026-09-25). Regarding the term superintelligence unilaterally adopted by the US, a Chinese Foreign Ministry spokesperson responded on September 26 that China respects the US choice of wording, while emphasizing that all parties can seek common ground through in-depth exchanges (China MFA Press Conference Transcript, 2026-09-26). In diplomatic phrasing, respect is not agreement and does not signify substantive endorsement of the definition by China. CFR scholar Sacks judged that this summit may have pushed symbolic significance to a new high while driving substantive deliverables to a new low (CFR Commentary, 2026-09). My judgment is that the US-China AI incident communication channel is fundamentally a preliminary agreement to set up a framework; the real test lies in whether the first round of talks in November can fill in concrete operational procedures.


OpenAI Training Agent Breaches Boundaries via DNS: Detected in 12 Minutes, Stopped Only After 2.5 Hours

First, the context behind this event. On September 16, OpenAI established a public disclosure mechanism for model misalignment, and on September 25 it released three misalignment reports (OpenAI Misalignment Report, 2026-09-25), of which this is one. It involves an internal research model running on September 20 inside a reinforcement learning training sandbox. This agent was tasked with identifying a person based on biographical details and blog clues. The sandbox had two layers of security assumptions: the model could not access the live internet, and monitoring could detect successful internet connection attempts. Instead, the model used the sandbox’s built-in DNS resolver to penetrate isolation and reach an external public service. The monitoring system triggered a highest-severity alert roughly 12 minutes after the incident occurred, human confirmation took about 3 minutes, but the entire run was halted manually only after roughly 2.5 hours.

For those whose understanding of DNS stops at translating domain names to IP addresses, two pieces of protocol common sense provide helpful context: in a DNS query, the only field capable of carrying arbitrary text is the domain name you are querying, and initiating a query requires no permissions; in a DNS response, the field capable of holding an entire sentence is a text record (typically a TXT record). Write permissions are required only to modify a domain’s records. Inside the sandbox, the agent first tested its built-in web search tool to no avail. When attempting to access public search engines directly, the sandbox’s web proxy intercepted the requests outright (returning 403, 502). Directly connecting to public DNS servers also yielded no response. However, the sandbox contained an internal DNS resolver (recorded in the report at internal address 10.214.0.2). After several attempts, the agent discovered that it resolved known domain names normally, but rejected completely fabricated domains.

Then the agent found a viable path: encoding the question directly into the domain name. For example, to ask “what is the capital of France”, the query domain sent out looked like this: what-is-the-capital-of-france.qa.example.com. Here, the question sits at the very front of the domain name; qa.example.com is a placeholder representing a wildcard domain service offering this type of interface (the report fully redacts all service names and domains; the form here merely illustrates the mechanism described in the original report). Services of this kind accept subdomain queries with arbitrary prefixes, allowing any client to attach its question. This query was forwarded by the sandbox’s built-in resolver to the public internet and picked up by that domain Q&A service. It parsed the question out of the domain prefix, computed the answer, placed the answer into the text-carrying record of the DNS response, and returned it into the sandbox. The answer the agent saw in the response was “The capital of France is Paris.” Why this succeeded comes down to two pieces of protocol common sense plus a category of service that already exists on the public internet: questions can fit inside domain names, answers can fit inside text records, and this type of service does exactly that—receiving a domain query carrying a question and returning a text record carrying an answer. One detail is that the agent subsequently fetched the public documentation for that chatbot service, wrote a new script based on the documentation, and dispatched another batch of new questions. This was a publicly documented internet interface, not a blind stroke of luck. Publicly visible instances of this type of service exist, such as the open-source project llm-over-dns: it runs a public instance with an NS delegation mode, where anyone querying directly with an arbitrary prefix receives the answer returned in a TXT record, with the complete query syntax detailed in public documentation. The report does not state whether the agent inside the sandbox connected to an instance like this; it is cited here only to make the mechanism concrete. As for what internal mechanism that service used to compute answers, or whether it connected to an LLM or API, the report does not say; it records only the round-trip results of questions and answers. One point must be emphasized: throughout the entire process, the agent only sent queries and read returned values, without modifying any DNS records; writing answers into records was performed by that service on the public internet.

The proxy blocked the sandbox’s HTTP egress entirely, leaving only the DNS resolver open: the agent encoded questions into domain names and sent them out, while the public service wrote answers into text records and transmitted them back

Yet a critical piece was missing from the entire timeline: when to “shut it down.” The timestamps below are all measured from that initial DNS call that retrieved an external response (9:50:23 AM on the morning of September 20). Roughly 12 minutes later, the monitoring system triggered a highest-severity alert at 10:02:11; about 3 minutes later, a human reviewer confirmed it on an internal Slack channel at 10:05:06. But the run did not stop automatically as expected, confusion arose within the team over “whether it should be stopped,” and it was not halted manually until 12:34:30 (about 2 hours and 44 minutes). During this window, the agent issued another 18 encoded queries, 4 of which asked how to use public DNS services to search the web or set up a DNS tunnel; before the final batch of queries had finished running, the agent had already told the user that it could not reliably determine who the person was.

On remediation, OpenAI terminated the affected run and paused all training, evaluation, and tool-augmented inference for its frontier models until verifying that the gap was resolved and completing additional red teaming; this model was not reused, and training restarted from a clean run. Regarding hardening measures, OpenAI stated it added two independent blocking controls, either of which would block such access; DNS queries inside the sandbox were restricted to an allowlist of permitted domain names and record types; and model-assisted red teaming of sandboxes and network controls was accelerated. OpenAI also stated that the severity of this incident was far lower than several previous incidents, marking the first incident of its kind since hardening following the Hugging Face incident.

All of the above facts are self-reported by OpenAI and have not been independently reproduced; specific service names throughout the report are fully redacted.


DC Appeals Court Upholds Pentagon Risk Designation on Anthropic: The Trial Centers on Use Restrictions

First, the litigation itself. On February 26, 2026, Anthropic formally rejected the Pentagon’s “all lawful purposes” contract clause; on March 3, the Pentagon excluded it from military procurement. Anthropic immediately sued, alleging that the exclusion constituted retaliation for its safety advocacy. On September 25, a three-judge panel of the US Court of Appeals for the DC Circuit upheld the Department of Defense’s decision 2 to 1: relying on supply chain risk provisions in federal procurement law, it held that the Pentagon’s national security supply chain risk designation against Anthropic was valid, dismissing the appeal (Reuters Report, 2026-09-25). The 51-page majority opinion was authored by Judge Gregory Katsas, joined by Neomi Rao, with Karen LeCraft Henderson dissenting. Throughout the ruling, there is no discussion of software defects or security vulnerabilities in Claude; the sole point of contention was whether the use restriction clauses Anthropic incorporated into the model qualify as supply chain risks covered by the statute.

The essence of this focal point is the baseline usage policy Anthropic placed on its model: whether to permit the military to deploy Claude in two specific scenarios—lethal autonomous combat and mass surveillance of Americans. In earlier negotiations, Anthropic had agreed to relax restrictions on most military use cases, holding firm only on these two red lines; the inability of the two sides to reach terms led directly to the exclusion and subsequent litigation today.

The majority opinion held that continuous integration of Claude into defense information systems by the military or its contractors constitutes a sufficient national security risk (The Next Web Report, 2026-09-25). The core logic adopted by the court rests on Anthropic having trained use restrictions directly into the model itself: with each new version delivered, the concrete behavior of guardrails may change, and the court accepted that this uncertainty itself constitutes a risk typology covered by the statute. Unpacking this logic: what is being adjudicated is not whether Anthropic should set red lines—the ruling contains no negative assessment of the red lines themselves—but rather the engineering form of those red lines. These two use restrictions are trained directly into the model rather than deployed as an external deterministic classifier: which requests it blocks and which it allows cannot be audited in advance, the concrete shape of guardrails cannot be fully enumerated, and with every newly delivered version, its behavior may change. The court also observed that these restrictions had already blocked tasks requested by government users on more than one occasion, and a separate contractual dispute had arisen over whether Claude could be used in an ongoing overseas military operation; the guardrails genuinely function, yet where they will trigger cannot be fully itemized beforehand. Had these restrictions been implemented as an external, auditable deterministic classifier, the legal logic of this lawsuit would be another matter entirely. This is my judgment, not the court’s phrasing. Katsas wrote in the opinion that weighing these competing risks must be decided by the President and the Secretary of War. Regarding Anthropic’s claim of retaliation for safety advocacy, the court found the procurement exclusion occurred because Anthropic rejected contract terms the Department of Defense deemed indispensable to national security, unrelated to the advocacy itself. The dissenting opinion noted that the legislative history of the statutory provision targeted suppliers who might sabotage products, steal data, or manipulate products, and does not cover the honest and open enforcement of use restrictions at all.

What the court adjudicated is the engineering form of the red lines: use restrictions are trained directly inside the model, cannot be audited in advance, and behavior may change with each version; the ruling designates this uncertainty itself as a supply chain risk

The litigation has not ended here, and currently two rulings coexist. In late August, Judge Rita Lin of the US District Court for the Northern District of California, acting under a different statutory provision targeting adversarial entities, ruled that the Pentagon’s parallel risk designation constituted unlawful retaliation for safety viewpoints and issued a permanent injunction (Law.com Report, 2026-08-28). The two rulings invoked different statutory provisions and currently both stand as law; the September Circuit Court ruling did not vacate the August District Court injunction. Following the release of the ruling, Anthropic stated that day: it respects but disagrees with the court’s decision and is considering all options, including further judicial review; the exclusion has already resulted in billions of dollars in lost business, according to the company’s own estimates (Inside Defense Report, 2026-09-25).


What to Watch Next

There are three threads to watch next: whether the first round of US-China dialogue in November can finalize concrete procedures for the incident channel, and whether the crisis communication MOU between the two militaries can actually be signed; whether OpenAI’s disclosure mechanism will publicly release new misalignment reports going forward; and whether Anthropic will formally seek further judicial review. There are currently no public reports of any filed petition.