Industry & CompetitionGovernance & ComplianceScience & Tech Frontiers

Four Observations on AI: AWS Silicon Partners, Autonomous Weapons, Mathematical Methods, and Model Similarity

Recently, I read about four distinct developments. On September 8, Qualcomm announced an expanded data center custom silicon collaboration with Amazon; Anthropic disclosed that developers had used Claude Code to write software for autonomous attack drones; twenty-five Fields Medallists jointly signed a call warning against turning mathematical research into a benchmark-chasing contest; and Arena published a quantitative measurement of response similarity across mainstream large models. These developments lie in different technical areas and are independent of one another; below is a separate record of each.

1. AWS Expands Silicon Partners: Broadening Paths Beyond In-House Development

AWS has long been developing in-house silicon, but in looking at this Qualcomm partnership, what caught my attention is that it continues to expand its circle of partners. In many technical discussions, people tend to view a cloud provider’s in-house hardware and its external partnerships as mutually exclusive paths. In terms of actual computing capacity build-out, in-house hardware and external custom partnerships can proceed entirely in parallel. Facing rapidly accumulating online computing requests, laying out multiple supply channels gives the infrastructure team broader room for scheduling.

According to Qualcomm’s official press release published on September 8, the two parties plan a multi-generational product collaboration to jointly develop custom data center silicon for AWS, focusing on trained AI inference workloads. In the release, AWS Vice President Prasad Kalyanaraman stated that the collaboration on custom silicon and advanced connectivity will deliver more performant, efficient, and cost-effective infrastructure for customers. AWS itself has an in-house chip lineup including Trainium and Inferentia, covering model training and inference tasks. Reaching a multi-generational collaboration with Qualcomm indicates that while advancing its in-house hardware, the platform continues to broaden external custom partnerships.

Beyond computing silicon itself, this collaboration extends to high-speed interconnect solutions for data center networking. The two parties plan to develop optical interconnect solutions extending up to 1.6 trillion bits per second, involving underlying modules such as data transmission and optical digital signal processing. According to the press release, these solutions are designed to support the growing scale and bandwidth demands of AI infrastructure.

What I find interesting here is the interaction between the two parties at the software and tools level. Qualcomm mentioned that its team plans to expand the use of AWS AI infrastructure, including Amazon Bedrock, for electronic design automation. Qualcomm hopes to reduce chip design cycles through this, though the press release only outlines usage plans without providing quantitative data on specific cycle time reductions. The two companies have thus formed a two-way business relationship: Qualcomm plans to expand cloud-based chip design work, while AWS collaborates with Qualcomm on custom data center silicon.

Reviewing the history between the two companies, Qualcomm had already entered the AWS ecosystem earlier. Prior to this, AWS had already launched Amazon EC2 DL2q instances powered by Qualcomm Cloud AI 100 chips, providing developers with a low-power inference environment. The multi-generational collaboration reached this time is a further deepening of their existing business relationship, indicating a clear consensus between both sides on long-term joint customization.

The contents announced in Qualcomm’s press release primarily establish a long-term collaborative framework. The release did not disclose specific chip model names, nor did it provide a delivery timeline or third-party benchmark results for the initial chips. When the silicon will ultimately tape out and land in data centers, and how it will coordinate in production workloads alongside the Trainium series chips, remains to be verified once specific products appear in the future.

What can be seen from this is that AWS is not staking its choices on a single path. It already has its in-house Trainium and Inferentia chip lines, and when facing silicon partners it has already worked with, it is also willing to sit down and develop multi-generational products together. Walking both paths simultaneously leaves the platform more maneuverability when allocating compute hardware.

AWS simultaneously advances in-house development and external custom partnerships; connections in the diagram indicate collaboration paths, not that new chips have been delivered.

2. Claude and Autonomous Weapons Software R&D

In its security threat report published on September 10, Anthropic documented a concrete engineering R&D case. To develop software for autonomous attack drones, a technical team directly invoked the general coding assistant Claude Code to write and maintain engineering code.

According to the disclosures on pages 117 to 119 of the report, the involved group was designated GTG-27005, operating in Russia under the project names DronDoc or Serafim. Anthropic assessed that this was a freelance technical team doing a mix of civilian and military work, not a Russian state entity. The group claimed to have received funding support from military and other institutions, though Anthropic noted in the report that it could not independently verify these funding claims. In practice, the group used Claude Code to write and test code, saving the generated code directly into their own project files.

What drew attention in this case was the goal set for the software. Anthropic noted in the report that the system was designed for autonomous lethal engagement. The team designed the onboard model to recognize various target classes, including personnel, and issue attack commands autonomously without requiring step-by-step confirmation from human operators. Placing target filtering and firing decisions into software logic brought code directly into the core loop of engagement.

One technical detail documented in the report was that the team integrated real hardware development boards into their testing pipeline, attempting co-debugging between software code and underlying hardware. Hardware-in-the-loop testing indicates that the code had begun moving beyond pure theoretical text toward physical verification. Among the six rows of systems or subsystems listed, Anthropic rated the technology readiness level of five rows as TRL 3–4, with the remaining row designated as doctrine and simulation. This maturity assessment reflects the vendor’s own technical judgment; the report provided no certification from independent external bodies.

However, it must also be recognized that this is currently an R&D record rather than evidence of combat deployment, and the sole information source is Anthropic’s own security report. The report demonstrates that actors invoked general-purpose programming tools to assist in writing weapon code, but it provides no field evidence that the system has been fully manufactured or deployed to the battlefield. Keeping this clear prevents mistaking software debugging in an experimental phase for operational combat equipment.

The report also mentioned another Yemen case involving the Middle East. A team with hardware engineering backgrounds returned to the Claude platform within hours of a suspected failed guided rocket test-fire to analyze telemetry data and troubleshoot technical issues. Anthropic noted that it had no evidence the group possessed operational equipment ready for combat. This example illustrates the model’s assistive role in complex engineering troubleshooting, but this record did not specify whether the project involved autonomous attack decisions; it cannot be used to corroborate autonomous attributes of weapons, nor should it be conflated with the preceding drone project.

The immediate reminder this brings is very realistic. When people use everyday coding tools for such R&D, risks do not have to wait for models to exhibit spontaneous malice. As long as someone inputs requirements and organizes code, existing general-purpose tools can already become entangled in dangerous physical projects.

3. The Purpose of Mathematical Research: Preserving Methods or Chasing Answers

Large language models have repeatedly broken records in math competitions and automated theorem proving, frequently sparking enthusiastic discussions in technical communities. On September 11, mathematician Terence Tao reposted on his personal blog a joint declaration published on mathandai.org, titled “A Severe Misalignment of AI in Mathematics.” This declaration was supported by twenty-five Fields Medallists as initial signatories, including Terence Tao.

The signing mathematicians did not dismiss the value of technology itself. In the declaration, they acknowledged the progress of large models in mathematical problem-solving and recognized the potential of tools to enhance mathematical understanding. Where the scholars took issue was the gamified practice by commercial institutions of packaging the conquest of isolated propositions as model benchmark leaderboards. The declaration pointed out that solving problems is only a tool and proxy for achieving conceptual understanding and scholarly insight. An excessive bias toward mass-producing true-or-false conclusions risks damaging the academic environment that fosters new ideas.

In rigorous mathematical exploration, derivations certainly require logical correctness, but that does not mean the conclusion itself is the finish line. In many cases, even when a famous conjecture is proven or disproven, related exploration does not end. The new mathematical formulations created by researchers during derivation, along with the analytical tools and theoretical frameworks abstracted, can often transfer to other mathematical branches and even other scientific fields. If the evaluation of the entire field revolves around final true/false labels, foundational exploration capable of opening new domains may be neglected.

We can consider a hypothetical scenario: an automated system discovers a set of elaborately constructed counterexamples that happen to disprove a prominent conjecture, and benchmark leaderboards immediately mark the problem as solved. If people consider the work finished simply because the counterexample holds, the mechanistic analysis, property classification, and new tool building that might otherwise unfold around the counterexample could lose motivation for sustained investment. Good counterexamples often inspire entirely new theoretical directions; what matters is whether, when faced with a counterexample, people stop at recording the score of disproving a conclusion, or continue questioning the mathematical mechanisms behind the counterexample.

Placing all attention on the final answer is analogous to engineering optimization overfitting to a single metric: one easily obtains the numbers while setting aside the conceptual foundation that requires long-term accumulation. This does not mean commercial labs produce no rigorous academic output at all; for instance, when OpenAI published its fluid dynamics research, it simultaneously provided both the academic manuscript and formal verification code. The trend mathematicians worry about is that pure leaderboard gamification dominates research evaluation, allowing the rush for answer labels to overshadow the accumulation of foundational theoretical frameworks.

Mathematical research should not only reach correct conclusions, but also leave behind methods that can be applied to other problems.

This divergence in philosophy manifested in an event at Caltech. Several current and former Caltech mathematics faculty and researchers published an open letter on Proofs & Prompts, criticizing the Caltech Mathathon originally slated to receive compute sponsorship from OpenAI and Anthropic. The event’s initial round was designed as a 40-hour extreme problem-solving sprint. According to the organizers’ response, only teams selected and advanced by judges would enter the second round, receiving a six-month verification and digestion research period. The open letter criticized this hackathon model for treating scholars as an uncompensated validation team, saddling them with the heavy burden of checking vast volumes of model output for free.

Facing feedback from the academic community, OpenAI quickly adjusted its sponsorship arrangements. On September 10, researcher Dan Roberts publicly announced on social media that, in light of the mathematical community’s concerns, the company had decided to withdraw OpenAI’s sponsorship of the Caltech Mathathon, expressing hope for more conversations and discussions in the future; this adjustment was subsequently reported by Business Insider. Dan Roberts announced the sponsorship withdrawal prior to Terence Tao’s repost of the joint declaration the following day. The organizers told Gizmodo in an interview that they were seeking alternative sponsorship and that the competition was not cancelled.

To me, what kind of methods remain after solving a problem is often more important than simply capturing a number. Seeing a method help make sense of another problem, or provide a handy tool for subsequent derivations, is often how mathematical research moves forward through accumulation. No matter how tools evolve, if they merely produce piles of true/false checkmarks for leaderboards, they risk missing what is truly inspiring in mathematics.

4. Similarity of Model Answers: Not Necessarily Bounded by Vendor Lines

Intuitively, people tend to assume that models developed by the same team will have closer output styles and content, while noticeable stylistic divergence will exist across different vendors. On September 11, Dawid Galarowicz and Peter Gostev published an analysis of model conceptual similarity in an official Arena article, where measurement results revealed a different factual picture. It should be noted that this article is a research insight shared by Arena’s official account rather than a peer-reviewed academic paper.

This analysis was based on 30,086 Text Arena comparison pairs on English prompts collected between May 1 and September 2. The sample size here corresponds to the total number of paired responses where two models answered the same user prompt, which cannot be directly equated with the total number of independent prompts. To measure content overlap between two models answering the same question, researchers introduced an LLM-as-a-judge. The judge extracted up to five text-supported core ideas from each model’s answer, then reviewed both original responses to compare the two sets of ideas, categorizing ideas shared by both models and those unique to each.

The calculation of similarity is straightforward: the number of shared ideas divided by the total number of distinct ideas after merging and deduplication. To reduce positional bias from presentation order, the algorithm repeated the evaluation with the input order of the two responses reversed, taking the average of the two passes. For example, given the same prompt, Model A and Model B each present three identical core insights, while Model A adds one unique perspective and Model B adds one unique angle, resulting in five distinct ideas in total across both responses. With three of those five ideas shared, their idea overlap is three-fifths, or 60%.

In the statistical chart published with the study, the research team listed overlap data for several representative model pairs. Among them, comparisons between Fable 5 and several other models showed an interesting distribution:

Model Pair Idea Overlap
Fable 5 × DeepSeek V4 Pro 59.2%
Fable 5 × GLM 5.3 55.9%
Fable 5 × Claude Opus 5 41.2%
Fable 5 × Claude Sonnet 4.6 37.9%
Across these four pairs, Fable 5 has higher idea overlap with DeepSeek and GLM than with Opus and Sonnet.

From these published figures, Fable 5’s idea overlap with external vendors’ DeepSeek V4 Pro reached 59.2%, and with GLM 5.3 reached 55.9%. By contrast, its overlap with Claude Opus 5, from the same Claude family, was 41.2%, and dropped further to 37.9% with Claude Sonnet 4.6. In this set of pairings, cross-vendor idea overlap noticeably exceeded internal intra-family overlap. While the original chart does not depict a complete matrix of all frontier models, these comparisons illustrate that model output similarity does not bear an inevitable correspondence to brand affiliation.

This high cross-vendor similarity also appeared in other tested pairs. As reported by the authors, the highest observed overlap pair was Kimi K3 and Muse Spark 1.2, reaching 63.5%. For comparison, between adjacent versions within the same family, Grok 4.5 and Grok 4.6 shared 59.7% of their ideas, and GPT 5.5 and GPT 5.6 shared 59.2%. Models from two completely different development teams exhibited higher idea overlap on identical questions than adjacent generations from the same vendor family.

Across the overall sample, the average idea overlap for all tested pairs was 43.1%. This figure represents the aggregate mean across all tested pairs; it does not mean that most models’ viewpoints converge heavily, much less that over 40% of responses are identical. Idea overlap varied across task categories: medicine and healthcare tasks averaged 51.1% overlap, while creative writing dropped to 34.9%. The authors’ explanation is that under scenarios with strong factual constraints, different models tend to extract consistent core insights, whereas in open-ended creative scenarios, each model’s choice of perspective and argumentation is more dispersed.

This measurement itself has clear limitations. The official article represents an observational snapshot that did not disclose the specific judge model selection, per-pair sample sizes, or confidence intervals, nor did it evaluate idea quality or correlation of errors. Similarity of output ideas reflects the intersection of answer content; it cannot be used to infer internal model architectural design or training data sources, much less be hastily taken as evidence of model distillation or code copying. Stepping back from brand labels alone, the phenomenon where a cross-vendor model like Fable 5 lands closer in answer content than its own family siblings provides an intriguing lens of observation.