On September 30, 2026, DeepSeek and Huawei open-sourced a suite of software components for Huawei Ascend chips, comprising six repositories that cover matrix multiplication, attention, a general operator library, Top-k selection, distributed communication, and compilation tools. Huawei also announced SuperPoD Flex, a 128-card Ascend 950 super node solution, for which these components were built. None of these are toy demos; they are the most compute-intensive, hardest-to-optimize operators in large model training and inference. Officially, they are categorized as compilation tools, high-performance computing libraries, and distributed communication libraries. There are no scheduling-layer modules.
When teams evaluate a secondary platform, they hold two rulers: one measures effort saved, and the other measures performance. Teams running all the way to final training are most sensitive to the second. In the past, everyone working on cross-platform portability asked the same question: how to modify a single codebase so it runs on both NVIDIA and other chips. That question was asked for years without ever yielding a good answer. DeepSeek asked a different question instead: between two chips, which parts must be identical, and which parts are allowed to differ? Broken down, this yields four sub-questions: Does model code need to change? How is data laid out in device memory? Who writes each operator? Are development tools shared? This batch of components delivers answers to each sub-question, and all answers are publicly verifiable.
Why the old question never yielded a good answer comes down to a physical reality. For the exact same matrix multiplication with the exact same numbers, the bytes stored in device memory are completely different: even if your function calls and arguments are identical, the way the two chips arrange this data is simply not the same. Forcing uniformity at the byte level inevitably means one chip must pay the overhead of data conversion, or even forfeit its core compute units. Reaching peak performance requires running close to the metal, but staying close to the metal precludes uniformity. This contradiction holds at the physical layer, regardless of software quality.
The entire industry has wrestled with this contradiction for years. Google, OpenAI, and Moore Threads each chose a different layer to unify; all succeeded on their home turf, and all hit a wall when moving to a secondary platform. Once AI-assisted coding became a reality, the ledger shifted: the labor cost of writing a second implementation plummeted, re-energizing chip vendors, framework teams, and model groups alike. Yet lowering costs only solved the willingness to write; regarding what to lock down and what to leave open across those four sub-questions, no one had ever provided a complete answer in public.
This batch of components puts the answers on the public table; our investigation found no prior open-source precedent of comparable scope. When porting real production operators to new hardware, what can be reused as-is, what must be rewritten, and how much engineering overhead rewriting entails now have line-by-line verifiable code evidence. Below, we first examine where each of the three legacy routes stalled, then see how physical constraints press down the ceiling, look at what AI coding changed and what it left unchanged, dissect how these components price compatibility across layers, and finally calculate the costs of this approach.
Let us first examine the three legacy routes and where each positioned itself within this dilemma. Over the past decade, to escape dependency on a single chip vendor, the community explored three archetypal approaches: Google placed its hopes on compilers, OpenAI on operator programming languages, and chip vendors on API compatibility. All three routes chose different balances between hardware performance and usability.
XLA/JAX: Signing the compatibility contract at the computational graph compiler layer. Just define the graph, and leave the details to the compiler. Co-designed with Google’s proprietary chips, it runs fast; but on general-purpose GPUs—especially for high-throughput matrix multiplication and sparse attention critical to LLMs—compiler-generated code performance has consistently lagged behind manual engineer optimization. In production pipelines, the hottest operators still rely on expert hand-tuning.
The second route is the programming model approach spearheaded by OpenAI, exemplified by Triton. It signs the contract at the operator programming model layer, allowing developers to write high-performance kernels using Python-like syntax. The abstractions Triton provides essentially map one-to-one to NVIDIA hardware. While virtually cost-free on native hardware, once moved to other chip architectures, those seemingly natural hardware assumptions turn into heavy adaptation burdens. The secondary platform’s operator ecosystem has consistently lagged behind the primary platform in both maturity and tuning depth.
The third route is the API cloning approach commonly adopted by chip vendors, such as Moore Threads’ MUSA. It sets the contract directly on the API surface of the mature ecosystem, aiming to compile and run existing code on the new chip without modifications. For customers eager to get legacy workloads up and running, this indeed offers the lowest migration barrier; however, the pipeline and microarchitecture of the new chip are often constrained at every turn to emulate the established behavior of legacy interfaces, making it difficult to unleash the native potential of the underlying hardware.
These three routes lock down things at different layers:
| Route | Championed by | Compatibility Unit Locked | Where It Stalls |
|---|---|---|---|
| XLA / JAX | Google, framework developers | Locks the computational graph; delegates kernels to the compiler | On the hardest LLM operators (FP8 GEMM, sparse attention), compiler-generated kernels fail to match hand-written performance; production hotspots on NVIDIA still follow the hand-written route |
| Triton | OpenAI, framework team | Locks a programming model with 1:1 mapping to NVIDIA hardware | This lock is free on NVIDIA but becomes an overhead on other chips; secondary platform kernel ecosystems have not approached the NVIDIA side |
| MUSA | Moore Threads, chip vendor | Locks the CUDA API surface, allowing existing CUDA code to run as-is | The new chip is forced to emulate CUDA behaviors, forfeiting the freedom to exploit its own microarchitecture |
| Boundary deduction | Model laboratories | Locks operation semantics, the minimal compatibility unit | Different routes lock different units; the larger the unit, the more freedom surrendered on new hardware. Model labs are locking operation semantics because only they possess both the algorithms and invocation call sites; our investigation found no prior open-source precedent of comparable scope. |
All three legacy routes hit a wall, and how they hit it demonstrates the exact same point: which layer compatibility is signed at directly dictates how much freedom remains for the secondary chip. Why this holds true comes down to the workloads of large models themselves. Throughput bottlenecks in training and inference are almost exclusively concentrated in three classes of operators: matrix multiplication, sparse attention, and cross-node expert communication. They must squeeze compute units and memory bandwidth to their physical limits, leaving zero room for waste.
Specifically, GPUs and Ascend embody different hardware design philosophies: GPUs feature dedicated high-speed interconnect structures, allowing data transfers between external memory and internal caches to overlap seamlessly with compute pipelines. Ascend separates matrix units and vector units into two distinct cores, relying on explicit programmatic scheduling and synchronization for on-chip data movement. Consequently, expecting to smooth over differences directly via microarchitectural instructions inevitably leads to pipeline scheduling breakdown.
A concrete example lies in popular FP8 low-precision matrix multiplication, where scaling factors are introduced to preserve precision. NVIDIA hardware requires packing four scaling factors into a 32-bit integer and storing them in device memory in a specific transposed layout to facilitate reads by data movement units; Ascend hardware rules, however, require packing two scaling factors into a 16-bit integer, with completely different row and column layout requirements.
If data layouts at the byte level are rigidly bound between both sides, the new chip is forced to spend massive clock cycles unpacking, transposing, and moving data, or even prevented from engaging its core high-speed compute modules. The more fine-grained the contract, the higher the cost for the new chip to mimic legacy behavior, and the lower the achievable throughput ceiling. Such losses cannot be compensated for by local code tweaks, because internal data path widths and layout rules are etched into silicon.
Objectively speaking, XLA’s success on TPUs and Triton’s high efficiency on its primary platform demonstrate that these methods are highly effective in their native domains. The problem arises when attempting to achieve peak hardware performance for the same operation across two different chips. This is a physically hard problem. No matter how cheap software becomes, it cannot alter how hardware organizes data.
A deadlock unyielding at the physical layer was pried loose from the cost side by AI programming: the labor cost of writing a second implementation became cheap, something directly visible in these repositories. The repository of compilation tool tilelang-ascend contains extensive engineering traces of AI-involved development: 21 skill files callable by agents, four agents chained into a three-stage state machine covering operator design, implementation, and performance tuning; CI automatically pulls examples to run benchmarks, and commit histories frequently feature model co-authored signatures. Huawei’s developer community has also publicly published specifications for agents writing operators, and officials claimed an agent could write an operator in half an hour. This figure covers only the code-writing phase; safety nets such as numerical verification, cross-version regression, and communication correctness remain the domain of senior engineers.
Deducing along two legacy paths reveals where AI’s boundaries lie. Continuing down MUSA’s path of API compatibility: no matter how fast AI writes compatibility layers, the output remains emulation; usability is addressed, but the performance ceiling does not budge an inch. Attempting to directly align with the low-level instructions of rival chips: that track shifts with every generation—when a competitor swaps out a compute component, all crafted adaptations become obsolete. AI can accelerate code generation, but only where agreements already exist; deciding at which layer to establish that agreement remains up to humans.
Thus, the question shifted from whether engineering bandwidth exists for two implementations to which layer the contract should be signed at. Once code generation became cheaper, teams had the luxury for the first time to calculate another balance sheet: if the upper layers remain unified while the lower layers are implemented separately per chip, can the maintenance and business ledgers balance out?
DeepSeek’s answer to this ledger is the code that follows. First, what are the six repositories: TileLang-Ascend (a DSL for writing kernels in Python), DeepGEMM-Ascend, TileKernels, FlashMLA, DeepSelect, and DeepEP-Ascend. Official announcements position them as compilation tools, computing libraries, and communication libraries, with no scheduling-layer modules; online serving scheduling and cluster load balancing were explicitly excluded test conditions and are not part of the open-source scope.
The intuition behind how this software handles differences between two chips boils down to a single sentence: fully guarantee upward commitments, and leave downward differences open. In terms of how models invoke these operators, interfaces on both sides are completely identical, requiring zero lines of model code modification; below the interface, the two chips operate entirely in their own ways. Broken down into four layers:
| Layer | Strategy | Evidence |
|---|---|---|
| Invocations, Python API | Locked down: identical names and parameters | Zero changes to model-side invocation code |
| Data bytes, scaling factor layouts, etc. | Intentionally differentiated: preserve semantics, not bytes | FP8 scaling factors: NVIDIA packs 4 into a 32-bit integer, Ascend packs 2 into a 16-bit integer, with differing row/column requirements handled by user-side conversion functions |
| Kernel implementations | Dual-tracked: same high-level domain-specific language, separate
hand-written _cuda.py/_asc.py |
engram_fused_weight: CUDA version uses thread binding; Ascend version uses core specialization and multi-version buffering |
| Development tooling, DeepJIT | Truly shared: one library used on both sides | NVIDIA versions of DeepGEMM/DeepEP adopted it starting from V2.5 |
How these four layers are implemented can be seen in three places across repository code.
The first is how data is arranged in device memory. For the same matrix multiplication and the same numbers, the byte layout differs on both sides—the scaling factor packing mentioned earlier is a prime example. Rather than smoothing over this discrepancy, DeepSeek treated it as a design starting point: format conversions are consolidated into a user-side function that rectifies formats as data enters before passing it to operators. Mathematical semantics remain unchanged, while byte layouts are determined independently by each hardware architecture.
The second is how each operator is written. The same operator comes in two implementations—one for NVIDIA and one for Ascend—each written in the same language. Comparing engram_fused_weight, the simplest operator in the repository, not a single detail matches across sides: thread division of labor, where data is placed, and how data is moved are all completely different. Two versions were written because the execution mechanisms of the two chips differ too drastically to accommodate with a single codebase.
The third place is the most telling: the only cross-platform component shared across the entire codebase is development infrastructure. Engineering steps including runtime compilation, hash caching, and loading were factored out into a standalone lightweight library, DeepJIT, shared by both sides; even NVIDIA’s own matrix multiplication and communication libraries adopted it starting from V2.5. Operators are written separately, while the tool that builds them is shared. The communication library follows the same philosophy: upper-layer interfaces align, while the lower layer switches to Huawei’s proprietary HCCL communication stack; unfinished features are publicly listed in documentation with placeholders in code. Looking across the four layers together: operator semantics are locked down for free, byte discrepancies cost a conversion function, peak performance demands custom implementations on each side, while tool-layer sharing actually saves one copy.
The cost of this approach is straightforward. Among the three heaviest components in this release—the core body of the matrix multiplication library, the attention library FlashMLA, and the Top-k selection library DeepSelect—the compute core uses no high-level languages whatsoever, relying directly on Ascend C for low-level implementations; ‘directly’ here refers to bypassing DSL generation, which is distinct from whether AI assistance was employed. Ascend C is Huawei’s low-level programming language for Ascend chips, analogous in role to CUDA C. Among them, the Top-k implementation is condensed into a single file of 54 KB and approximately 1,800 lines. Switching to a new chip means rewriting this code from scratch.
Having paid such high costs, how large is the actual return? This suite of components published a series of self-reported benchmark figures. Bottom line first: the numbers look impressive, but all stem from vendor self-benchmarks, each comes with a string of prerequisites, and none have been verified by third parties to date. Let us examine them one by one. The highest utilization figure comes from the matrix multiplication library. In DeepGEMM-Ascend’s published self-test table, a BF16 dense GEMM case on 950DT reports 431 TFLOPS (M=4096, N=7168, K=16384, CANN 9.20, cold L2); taking the 432 TFLOPS specification cited by maintainers as the denominator, this represents roughly 99.8%. Our verification obtained no independent replication, and GitHub issue #1 challenged the specification baseline. The test subject is a native Ascend C kernel written directly in low-level code without going through a DSL, so it cannot be used to prove that the DSL can also reach peak chip performance. Meanwhile, in the discussion thread of GitHub issue #1, developers pointed out that if calculated against a higher-spec floating-point tier as denominator, the operator’s actual utilization is approximately 85%; the issue remains open today. The most expensive part was not made cheaper by AI. On attention operators, the 9/30 release of FlashMLA reported 410 TFLOPS for prefill (95% of hardware peak) and 360 for decode (83%), V4.1 only, dav-3510, vendor self-tested.
Communication figures carry the heaviest prerequisites. Bandwidth figures for DeepEP-Ascend come from a PoC HDK provided to DeepSeek—a hardware development kit designated exclusively for internal testing—along with extra manual configurations not publicly distributed; the README lists a commercial HDK planned for release in mid-October 2026 as the subsequent deployment baseline, a date subject to the vendor’s roadmap. The complete metric specification is: in self-tests on 950DT with the designated PoC configuration, DeepEP-Ascend reports an FP8 dispatch bandwidth of 373–375 GB/s and a BF16 combine bandwidth of 345–347 GB/s for EP8; tests utilized netlayer 1, representing the Clos network external to the super node; timing includes issue and drain, excluding final epilogue. These numbers reflect the network layer external to the super node; high-speed interconnect within the node is another matter entirely.
End-to-end inference throughput represents yet another metric: Huawei reported that under its offline test setup with EP32, 128K context, and a Dspark speculative acceptance rate of 0.85, V4.1-Flash achieved 2,469 tok/s/card (TPOT 5 ms); tests explicitly excluded the impact of serving scheduling and framework load balancing, and our verification obtained no independent replication. Under the same offline test environment, throughput corresponding to a TPOT of 10 ms was reported at 5,102 tok/s/card.
Beyond these numbers lies a sobering fact: to date, there has been zero third-party verification. In the days following code release, all PRs submitted by external developers were labeled by the authors themselves as ‘not tested on actual NPUs’; there is no public evidence showing any third-party team has fully executed this software stack on an independent 950 cluster. Issue trackers across repositories also document early rough edges: the communication library launched without a LICENSE file, an issue flagged within 40 minutes and still unresolved (Issue #1); the matrix multiplication library is deeply coupled to mechanisms unique to the 950, which the vLLM-Ascend team found in testing would not run on previous-generation Atlas A3 hardware (Issue #3); and on the Top-k component, third parties raised priority disputes over algorithmic originality and citations (Issue #13). Regarding training, official announcements noted that operators used during training have corresponding high-performance implementations on Ascend, while also stating that ‘the TileLang route was first validated on NVIDIA’s mature platform’; to date, no public evidence indicates that related models have completed end-to-end full-scale pre-training on Ascend clusters.
Having surveyed the costs, the numbers, and the blank spots, we return to the opening question. What this batch of components truly delivers is a line-by-line auditable boundary: what must remain compatible, what can be left open, and what each is worth. DeepSeek’s Ascend components demonstrate a pragmatic route: preserve model operation interfaces and semantics, share development infrastructure, and allow critical kernels to be rewritten per hardware. What it unveils is an engineering boundary for layered reuse, not an accomplished breakthrough in universal performance portability. The key to breaking CUDA dependency is deciding what does not need to be compatible.