DeepSeek published two new Ascend libraries on September 30 with a disclosure baked into the same README as the numbers: the benchmarks ran on hardware the public cannot get.
Every bandwidth figure below is published by DeepSeek in its own README, and DeepSeek states there that the measurements were taken on a proof-of-concept hardware kit supplied to it, with manual configuration, which is not a publicly distributed release. This publication has run no benchmark and reproduced nothing. Nothing here says Nvidia is being replaced, and nothing here says DeepSeek trains on Ascend. Neither claim is supported by anything read for this piece. Nothing here is investment advice.
What Happened
On September 29–30, 2026, DeepSeek published two new repositories — DeepGEMM-Ascend and DeepEP-Ascend — alongside Ascend-targeted updates pushed to three projects that already existed: TileKernels (created April 2026), DeepSelect (created September 2026), and FlashMLA (created February 2025). This is two new libraries, not six new projects; the GitHub API creation dates confirm the distinction. DeepSeek’s WeChat post of the same day, as reported by China Securities Journal, said every TileLang operator used in DeepSeek’s training has a corresponding high-performance Ascend implementation, and that the company hopes the work sets an example for building usable software ecosystems on more AI chips.
The DeepEP-Ascend README publishes expert-parallelism dispatch and combine bandwidth figures measured on Ascend 950DT with CANN 9.2.0. At EP8, dispatch runs 373–375 GB/s and combine 345–347 GB/s. At EP16, dispatch is 348–352 and combine 338–341. At EP32, 335–340 and 320–324. At EP64, 323–327 and 294–298. At EP128, dispatch falls to 313–320 and combine to 272–278. Those figures were collected using a configuration specific to a proof-of-concept HDK supplied to DeepSeek — a hardware kit that is not a publicly distributed release. DeepSeek says this in the same README that carries the numbers.
The commercial firmware release for Atlas 850E is described as planned for around mid-October 2026 — approximately October 15 — with the README noting that this is a vendor release plan and availability is subject to Huawei’s publication schedule. A third caveat appears alongside both of these: kernel support on other Ascend generations or CANN versions has not been established by these measurements.
The key insight: DeepSeek published the limits of its own benchmark inside the same file as the benchmark. A vendor that does that is not overclaiming — it is doing the opposite. The caveats quoted throughout this piece are DeepSeek’s own words, not this publication’s reservations. The discipline with which the scoping is stated is itself the signal worth reading.

The Structural Read
The Map of AI separates infrastructure layers cleanly: silicon sits at the bottom, firmware and driver software above it, then the kernel and operator layer, then training frameworks, then models. What DeepSeek shipped on September 30 lives entirely in the kernel and operator layer. It is a software-enablement story at a specific altitude — and understanding the altitude tells you what the release can and cannot do.
The API claim is the strong one, and it is checkable. DeepGEMM-Ascend describes itself as fully API-compatible with DeepGEMM and uses the same Python package name. DeepEP-Ascend says its public buffer APIs are aligned with the Nvidia version of DeepEP and likewise keeps the package name. That is the actual portability mechanism: code written against the Nvidia stack can target Ascend without rewriting the calling code. That compounds over time because it lowers the marginal cost of every future operator that gets added to the stack.
The performance claim is scoped far more narrowly than the framing typically applied to it. One chip generation, one CANN version, a firmware kit supplied to DeepSeek and not publicly distributed, a commercial release only planned for mid-October, and an explicit statement that support on other Ascend generations has not been established. That is not a failure — it is an honest baseline. The EP8 dispatch figure is the headline number, and it is also the one furthest from how large mixture-of-experts models are run.
DeepEP-Ascend README — 30 September 2026
“The performance results in this README were collected using a PoC HDK supplied to DeepSeek, with additional manual configuration. That configuration is not a publicly distributed release.”
Expert parallelism is the axis a large mixture-of-experts model scales along, and the published table shows what happens along that axis. The README itself says sustained dispatch reaches roughly 90–95 percent of the physical payload bandwidth limit for EP sizes up to 32, and that larger EP sizes and combine remain under optimisation — combine carries additional local reduction overhead and memory contention. By EP128, dispatch has dropped to 313–320 GB/s and combine to 272–278 GB/s, roughly 15 percent and 20 percent below their EP8 midpoints.
The regime where the numbers are strongest is not the regime where the largest models run. That is a description of the table as published.
There are also real gaps named in the Ongoing section of the README that matter for anyone assessing production readiness: Ascend reduce-scatter and all-reduce kernels are still being built; expert load-balancing communication kernels are not implemented; pipeline parallelism and Engram are experimental; hybrid communication, CPU-backed Engram storage, and graph capture are unsupported. And one thing inside the API-compatible library is not identical to the Nvidia version — the README says so directly: the scaling factor format on Ascend differs, with each pair of UE8M0 scaling factors along K packed into an int16 and stored in MN-major order. That is a surface detail but it matters for anyone porting weight-handling code verbatim.
Map of AI — Layer Read
Where This Release Sits in the Stack
DeepSeek’s September 30 work occupies the kernel and operator layer — above silicon and firmware, below training frameworks and models. API parity at this layer is a structural lever: it means the layer above (training code) does not need to change to address the layer below (Ascend silicon). Performance parity at this layer is a separate claim, scoped to one chip, one CANN version, and hardware not yet publicly distributed. Both claims are real. They are not the same claim.
Three Implications
PORTABILITY IS THE DURABLE ASSET
API compatibility at the kernel layer means training code written against the Nvidia stack can target Ascend without being rewritten. That is not a performance claim — it is a switching-cost reduction. Every operator added to the stack in future compounds this value, because the calling code stays the same. The portability mechanism is already in place; the question is how many operators fill it out.
THE BENCHMARK CEILING IS NOT THE PRODUCTION FLOOR
EP8 dispatch of 373–375 GB/s is the headline figure. But real mixture-of-experts deployments run at higher expert-parallelism degrees, and the table shows a monotonic decline across every step from EP8 to EP128. The README attributes this to local reduction overhead and memory contention at scale, and says larger EP sizes remain under optimisation. Any performance assessment should use the EP range relevant to the target model, not the EP8 peak.
THE FIRMWARE GATE IS STILL CLOSED
The commercial firmware release for Atlas 850E is a vendor plan for approximately mid-October 2026, subject to Huawei’s publication schedule. Until that releases, the benchmark conditions are not reproducible outside of DeepSeek’s own environment. The software layer shipped on September 30; the hardware enablement layer it depends on has not. That gap is worth tracking, not dismissing — it is the difference between what is published and what is usable.
The Bottom Line
DeepSeek shipped a real piece of engineering on September 30: API-compatible kernel libraries that let training code address Ascend silicon without being rewritten. That is the claim that holds. The bandwidth numbers — measured on a PoC hardware kit supplied to DeepSeek, on one chip generation, under one CANN version, with commercial firmware not yet publicly available — are a baseline for a specific configuration, not a production benchmark, and DeepSeek’s own README makes every one of those limits explicit.
Two new repositories exist. The firmware gate is still closed. The performance curve degrades at the EP ranges where large models actually run. All of that is in the source. Read the source.
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
Every bandwidth figure above is published by DeepSeek in the DeepEP-Ascend README, read directly on 30 September 2026. DeepSeek states in that same README that the measurements were collected using a proof-of-concept hardware kit supplied to it, with additional manual configuration, and that this configuration is not a publicly distributed release. It also states that users on earlier proof-of-concept versions may observe lower bandwidth, and that kernel support on other Ascend generations or CANN versions has not been established by these measurements.
This publication has run no benchmark and has reproduced none of these results. The mid-October 2026 date for Huawei’s commercial Atlas 850E firmware release is described in the README as a vendor release plan subject to Huawei’s publication schedule. It is not a commitment and nothing above treats it as one. Nothing above claims that Nvidia is being replaced or displaced, or that the CUDA software advantage has ended.
No source read here supports any such claim, and this is a software-enablement story rather than a hardware-share one. Nothing above claims DeepSeek trains production models on Ascend. The China Securities Journal wire, reporting DeepSeek’s WeChat post of 30 September, says that every TileLang operator used in DeepSeek’s training has a corresponding high-performance implementation on Ascend. That is a statement about operator coverage and not about where training runs.
The WeChat post itself was not read for this piece. Repository creation dates and star counts come from the GitHub API on 30 September 2026. No DeepGEMM-Ascend performance figure appears above because none was verified. Nothing above predicts anything, and nothing here is investment advice.
Sources: github.com · openrouter.ai · fourweekmba.com · github.com · api.github.com









