Chips & Infrastructure · · Team LATENT
DeepSeek releases software for AI computation and communication on Huawei chips
Running AI on chips other than NVIDIA's requires adapting computation, memory layouts, and communication between chips. The releases of DeepGEMM-Ascend and DeepEP-Ascend explain why, and clarify the scope of API compatibility. The performance figures are the developer's measurements, and the communication benchmarks used a PoC HDK that has not been publicly distributed.
On September 30, 2026, DeepSeek made the initial release of DeepGEMM-Ascend for Huawei's Ascend 950 AI chips. The software is designed to speed up the large volume of matrix calculations AI performs on Ascend. DeepEP-Ascend, which handles communication when AI runs across multiple chips, is also available. DeepGEMM release information · DeepEP overview
The release illustrates why changing AI chips also requires reworking computation and data movement. Putting the same model on a chip other than NVIDIA's will not deliver its performance potential if the chip spends its time waiting for the numbers it needs. DeepSeek's implementation shows what remains shared and what has been adapted to Ascend.
Contents
- Matrix operations need computation tailored to the chip
- The same numbers need a different layout in memory
- MoE requires communication between chips before and after computation
- Matching APIs can still leave differences in supported features
- Communication performance was measured on a setup not yet publicly distributed
- Deployment decisions still require testing the whole model
Matrix operations need computation tailored to the chip
AI processes text by converting it into sequences of numbers. It repeatedly multiplies and adds those numbers with values learned during training, known as weights. A table of numbers arranged in rows and columns is a matrix, and the operation that multiplies matrices is called GEMM. DeepGEMM is a library that provides these operations for other software to call.
For example, when processing many texts at once, matrices are divided into smaller blocks and passed to the chip's compute units. If the block sizes and execution order do not suit the chip's architecture, waiting time increases even when compute capacity is available. A small program that performs such operations is called a kernel.
DeepGEMM-Ascend uses Ascend's matrix multiply-and-add instructions. It supports the BF16, FP8, and FP4 number formats and incorporates Ascend-specific optimizations, such as overlapping data loading with computation. The 16, 8, and 4 in the names indicate how many bits each format uses to represent a number. Fewer bits make the data smaller, but also constrain the values and precision that can be represented. DeepGEMM overview and supported formats
Tools for writing these computations are also needed. Some operations in DeepGEMM-Ascend use TileLang, a language for developing kernels. TileLang's Ascend implementation organizes data copies, execution order, and synchronization to suit the chip. Adapting the software therefore also involves the layer that translates the computation specified by a developer into actual instructions. DeepGEMM dependencies · TileLang's Ascend documentation
The same numbers need a different layout in memory
Faster matrix computation also calls for careful handling of memory reads. Storing the required numbers in groups the chip can read efficiently makes it easier to feed its compute units. Even when a matrix has the same dimensions, the order in which its values are stored in memory is a separate concern.
DeepGEMM-Ascend has a concrete difference from the NVIDIA implementation: how it stores scaling factors, the auxiliary data that represents the scale of low-precision values. On Ascend, pairs of scaling factors are packed into 16-bit integers and stored in a prescribed order. The README explicitly notes this difference and provides functions for converting the factors to the required layout. Scaling factor layouts and conversion functions
A shared way to call the software does not necessarily mean input preparation is identical. This example pinpoints something developers need to check when changing chips. Beyond matching function names, they must also adapt the format and layout of the data passed to those functions.
MoE requires communication between chips before and after computation
MoE is an approach in which a model contains multiple computational components and selects which ones to use for each input. These components are called experts. The model selects experts for each token, a unit used to process text, and sends the numerical data to the experts it needs.
When experts are spread across multiple chips, communication is needed whenever a selected expert resides on another chip. This arrangement is called expert parallelism, or EP. Sending data to its destinations is called dispatch, while gathering and merging the results is called combine. DeepEP-Ascend handles this two-way communication. DeepEP features and usage examples
| MoE stage | What the software does |
|---|---|
| Distribute the inputs | Send data to the chips hosting the selected experts |
| Compute the matrices | Arrange the received data appropriately and perform each expert's computation |
| Gather the results | Return the computation results to their originating participants and merge them |
Even a chip that computes quickly cannot start an expert's computation until its inputs arrive. Collecting the results also takes time. Both the computation optimizations handled by DeepGEMM and the transfers between chips handled by DeepEP therefore need to be in place.
Matching APIs can still leave differences in supported features
An API is a set of rules for calling functionality from another program. DeepSeek describes DeepGEMM-Ascend as fully compatible with DeepGEMM's API, using the same package name and development workflow. That claim, however, applies to the DeepGEMM API. It is not an announcement that the entire CUDA development platform for NVIDIA GPUs has been replaced. DeepGEMM's compatibility statement
DeepEP-Ascend also aligns its APIs with the NVIDIA version of DeepEP, but it has Ascend-specific operating requirements. In some cases, an API is present even though the underlying communication operation has not yet been implemented. DeepEP's API support table
Bucket functionality, which gathers or aggregates data across chips, is experimental. The all-gather operation is available, but reduce-scatter and all-reduce, which aggregate values, are still under development. PP, which divides computation stages across chips, and Engram, which accesses remote memory, are also experimental, with correctness and performance validation ongoing. Some APIs for balancing the load across experts still lack Ascend communication kernels. Work in progress
Aligning APIs helps make existing code easier to port. Whether the required features are implemented, whether they produce the same results, and whether they are fast enough each require separate checks. API compatibility alone does not establish equivalence in every feature or in performance.
Communication performance was measured on a setup not yet publicly distributed
The DeepGEMM-Ascend README reports matrix computation reaching up to 99.8% of the hardware limit on Ascend 950DT. This figure comes from a particular BF16 matrix size. The measurement used CANN 9.20 and a dedicated benchmarking function, with the data not already present in the L2 cache. It does not mean an entire model runs at 99.8% of theoretical performance. DeepGEMM benchmark conditions and results
DeepEP-Ascend publishes results with the number of ranks, the participants in communication, ranging from 8 to 128. EP8 distributes the work across 8 ranks, while EP128 uses 128 ranks. Bandwidth measures the amount of data that can be transferred per second.
| EP size | Input dispatch bandwidth | Result combine bandwidth |
|---|---|---|
| EP8 | 373〜375 GB/s | 345〜347 GB/s |
| EP16 | 348〜352 GB/s | 338〜341 GB/s |
| EP32 | 335〜340 GB/s | 320〜324 GB/s |
| EP64 | 323〜327 GB/s | 294〜298 GB/s |
| EP128 | 313〜320 GB/s | 272〜278 GB/s |
Source: DeepEP's performance table. The ranges cover all participating ranks. These are not comparisons with NVIDIA chips.
The measurements used Ascend 950DT and the CANN 9.2.0 software stack. Each token was routed to 6 of 256 experts, with FP8 used for dispatch and BF16 for combine. Capacity was 16,384 tokens per rank, with a numerical vector width of 7,168, using the supernode's external Clos network. Each rank ran 10 warmups and 50 measurement samples with cache flushing. Timings include issuing the transfers and waiting for them to complete, but exclude the final post-processing.
These communication measurements used a proof-of-concept PoC HDK supplied to DeepSeek, along with manual configuration. An HDK is the collection of software that operates the hardware, including drivers and firmware. The measured configuration has not been publicly distributed, so installing the published code alone does not necessarily reproduce the same bandwidth. HDK and firmware requirements
The README recommends Huawei's Q3 commercial HDK for Atlas 850E. It says public availability is planned for mid-October 2026, around October 15. This is Huawei's release plan, not a guaranteed date. The figures above were not measured on that unreleased commercial version.
Deployment decisions still require testing the whole model
The release makes the code for computation and communication on Ascend available alongside concrete details of supported features and measurement conditions. Developers seeking to run AI on chips other than NVIDIA's can now check whether the operations their models need are covered.
Before deployment, developers need to check the functions and data formats their model uses and align the versions of drivers and development tools. They then need to measure response quality, processing time, and the number of concurrent requests the model can handle as a whole. These documents alone do not establish whether quality or operating costs will match an NVIDIA setup.
For users, this also provides a way to assess news that AI has run on a different chip. Looking beyond support for computation to memory layouts, communication between chips, and a publicly available runtime environment makes it easier to judge how far practical deployment has progressed.
The article is dated October 2, 2026, which is also the date the sources were checked. Performance figures are the developer's measurements from DeepSeek's public materials; LATENT has not reproduced them on hardware.