DeepGEMM Ascend runs DeepGEMM kernels on Huawei NPUs
DeepSeek's DeepGEMM Ascend keeps the DeepGEMM API on Huawei Ascend NPUs, with dense GEMM reaching up to 99.8% of the hardware limit.
Published · on github.com · 2 min read

DeepSeek has published DeepGEMM Ascend, a port of its DeepGEMM matrix multiplication library to Huawei's Ascend NPU platform, on github.com. The repository states that the port is fully API-compatible with DeepGEMM and supports BF16, FP8 and FP4 GEMM, MQA logits and MegaMoE. On Ascend, the same package name, the same APIs and the same development workflow apply as on the platforms DeepGEMM already supported.
The library is a lightweight abstraction over the Ascend MAD (matrix multiply-add) primitives. That layer hides fractal layouts, alignment constraints, address calculations and the verbose low-level parameters that usually make GEMM kernels hard to write, so the kernels stay concise. The implementations use Ascend-specific techniques, including sparse data loading and coroutine-based pipelining, to approach the hardware's performance limits, and the repository presents them as references for extreme performance optimisation on the platform.
Dense GEMM reaches 99.8% of the hardware limit
The published benchmarks were measured on an Ascend 950DT with the CANN 9.20 toolkit, using bench_msprof with cold L2. Shapes follow the DeepGEMM test suite and cover inference and training workloads from the DeepSeek model series. Dense GEMM reaches up to 99.8% of the hardware limit across types: BF16×BF16 at 4096×7168×16384 runs in 2229.9 microseconds at 431 TFLOPS against a 432 TFLOPS limit.
Other kernels are measured separately. The MQA logits kernel for the DeepSeek Lightning Indexer is described as FIX-pipe bound rather than compute bound, saturating that pipe at 99% utilisation. MegaMoE fuses expert-parallel dispatch, two grouped GEMMs, SwiGLU and combine, benchmarked over EP8 with top-k=6 and one shared expert, averaged across eight ranks. The HC prenorm GEMM for the mHC module nearly saturates HBM write bandwidth.
What the release requires
The initial release supports Ascend 950 devices. Requirements listed are a HUAWEI Ascend NPU, the CANN 9.20 toolkit providing bin/bisheng and bin/ld.lld, the torch_npu package, Python 3.10 or higher, and compilers with C++20 format support. The scaling factor layout differs from NVIDIA's: each pair of UE8M0 scaling factors along the K dimension is packed into an int16 and stored in MN-major order.
The repository is released under the MIT licence. It credits the upstream DeepGEMM project, CUTLASS, DeepJIT and Tilelang, and acknowledges Huawei for technical support during development. The concrete figure to hold is 99.8% of the hardware limit on BF16 dense GEMM, measured on Ascend 950DT.
Source: github.com — BARGO’s commentary on the linked source.