Huawei's Ascend AI-chips outperform Nvidia processors in running DeepSeek R1

Huawei Technologies’ advanced data center architecture CloudMatrix 384 has enabled the company’s Ascend chips to surpass the performance of Nvidia’s H800 graphics processing units (GPUs) in running DeepSeek’s R1 artificial intelligence (AI) model, according to a technical paper jointly written by researchers from Huawei and Chinese AI infrastructure start-up SiliconFlow. The paper described CloudMatrix 384 as a specialized “AI supernode” that is purpose-built for handling extensive AI workloads. Huawei expected CloudMatrix “to reshape the foundation of AI infrastructure”, according to the paper. It consists of 384 Ascend 910C neural processing units (NPUs) and 192 Kunpeng server central processing units, which are interconnected through a unified bus providing ultra-high bandwidth and low latency. The advanced large language model (LLM) serving solution, dubbed CloudMatrix-Infer, leverages that infrastructure. It surpassed the performance of some of the world’s most prominent systems in running DeepSeek’s 671-billion-parameter R1 reasoning model.

The architecture reflects Huawei’s efforts to overcome U.S. tech control measures and sanctions, as the company pushes the boundaries of AI system performance. Data centers are facilities that house large-capacity servers and data-storage systems, with multiple power sources and high-bandwidth internet connections. In a so-called prefill phase involving the initial processing of prompts, CloudMatrix-Infer reached a throughput of 6,688 tokens per second per NPU for a 4,000-token prompt length. That equates to a computational efficiency of 4.45 tokens per second per trillion floating-point operations per second (TFLOPs). Tokens are the basic units that LLMs – the technology behind generative AI services like ChatGPT – use to process text. Token length directly impacts cost, processing time and an AI model’s ability to understand and respond to complex instructions or narratives. TFLOPS is a measure of a computer’s processing speed – specifically, its ability to perform complex calculations in tasks such as training AI systems.

In the subsequent decode phase that generates output from an AI model, the paper’s findings showed that CloudMatrix recorded 1,943 tokens per second per NPU for a 4,000-length key-value cache – a memory structure that enables more efficient use of AI processors. The same phase showed output generation times consistently below 50 milliseconds per token, yielding an efficiency of 1.29 tokens per second per TFLOPS. According to the paper, those metrics surpassed the performance of Nvidia’s SGLang fast-serving framework for LLMs based on the U.S. firm’s flagship H100 GPU and another system running DeepSeek’s R1 using the H800 processor. The paper marked the first time that Shenzhen-based Huawei formally provided details about the capabilities of its flagship Ascend 910C AI accelerator, the South China Morning Post reports.