A Huawei-led consortium’s July 22 technical report documents full-parameter post-training of DeepSeek’s V4 family on Huawei Ascend chips, supplying efficiency numbers that were absent when the claim surfaced in June. The team’s own measurement puts the run at 34.22% model FLOPs utilization, or MFU. Independent experts say Nvidia hardware likely still did the heavy lifting and read the result as evidence for post-training and inference rather than proof that Huawei can replace Nvidia for frontier training.
What Changed
- A Huawei-led consortium's July 22 arXiv report documents full-parameter post-training of DeepSeek's V4 family on Huawei Ascend chips and reports 34.22% model FLOPs utilization, a 2.93-fold efficiency gain over the open-source baseline recipe.
- The paper covers post-training only and does not establish which hardware trained the 1.6-trillion-parameter V4-Pro originally; Tsinghua's Liu Zhiyuan told MIT Technology Review the model may still have been trained mainly on Nvidia hardware.
- A June account sourced to the Shenzhen municipal government put the cluster at at least 1,000 Ascend 910C chips, a part that returned roughly 60% of an Nvidia H100's inference performance in earlier DeepSeek testing.
- The report lands as Washington investigates Chinese firms' chip access and Beijing weighs its own export controls on AI models, weeks before Xi Jinping's planned Washington visit.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
The July 22 report
Wednesday’s technical paper documents end-to-end optimization of full-parameter post-training for the V4 family, including V4-Flash and the 1.6-trillion-parameter V4-Pro, on Huawei’s Ascend NPU SuperPOD.
The authors’ own measurements compare the optimized Ascend system with an open-source baseline recipe and put the efficiency improvement at 2.93-fold. They also state that training remained stable.
But the paper covers post-training only and does not establish which hardware was used for V4’s original pre-training.
The gap in the June 6 claim
Tom’s Hardware wrote that the original announcement “carries no benchmarks” and supplied no run duration, direct Nvidia comparison or measure of cluster efficiency. The publication called it part of “a series of dubious claims” and noted that DeepSeek had not commented. Wednesday’s paper adds a utilization rate and a baseline comparison, both measured by the Huawei-led team.
According to the June account, the group used at least 1,000 Ascend 910C chips. That figure came from the Shenzhen municipal government through the South China Morning Post, rather than from the report itself. Tom’s Hardware described the 910C as a dual-die accelerator using SMIC’s second-generation 7-nanometer process and Huawei’s Da Vinci architecture. It said earlier DeepSeek testing returned roughly 60% of an Nvidia H100’s inference performance, a figure that applies only to inference.
Assessments of Nvidia’s role
Tsinghua computer-science professor Liu Zhiyuan told MIT Technology Review that DeepSeek “appears to have adapted only part of V4’s training process for Chinese chips” and may still have trained the model mainly on Nvidia hardware. Multiple anonymous sources told the publication that Chinese processors remain better suited to inference than to training.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
In one provenance dispute, a senior Trump administration official alleged that V4 was trained on smuggled Blackwell chips; DeepSeek denied the account and cited H800 GPUs and Ascend 910Cs, while Nvidia called the claim “far-fetched.”
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
The DeepSeek-Huawei collaboration indicated the models could achieve similar performance on Huawei and Nvidia systems, Omdia chief analyst Lian Jye Su said in April, as cited by Reuters. The May 24 aiproem analysis described V4’s strongest implication as applying to inference and stated that the model does not prove Huawei can replace Nvidia for frontier training.
V4 had been validated on Nvidia GPUs and Huawei Ascend processors, Thechinaacademy reported. The publication also wrote that DeepSeek refused Nvidia and AMD early access to optimize for V4, yet the model ran on CUDA at launch. Nvidia published a same-day demonstration of V4-Pro on Blackwell hardware at more than 150 tokens per second per user.
Export controls before Xi’s visit
China was considering tighter export controls on AI models and chips, Reuters reported on July 21, citing the Financial Times. On July 23, Semafor wrote, citing The Information, that the United States was investigating Chinese AI companies’ access to advanced processors and had warned that sanctions or blacklisting could follow. China’s Ministry of Commerce had met Alibaba, ByteDance and Z.ai about restricting overseas access to domestic models, including unreleased systems, a mid-July Reuters analysis said.
The U.S. Commerce Department’s June order barring foreign nationals from Anthropic’s Fable and Mythos models is the template Beijing is now weighing for its own models, according to that analysis. Those U.S. investigations, Semafor reported, come weeks before Xi Jinping’s planned visit to Washington.
Frequently Asked Questions
What did the July 22 report show?
A Huawei-led team documented full-parameter post-training of DeepSeek's V4 family on Huawei's Ascend NPU SuperPOD, reporting 34.22% model FLOPs utilization and a 2.93-fold efficiency improvement over the open-source baseline recipe. The figures are the team's own measurements.
Does this prove China trained a frontier model without Nvidia?
No. The paper covers post-training only and does not establish which hardware trained V4 originally. Tsinghua professor Liu Zhiyuan told MIT Technology Review the model may still have been trained mainly on Nvidia hardware, and anonymous sources said Chinese chips remain better suited to inference than training.
How many Huawei chips were used?
A June account sourced to the Shenzhen municipal government through the South China Morning Post said the group used at least 1,000 Ascend 910C chips. That figure came from the local government, not from the report itself.
How does the Ascend 910C compare with Nvidia?
Tom's Hardware described the 910C as a dual-die accelerator built on SMIC's second-generation 7-nanometer process, and said earlier DeepSeek testing returned roughly 60% of an Nvidia H100's inference performance, a figure that applies only to inference.
Why does the timing matter?
It arrives as the U.S. investigates Chinese firms' access to advanced processors and Beijing weighs its own controls on AI models, mirroring a June U.S. order barring foreign nationals from Anthropic's Fable and Mythos models, weeks before Xi Jinping's planned Washington visit.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR