Z.ai said Wednesday that the week-long anonymous preview of GLM-5.3-Flash ran entirely on domestically developed Chinese chips and reached per-token costs comparable to mainstream Nvidia GPUs. A custom serving stack handled the preview across tens of thousands of domestic accelerators. The serving result does not establish that the same hardware can complete a frontier-model training run.
What Changed
- Z.ai says the week-long anonymous GLM-5.3-Flash preview ran across tens of thousands of domestic Chinese accelerators.
- The company reported per-token cost comparable to mainstream Nvidia GPUs, but named no chip vendor and published no throughput, power or utilization figures.
- Serving is not training. DeepSeek's R2 run never completed on Ascend hardware, even after Huawei sent engineers on site.
- Supply is the tighter constraint. SemiAnalysis puts CXMT's 2026 HBM output at enough for only 250,000 to 300,000 Ascend 910C-equivalent packages.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
A serving stack built around memory
The technical disclosure describes a custom inference engine built on SGLang for chips constrained mainly by memory capacity and bandwidth. The pressure increases when prompts approach the model’s one-million-token context limit, disclosed at the Aug. 26 launch.
Z.ai split multimodal encoding, prompt prefill and token-by-token decoding into separately scheduled worker pools. That Encode-Prefill-Decode design lets operators add capacity where a workload is backing up instead of asking every accelerator to handle the full sequence.
Other measures compressed model weights and the key-value cache so a memory-limited chip could hold more of the workload. The company used that arrangement throughout the preview disclosed at the same launch.
The model released that day contains 320 billion parameters but activates 18 billion for each token. Against GLM-5.3, Z.ai measured three times less attention computation and a 4.4-fold smaller key-value cache at launch.
A GLM-5.3 infrastructure agent also helped engineers develop kernels and diagnose bottlenecks, “creating a feedback loop in which the model helped optimize the system serving the model itself.”
What the claim leaves out
Z.ai has not named the chip model or vendor. Huawei, the largest domestic supplier, shipped 812,000 AI accelerators in 2025 for a 20.3% share of China’s roughly four-million-unit market, compared with Nvidia’s 2.2 million units and 55%. Z.ai published no power consumption, exact throughput, utilization rate or normalized Nvidia comparison. The cost-parity assertion cannot be checked from the material released Wednesday, and none of the serving results has been independently audited.
Served is not trained. Z.ai says the preview and subsequent inference ran on Chinese accelerators. It does not say GLM-5.3-Flash was trained entirely on them.
FREE WEEKDAY MORNING BRIEFING
Don’t miss the next AI story that matters.
The Implicator Morning Briefing filters the AI news cycle to the stories worth your attention and explains their consequences. From San Francisco, every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.
About five minutes. No hype. No spam.
Huawei’s Ascend processors can run inference workloads, yet DeepSeek’s R2 training effort failed to finish reliably on them. Even after Huawei sent its own engineers on site, the team could not complete a successful full training run on Ascend hardware, pushing R2’s launch back by months. DeepSeek switched back to Nvidia chips for training and kept Ascend hardware for inference.
The price attached to the claim
On Aug. 26, 2026, Z.ai priced standard GLM-5.3-Flash access at $0.15 per million input tokens, $0.03 per million cached input tokens and $0.50 per million output tokens. A launch discount halves those rates through Sept. 9, 2026. At the same launch, GLM-5.3 cost $1.40 for input and $4.40 for output per million tokens.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
In its launch materials, Z.ai reported an Artificial Analysis Intelligence Index version 4.1.1 score of 57 at $0.045 per task on the discounted tier. The figure had not yet appeared on the Artificial Analysis leaderboard. Z.ai said that score had previously cost about ten times more.
Supply is the harder constraint
TrendForce projects domestic accelerators will take nearly 90% of China’s high-end market in 2026, up from 45% in 2025. Holding the market near its 2025 base of four million units would require Chinese suppliers to replace about 1.96 million foreign accelerators in 2026 and increase their own output 2.2-fold from 2025.
High-bandwidth memory and advanced packaging capacity remain the choke points. Chinese firms had stockpiled roughly 13 million HBM stacks from Samsung before export controls tightened. SemiAnalysis estimates CXMT’s 2026 output at about two million stacks, enough for only 250,000 to 300,000 Ascend 910C-equivalent packages. That ceiling holds regardless of how much logic-die capacity SMIC has.
Because the weights are public under an MIT license, outside developers can now download GLM-5.3-Flash and measure its serving efficiency themselves.
Frequently Asked Questions
What did Z.ai actually claim about Chinese chips?
That the week-long anonymous preview of GLM-5.3-Flash ran entirely on domestically developed Chinese accelerators, across tens of thousands of them, at per-token cost comparable to mainstream Nvidia GPUs.
Which Chinese chips were used?
Z.ai has not said. It named no chip model or vendor. Huawei is the largest domestic supplier, shipping 812,000 AI accelerators in 2025 for a 20.3% share of China's roughly four-million-unit market, against Nvidia's 2.2 million units and 55%.
Was GLM-5.3-Flash trained on Chinese chips too?
Z.ai says the preview and subsequent inference ran on Chinese accelerators. It does not say the model was trained entirely on them. DeepSeek's R2 training effort failed to finish on Huawei Ascend hardware and switched back to Nvidia chips for training.
What does GLM-5.3-Flash cost?
On Aug. 26, 2026, standard access was $0.15 per million input tokens, $0.03 per million cached input tokens and $0.50 per million output tokens. A launch discount halves those rates through Sept. 9, 2026. GLM-5.3 cost $1.40 for input and $4.40 for output per million.
Has anyone verified the serving claims?
Not yet. None of the serving results has been independently audited, and the Artificial Analysis Intelligence Index score of 57 came from Z.ai's own launch materials rather than the Artificial Analysis leaderboard. The weights are public under an MIT license, so outside developers can now measure it themselves.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR