DeepSeek’s September release narrowed the US lead over China to roughly 3% in a LiveBench score comparison. The V4.1 Flash model moved close to Anthropic’s leading system on LiveBench’s overall tests, which include reasoning and coding. Bloomberg Intelligence senior analyst Robert Lea expects the improved performance to bring Chinese developers further market share gains.

In the leaderboard captured on October 4, DeepSeek V4.1 Flash Max Effort scored 81.1 against 83.4 for Anthropic’s Claude Fable 5.1 Max Effort. That is a difference of 2.3 score points, or about 2.8% of Anthropic’s score. Lea’s comparison put the gap at about 9% in May and 15% earlier in 2026.

The scores differ by task. In the October 4 snapshot, DeepSeek scored 77.3 on agentic coding, compared with 66.1 for Claude Fable 5.1 Max Effort.

The US still holds the lead in this comparison. These benchmark results do not establish national AI leadership or show whether US chip export restrictions are effective.

What Changed

AI-generated summary, reviewed by an editor. More on our AI guidelines.

DeepSeek’s efficiency changes

DeepSeek’s September 10 release notes describe a design intended to cut the computing and memory needed to serve answers. Only a small portion of the model’s stored parameters, the values learned during training, is active when it reads input or generates output. The company says it also shrank the cache that holds information used during a conversation.

Those efficiency claims come from DeepSeek. The release supports native visual understanding and continues to offer off-peak API rates at half the peak price, under the new pricing that took effect on September 10. The benchmark gain does not by itself establish how reliably the model will perform inside a company or whether it can be deployed securely.

OpenRouter usage and enterprise demand

Chinese models processed more tokens than US models on OpenRouter in the trailing week to September 7. But US models handled slightly more requests in that period, the usage window examined in JPMorgan Asset Management’s September 9 analysis. OpenRouter is a service developers use to send requests to different AI systems.

Token share is not market share. Automated agents can produce large volumes of tokens while working through a task, so heavy token consumption can make a provider look more widely used than a count of requests does.

OpenRouter also underrepresents enterprise activity. Most of those workloads run through major cloud platforms or directly with developers such as OpenAI and Anthropic. Large US companies still predominantly use US models, while large Chinese companies tend to use domestic ones.

Lower token prices do not always mean cheaper completed tasks. Models consume different amounts of text before reaching an answer, and downloadable models still incur serving costs. A price comparison based on tokens alone can therefore miss the cost of completing the work.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

Enterprise procurement also depends on regulation and data governance. Downloading model weights permits private hosting, but does not reveal the training data or resolve intellectual property questions.

Profitability remains uncertain

Lea forecasts that China’s AI industry could remain unprofitable until 2030 despite the benchmark gains. In his assessment, low-margin token supply and a domestic price war may prevent developers from building a lasting commercial advantage.

China’s market is now crowded with more than 1,100 large language models.

Lea identifies ByteDance’s Doubao as the frontrunner in AI app monetization, while rival chatbots from DeepSeek and Tencent remain free.

“Putting China’s AI sector on a sustainable profit footing will require a cooling of competitive pressures, an industry shakeout, and a more rational approach to pricing,” he said.

Frequently Asked Questions

How close is DeepSeek to the US leader on LiveBench?

The October 4 snapshot showed DeepSeek V4.1 Flash Max Effort at 81.1 overall and Anthropic’s Claude Fable 5.1 Max Effort at 83.4. The 2.3-point difference is about 2.8% of Anthropic’s score, rounded to roughly 3% in the comparison.

Does DeepSeek lead on every task?

No. DeepSeek scored 77.3 on agentic coding against 66.1 for Claude Fable 5.1 Max Effort in the October 4 snapshot. Anthropic held the higher overall score.

What changed in DeepSeek’s September release?

DeepSeek described a design intended to reduce the computing and memory needed to serve answers, using only a small part of its stored parameters at a time and a smaller conversation cache. These efficiency claims come from the company.

Does Chinese token traffic prove market leadership?

No. JPMorgan’s September 9 analysis found more Chinese-origin token traffic on OpenRouter but slightly more US-origin requests. The service underrepresents enterprise workloads, and agents can consume many tokens per task.

When could China’s AI industry become profitable?

Robert Lea forecasts that the industry could remain unprofitable until 2030. That is an analyst forecast, tied to competitive pressure and pricing, rather than a demonstrated outcome.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Stanford AI Index 2026 pegs US-China AI gap at 2.7%
Stanford's 2026 AI Index closes the US-China performance gap to 2.7%. But the Chinese labs that closed it are pivoting to closed source, and DeepSeek V4 has gone silent on Huawei silicon. China caught up the week Chinese labs quit the game that got them there.
Z.ai Delays GLM-5.3 Weights After CyberGym Score Tops Mythos
GLM-5.3 scored 84.5% on CyberGym, edging Anthropic's restricted Mythos 5, and Z.ai responded by holding its downloadable weights until around August 28. The lead vanishes on exploitation benchmarks, and every figure came from Z.ai's own harness.
AI News

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai