> ## Content Index
> Fetch the complete content index at: https://www.implicator.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# DeepSeek Narrows US AI Benchmark Lead to 3% After September Release
- URL: https://www.implicator.ai/deepseek-us-ai-benchmark-gap-3-percent/
- Published: 2026-10-04T22:45:24.000Z
- Updated: 2026-10-04T22:45:24.000Z
- Description: DeepSeek has brought China within about 3% of the US leader in a LiveBench comparison. Its agentic coding score exceeds Anthropic’s leading overall entry, but token traffic and profitability forecasts show how much the benchmark leaves out of the commercial contest.
- Author: Marcus Schuler
- Tags: AI News, #home-companion

DeepSeek’s September release narrowed the US lead over China to [roughly 3% in a LiveBench score comparison](https://www.bloomberg.com/news/articles/2026-10-04/us-lead-in-ai-over-china-narrows-after-deepseek-gains-bi-says?ref=implicator.ai). The V4.1 Flash model moved close to Anthropic’s leading system on [LiveBench’s overall tests](https://livebench.ai/?ref=implicator.ai#/), which include reasoning and coding. Bloomberg Intelligence senior analyst Robert Lea expects the improved performance to bring Chinese developers further market share gains.

In the leaderboard captured on October 4, DeepSeek V4.1 Flash Max Effort scored 81.1 against 83.4 for Anthropic’s Claude Fable 5.1 Max Effort. That is a difference of 2.3 score points, or about 2.8% of Anthropic’s score. Lea’s comparison put the gap at about 9% in May and 15% earlier in 2026.

The scores differ by task. In the October 4 snapshot, DeepSeek scored 77.3 on agentic coding, compared with 66.1 for Claude Fable 5.1 Max Effort.

The US still holds the lead in this comparison. These benchmark results do not establish national AI leadership or show whether US chip export restrictions are effective.

What Changed

- DeepSeek narrowed China’s gap in a LiveBench comparison to about 3%, from roughly 9% in May.
- The October 4 leaderboard showed DeepSeek at 81.1 overall, against Anthropic’s 83.4.
- OpenRouter token traffic does not measure the whole AI market and underrepresents enterprise workloads.
- Analyst Robert Lea forecasts that China’s AI industry could remain unprofitable until 2030.

AI-generated summary, reviewed by an editor. [More on our AI guidelines](https://www.implicator.ai/about/).

## DeepSeek’s efficiency changes

DeepSeek’s [September 10 release notes](https://api-docs.deepseek.com/news/news260910/?ref=implicator.ai) describe a design intended to cut the computing and memory needed to serve answers. Only a small portion of the model’s stored parameters, the values learned during training, is active when it reads input or generates output. The company says it also shrank the cache that holds information used during a conversation.

FREE AI BRIEFING · WEEKDAYS

Follow the changing economics of AI models.

Get the AI stories shaping the day, with concise context from San Francisco. The briefing takes about five minutes and arrives at 4:45 a.m. Pacific, 7:45 a.m. Eastern.

Email address 

Keep me informed 

Check your inbox for the confirmation link.

Free. No hype. Unsubscribe anytime.

Those efficiency claims come from DeepSeek. The release supports native visual understanding and continues to offer off-peak API rates at half the peak price, under the new pricing that took effect on September 10\. The benchmark gain does not by itself establish how reliably the model will perform inside a company or whether it can be deployed securely.

## OpenRouter usage and enterprise demand

Chinese models processed more tokens than US models on OpenRouter in the trailing week to September 7\. But US models handled slightly more requests in that period, the usage window examined in [JPMorgan Asset Management’s September 9 analysis](https://am.jpmorgan.com/us/en/asset-management/institutional/insights/market-insights/market-updates/on-the-minds-of-investors/how-is-the-us-china-ai-race-evolving/?ref=implicator.ai). OpenRouter is a service developers use to send requests to different AI systems.

Token share is not market share. Automated agents can produce large volumes of tokens while working through a task, so heavy token consumption can make a provider look more widely used than a count of requests does.

OpenRouter also underrepresents enterprise activity. Most of those workloads run through major cloud platforms or directly with developers such as OpenAI and Anthropic. Large US companies still predominantly use US models, while large Chinese companies tend to use domestic ones.

Lower token prices do not always mean cheaper completed tasks. Models consume different amounts of text before reaching an answer, and downloadable models still incur serving costs. A price comparison based on tokens alone can therefore miss the cost of completing the work.

Know someone who'd find this useful? [✉️ Email it to a friend in one click](mailto:?subject=A%20newsletter%20I%20think%20you%27d%20like&body=This%20is%20one%20of%20maybe%20three%20newsletters%20I%20actually%20read.%20The%20rest%20just%20pile%20up%2C%20unread%2C%20judging%20me.%0A%0AAnd%20yes%2C%20this%20email%20mostly%20wrote%20itself%2C%20which%20is%20a%20little%20on%20the%20nose%20for%20an%20AI%20newsletter.%20Doesn%27t%20make%20it%20wrong.%20implicator.ai%20is%20good.%0A%0ASubscribe%20free%3A%20https%3A%2F%2Fwww.implicator.ai%2Fsubscribe%2F%3Futm%5Fsource%3Dnewsletter%26utm%5Fmedium%3Dforward%26utm%5Fcampaign%3Demail%5Fforward), or they can [subscribe free here](https://www.implicator.ai/subscribe/?utm%5Fsource=newsletter&utm%5Fmedium=forward&utm%5Fcampaign=forward%5Fto%5Fcolleague).

Enterprise procurement also depends on regulation and data governance. Downloading model weights permits private hosting, but does not reveal the training data or resolve intellectual property questions.

## Profitability remains uncertain

Lea forecasts that China’s AI industry could remain unprofitable until 2030 despite the benchmark gains. In his assessment, low-margin token supply and a domestic price war may prevent developers from building a lasting commercial advantage.

China’s market is now crowded with more than 1,100 large language models.

Lea identifies ByteDance’s Doubao as the frontrunner in AI app monetization, while rival chatbots from DeepSeek and Tencent remain free.

“Putting China’s AI sector on a sustainable profit footing will require a cooling of competitive pressures, an industry shakeout, and a more rational approach to pricing,” he said.

Frequently Asked Questions

How close is DeepSeek to the US leader on LiveBench?

The October 4 snapshot showed DeepSeek V4.1 Flash Max Effort at 81.1 overall and Anthropic’s Claude Fable 5.1 Max Effort at 83.4\. The 2.3-point difference is about 2.8% of Anthropic’s score, rounded to roughly 3% in the comparison.

Does DeepSeek lead on every task?

No. DeepSeek scored 77.3 on agentic coding against 66.1 for Claude Fable 5.1 Max Effort in the October 4 snapshot. Anthropic held the higher overall score.

What changed in DeepSeek’s September release?

DeepSeek described a design intended to reduce the computing and memory needed to serve answers, using only a small part of its stored parameters at a time and a smaller conversation cache. These efficiency claims come from the company.

Does Chinese token traffic prove market leadership?

No. JPMorgan’s September 9 analysis found more Chinese-origin token traffic on OpenRouter but slightly more US-origin requests. The service underrepresents enterprise workloads, and agents can consume many tokens per task.

When could China’s AI industry become profitable?

Robert Lea forecasts that the industry could remain unprofitable until 2030\. That is an analyst forecast, tied to competitive pressure and pricing, rather than a demonstrated outcome.

AI-generated summary, reviewed by an editor. [More on our AI guidelines](https://www.implicator.ai/about/).

## Related stories

[Stanford AI Index 2026 pegs US-China AI gap at 2.7%Stanford's 2026 AI Index closes the US-China performance gap to 2.7%. But the Chinese labs that closed it are pivoting to closed source, and DeepSeek V4 has gone silent on Huawei silicon. China caught up the week Chinese labs quit the game that got them there.Implicator.ai![](https://www.implicator.ai/content/images/2026/04/2026-04-13-09.15.19-ai_scoreboard@2x.webp)](https://www.implicator.ai/stanfords-2026-ai-index-puts-us-lead-over-china-at-2-7-as-deepseek-v4-stalls/)

[Z.ai Delays GLM-5.3 Weights After CyberGym Score Tops MythosGLM-5.3 scored 84.5% on CyberGym, edging Anthropic's restricted Mythos 5, and Z.ai responded by holding its downloadable weights until around August 28\. The lead vanishes on exploitation benchmarks, and every figure came from Z.ai's own harness.Implicator.ai![](https://www.implicator.ai/content/images/2026/08/2026-08-15-03.27.10-z-ai-delays-glm-5-3-weights-cyber-score-mythos-5@2x.webp)](https://www.implicator.ai/z-ai-delays-glm-5-3-weights-two-weeks-after-cyber-score-beats-mythos-5/)