DeepSeek released V4-Flash-0731 and placed its official API in public beta on July 31, while Artificial Analysis scored the model 50 on its Intelligence Index, 10 points above the April preview. The mixture-of-experts architecture and model size were unchanged, with the improvement coming from re-post-training rather than a new design. DeepSeek's model card identifies V4-Flash-0731 as the official release superseding the April preview.

Artificial Analysis's GDPval-AA v2 evaluates agentic real-world work tasks. The model's rating rose to 1,559 from 1,189, placing it below Kimi K3 and above GLM-5.2 in that evaluation. Terminal-Bench 2.1 increased 17 points to 79%, and Tau-cubed-Bench Banking gained eight points to 31%. The same benchmark run used about 206 million output tokens, 12% fewer than its predecessor.

What Changed

AI-generated summary, reviewed by an editor. More on our AI guidelines.

The base model retains 284 billion total parameters, with 13 billion active at inference, and a 1-million-token context window. The Hugging Face repository's 304-billion figure includes the attached DSpark speculative-decoding module.

The broader index placed V4 Flash level with Gemini 3.6 Flash and one point behind GPT-5.6 Luna at maximum reasoning. The evaluator calculated that cost per task was about 60% lower than GPT-5.6 Luna at maximum reasoning, even after OpenAI cut Luna's price. It attributed much of the difference to DeepSeek's roughly 98% cache-hit discount, compared with 90% at most providers. DeepSeek published the weights under an MIT license and kept first-party API prices at the preview's levels: $0.14 per million input tokens, $0.28 per million output tokens and $0.0028 for cached input.

Its AA-Omniscience score improved by seven points to negative 16 because the hallucination rate fell, according to the evaluation. Accuracy stayed at 37%, while the hallucination rate declined to 84%, comparable in the same evaluation to GPT-5.6 Terra at 85% and Mistral Medium 3.5 at 82%.

DeepSeek's own model card reported higher agent scores in its evaluation setup. Terminal-Bench 2.1 reached 82.7, up from 61.8 for the preview, and DeepSWE rose to 54.4 from 7.3. Asif Razzaq of MarkTechPost noted that the largest self-reported jumps came from a setup outsiders cannot yet reproduce: the company disclosed that its public code-agent tests used an unreleased minimal mode of DeepSeek Harness, and two DSBench results came from internal test sets. He advised developers to run their own evaluations first.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

In a July 31 post, Benjamin Marie of The Kaitchup contrasted DeepSeek's immediate weight release with waiting periods at MiniMax and Qwen. Reviewing the evaluator's additional results, Marie found gains on other benchmarks, including GPQA Diamond, which he described as very different from agentic coding, and no reported regressions.

Developer Simon Willison wrote that the model may offer the best value per unit of measured intelligence, but his small OpenRouter test showed how settings can affect practical output. Using his saved pelican prompt in an informal image-generation test, Willison got a disappointing result at the default reasoning level. Raising reasoning_effort to high produced a much better result, matching the model card's low, high and max control levels.

The update applies to the V4-Flash API, which now supports the Responses API and is adapted for Codex. DeepSeek's V4-Pro API and the models used in its app and website were not updated. The company said the official V4-Pro release "will follow soon."

Frequently Asked Questions

What score did DeepSeek V4-Flash-0731 get on the Artificial Analysis Intelligence Index?

Artificial Analysis scored it 50, ten points above the April 2026 V4 Flash preview. That places it level with Gemini 3.6 Flash and one point behind GPT-5.6 Luna at maximum reasoning.

Did DeepSeek change the model's architecture for this release?

No. The mixture-of-experts architecture and model size were unchanged. The base model keeps 284 billion total parameters with 13 billion active at inference and a 1-million-token context window. The gain came from re-post-training rather than a new design.

How much does the V4-Flash API cost?

DeepSeek kept first-party API prices at the preview's levels: $0.14 per million input tokens, $0.28 per million output tokens and $0.0028 for cached input. The evaluator put cost per task about 60% below GPT-5.6 Luna at maximum reasoning, even after OpenAI cut Luna's price.

Are DeepSeek's published agent benchmark scores independently verifiable?

Not yet. Asif Razzaq of MarkTechPost noted the largest self-reported jumps came from a setup outsiders cannot reproduce, because the public code-agent tests used an unreleased minimal mode of DeepSeek Harness and two DSBench results came from internal test sets. He advised developers to run their own evaluations first.

Are the model weights available?

Yes. DeepSeek published the weights on Hugging Face under an MIT license. Benjamin Marie of The Kaitchup contrasted that immediate release with the waiting periods seen at MiniMax and Qwen.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Moonshot Releases Kimi K3 Weights as Amodei Rejects Open-Weight Ban
Moonshot AI released the weights for its Kimi K3 model on Monday, a system with 2.8 trillion parameters. Developers can now download, modify and self-host K3, which the report described as the world's
Germany's Soofi S AI Model Tops All Open-Source Rivals on German Benchmarks
A German research consortium coordinated by the KI Bundesverband released Soofi S, an open-source German-English foundation model, this week, according to its pretraining report. In the team's tests,
Moonshot Launches Kimi K3 With 2.8 Trillion Parameters and 1M Context
Moonshot AI said Thursday that it had launched Kimi K3, a model built with 2.8 trillion total parameters. K3 is available now through Moonshot’s products and API. Full weights remain unavailable; Moon
AI News

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai