Google introduced Gemini 3.8 Flash and the access-restricted Gemini 3.8 Flash Cyber on September 2, with the general model’s independently measured score of 59 sitting behind Claude Fable 5.1 and GPT-5.6 Sol. At $0.58 per measured task, Gemini cost less than both higher-scoring systems. Developers choosing a model for agent work must weigh lower task cost and faster generation against lower benchmark performance.
The Breakdown
- Gemini 3.8 Flash scored 59 on the September 2 Artificial Analysis snapshot, behind GPT-5.6 Sol at 61 and Claude Fable 5.1 at 66, but ahead of GLM-5.3 Flash at 57.
- Its measured cost was $0.58 per Intelligence Index task, below Sol at $0.95 and Fable at $3.69, but above open-weights GLM at $0.09.
- Google kept the introductory token rates at $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, while measured task cost rose from $0.40 for Gemini 3.7 Flash to $0.58 for 3.8.
- Gemini 3.8 Flash Cyber scored 47.2% on CWE-Bench, but the external leaderboard lists no cost for it and access is restricted through Google’s Fairwind program.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
Two points behind Sol
On the Artificial Analysis Intelligence Index, accessed September 2, Gemini 3.8 Flash at high effort scored 59 and cost $0.58 per task. GPT-5.6 Sol at maximum effort scored 61 and cost $0.95 per task, while Claude Fable 5.1 at maximum adaptive effort scored 66 and cost $3.69. Fable’s default fallback means its result is not a pure test of one model.
The task prices are weighted averages across the index, not provider list prices. Full-run totals also depend on each system’s token use, so models with identical per-token rates can produce different bills on the same evaluation.
Gemini generated 304.6 output tokens a second and produced 120 million output tokens across the evaluation; the full run cost $825.83. Sol generated 72.4 tokens a second, produced 70 million output tokens and cost $2,017.29. Fable generated 66.4 tokens a second, and its evaluation cost $8,523.16.
In the index snapshot accessed September 2, open-weights GLM-5.3 Flash was cheaper and scored lower. It scored 57, two points below Gemini, but cost $0.09 per task and $138.02 for the full run. Gemini was much faster than GLM’s 42.8 output tokens a second and used fewer than GLM’s 150 million output tokens, yet Gemini’s $0.58 measured task cost was about 6.4 times GLM’s $0.09 cost.
Unchanged rates, higher task cost
In the index snapshot accessed September 2, Gemini 3.8 improved on Gemini 3.7 Flash’s high-effort score of 56. The gain came with a rise in cost per task from $0.40 to $0.58 and an increase in output volume from 64 million to 120 million tokens across the respective evaluations.
FREE WEEKDAY MORNING BRIEFING
Track what each new AI model really costs.
The Implicator Morning Briefing filters the AI news cycle to the stories worth your attention and explains their consequences. From San Francisco, every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.
About five minutes. No hype. No spam.
The API rates did not change between Gemini 3.7 Flash and Gemini 3.8 Flash: $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Those rates double on January 1, 2027. Stable token prices therefore did not produce a stable task cost in the outside test. Actual results can also differ at the model’s default medium effort, since this comparison used high effort.
Google’s model card states the tradeoff directly: “At times, the model might use more tokens to maximize performance, especially at higher effort levels.” Developers can lower the effort setting or continue using 3.7 for work where compute efficiency matters more.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
In Finance Agent v2, accessed September 2, 3.8 led three of nine categories, but 3.7 still led general quantitative analysis. On the held-out Harvey legal test that day, 3.8 resolved 10.0% of complete tasks. Muse Spark 1.2 led complete-task resolution at 25.42%.
A separate cyber release
On CWE-Bench, accessed September 2, Gemini 3.8 Flash Cyber passed 47.2% of 100 held-out audit-and-patch tasks, just below Claude Fable 5 at 47.8% (the general Intelligence Index uses Claude Fable 5.1) and above GPT-5.6 Sol at 44.2% and Gemini 3.7 Flash at 44.0%. The deterministic verifier accepts a patch only when the exploit stops working and existing tests still pass.
Google’s own tests put the cyber model above 70% success across 20 programming languages and found 2.6 times as many correct Chrome patches as unnamed larger commercial models. Google did not disclose the underlying test sets or the competitor identities.
The cyber model is not available through the general Gemini API. Fairwind access is limited to approved governments, infrastructure operators, software maintainers and research teams, with named-user controls. Partners may use it for permitted defensive or academic research and cannot share, resell or redistribute access.
As of September 2, CWE-Bench publishes average rollout costs of $10.27 for Fable 5, $2.29 for Sol and $1.43 for Gemini 3.7, but no cost for 3.8 Cyber. Google’s claim that the new model costs “significantly” less therefore has no external figure to test.
Frequently Asked Questions
How did Gemini 3.8 Flash score against its main competitors?
In the Artificial Analysis snapshot accessed September 2, Gemini 3.8 Flash scored 59 at high effort. GPT-5.6 Sol scored 61, Claude Fable 5.1 with its default fallback scored 66, and open-weights GLM-5.3 Flash scored 57.
Is Gemini 3.8 Flash cheaper than GPT-5.6 Sol and Claude Fable 5.1?
On the Intelligence Index, Gemini cost $0.58 per measured task, compared with $0.95 for GPT-5.6 Sol and $3.69 for Claude Fable 5.1. GLM-5.3 Flash was cheaper than all three at $0.09 per task.
Why did Gemini’s task cost rise if its token price stayed the same?
Gemini 3.8 Flash produced 120 million output tokens across the high-effort evaluation, up from 64 million for Gemini 3.7 Flash. The measured cost per task rose from $0.40 to $0.58 even though Google kept the introductory input and output rates unchanged.
How did Gemini 3.8 Flash Cyber perform on CWE-Bench?
It passed 47.2% of 100 held-out audit-and-patch tasks. Claude Fable 5 scored 47.8%, GPT-5.6 Sol scored 44.2%, and Gemini 3.7 Flash scored 44.0%. CWE-Bench does not list an average rollout cost for Gemini 3.8 Flash Cyber.
Can any developer use Gemini 3.8 Flash Cyber?
No. It is not offered through the general Gemini API. Google limits access through Fairwind to approved governments, infrastructure operators, software maintainers and research teams, with named-user controls and no redistribution.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
Related stories



IMPLICATOR