> ## Content Index
> Fetch the complete content index at: https://www.implicator.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Google Argon: The eternal third and why some Google staffers say its coding falls short of its benchmarks
- URL: https://www.implicator.ai/google-gemini-4-argon-staff-doubt-coding/
- Published: 2026-10-01T05:27:44.000Z
- Updated: 2026-10-01T05:27:44.000Z
- Description: Gemini 4 Argon, Google's first frontier model in more than seven months, ties GPT-6 Astra on an independent index and trails Anthropic's newest Claude models. Some Google employees say it struggles with real coding work. Google calls that inaccurate.
- Author: Marcus Schuler
- Tags: AI News

Koray Kavukcuoglu, a longtime lieutenant of Demis Hassabis who took over Google DeepMind's day-to-day operations in August when Hassabis became chairman, told a conference audience last week he was encouraged by the model's performance. "I have the utmost trust in the team," he said. "In my mind, it's a certainty that we are always gonna be at the frontier."

Some Google employees who have used the model say it does less well on real work, coding in particular, than its benchmark scores suggest, [Bloomberg reported](https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about-new-gemini-model?ref=implicator.ai).

Google [unveiled Gemini 4 Argon](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/?ref=implicator.ai) on Wednesday, initially for selected cybersecurity partners, with leading scores on several software engineering and professional-work benchmarks. Argon scored 53 on the Artificial Analysis Intelligence Index, level with OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 and trailing Anthropic's newest models.

The Breakdown

- Google unveiled Gemini 4 Argon on Sept. 30, first only to cybersecurity partners in its Fairwind program; paid API customers and Google AI Ultra subscribers are next, with no date.
- Argon scored 53 on the Artificial Analysis Intelligence Index, level with GPT-6 Astra and Claude Fable 5.1 and behind Claude Opus 5.5 (58) and Claude Sonnet 5.5 (56).
- Some Google employees with direct access say it struggles with certain coding tasks, including front-end design, and two people said it appears affected by "benchmaxxing."
- Google said it would be inaccurate to say the model underperforms in coding; one employee cited "large consensus" that Argon is at the frontier.

AI-generated summary, reviewed by an editor. [More on our AI guidelines](https://www.implicator.ai/about/).

## The doubts inside Google

Employees with direct access to Argon said it struggled with certain coding tasks when they put it to work. One person singled out front-end design, the work that determines how an app or website looks and feels. Some employees believed Anthropic's Fable and OpenAI's Astra models were improving faster than Gemini and that Gemini 4, even at its best, would still lag behind those models in some areas. Others thought Google had caught up.

Alphabet shares rose about 1.7% in extended trading after the announcement; late in regular trading they had pared the day's earlier gains after the report on employee skepticism appeared.

Two people familiar with Argon said it appeared affected by "benchmaxxing." Experts describe this as concentrating engineering effort on getting a high test score, with less attention to whether the resulting product does the user's job well. They say labs prioritize those scores because customers often judge models by benchmark rankings.

Edwin Chen, founder of the AI startup Surge AI, described how reliance on benchmarks generally could lead a lab to produce code in a particular programming language while giving less attention to whether the finished app was easy to use or well designed.

"An analogy would be, 'Oh yeah, my kid got a really good score on the SAT', but the SAT doesn't translate into real-world performance," Chen said. "It's an incredibly pernicious problem."

FREE · ABOUT FIVE MINUTES

Track the frontier model race as it shifts.

Implicator cuts the flood to the developments worth your attention, with concise reporting on what changed and why it matters. Delivered every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.

Email address 

Get the free briefing 

Check your inbox for the confirmation link.

From San Francisco. No spam. Unsubscribe anytime.

## Google's case

The employees describing Argon's shortcomings were anonymous. Google disputed their account, saying it would be inaccurate to say the model underperformed in areas such as coding. An employee familiar with model development said there was "large consensus" internally that Argon was at the frontier, cited rigorous testing and denied difficulties with messy, real-world coding tasks.

Tulsee Doshi, who leads Gemini products at DeepMind, said, "We've seen strong performance up close as Googlers have put the model through its paces in recent weeks, with many relying on it for their hardest coding and research problems."

Google's launch announcement offered an example from its libgav1 video decoder. Argon agents replaced 32,000 lines of SIMD code in an existing Rust port. Google said the resulting decoder ran 2.7 times faster than that earlier Rust version with identical video output.

A person familiar with the matter said the model stood out at making sense of inputs beyond text, such as extracting metadata from video, and cited its safety, cybersecurity and ability to communicate clearly and naturally.

## What the independent tests show

On the [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/articles/gemini-4-argon-google-top-three-labs?ref=implicator.ai), Argon at its highest available reasoning setting scored 53 points, matching GPT-6 Astra and Claude Fable 5.1\. Claude Opus 5.5 scored 58 and Claude Sonnet 5.5 scored 56\. Google's previous non-Flash model, Gemini 3.1 Pro Preview, scored 30 on the same index. Artificial Analysis described Google as "one of the top three labs in intelligence achieved."

Coding rankings varied by test. In Google's launch comparison, Argon scored 77.9% on DeepSWE v1.1, which measures long-horizon software engineering, against Claude Opus 5.5's 74.2% and GPT-6 Astra's 74.1%. Across the 18 benchmarks Google disclosed, Argon led outright on 12 and tied for first on another.

On Artificial Analysis's Terminal Bench 4 evaluation, Argon scored 57%, behind Claude Sonnet 5.5 at 64%, Claude Opus 5.5 at 60% and GPT-6 Astra at 59%.

Know someone who'd find this useful? [✉️ Email it to a friend in one click](mailto:?subject=A%20newsletter%20I%20think%20you%27d%20like&body=This%20is%20one%20of%20maybe%20three%20newsletters%20I%20actually%20read.%20The%20rest%20just%20pile%20up%2C%20unread%2C%20judging%20me.%0A%0AAnd%20yes%2C%20this%20email%20mostly%20wrote%20itself%2C%20which%20is%20a%20little%20on%20the%20nose%20for%20an%20AI%20newsletter.%20Doesn%27t%20make%20it%20wrong.%20implicator.ai%20is%20good.%0A%0ASubscribe%20free%3A%20https%3A%2F%2Fwww.implicator.ai%2Fsubscribe%2F%3Futm%5Fsource%3Dnewsletter%26utm%5Fmedium%3Dforward%26utm%5Fcampaign%3Demail%5Fforward), or they can [subscribe free here](https://www.implicator.ai/subscribe/?utm%5Fsource=newsletter&utm%5Fmedium=forward&utm%5Fcampaign=forward%5Fto%5Fcolleague).

On Arena's web-development ranking, where people vote on outputs, Argon placed eighth with 1,679 points, up from 29th for Gemini 3.8 Flash.

The independent results also gave Google category leads. Argon topped Text Arena's human preference ranking with 1,525 points, 20 ahead of Claude Opus 4.6\. It [placed first on the Vals Index](https://www.vals.ai/models/google%5Fgemini-4-argon?ref=implicator.ai) at 68.9%, becoming the first Gemini model to lead that measure of professional work. The Vals ranking also puts Anthropic's mid-tier Claude Sonnet 5.5 above its top model, Claude Opus 5.5, a quirk to treat with caution.

## A year of catching up

Last November, Google debuted Gemini 3, a well-received model widely seen as a turning point in its efforts to keep up with OpenAI and Anthropic. Google announced Gemini 3.5 Pro at I/O in May, promised it for June and then abandoned it. Argon is its first proprietary model above the Flash class in more than seven months. Training runs to build such a model can cost as much as $400 million, according to Mandeep Singh, an analyst at Bloomberg Intelligence.

Google's worries about AI coding predate Argon. In April, three former employees said some DeepMind teams, including teams working on Gemini, [used Anthropic's Claude Code](https://www.bloomberg.com/news/articles/2026-04-21/google-struggles-to-gain-ground-in-ai-coding-as-rivals-advance?ref=implicator.ai). Kathy Korevec, who oversaw Google's Jules coding tool before leaving for OpenAI that month, described fragmented, overlapping developer tools in a post on X.

"That's not a talent problem. It's a systems problem," she wrote.

## Who can use it

Most developers cannot yet test Argon against their own work. Initial access is through [Google's Fairwind cybersecurity program](https://blog.google/innovation-and-ai/technology/safety-security/fairwind-program/?ref=implicator.ai), and the company is participating in the U.S. government's voluntary pre-release model access process.

Paid API customers and Google AI Ultra subscribers are next in line. Google has given them no release date, saying access will come "as soon as possible."

Frequently Asked Questions

What is Gemini 4 Argon?

Gemini 4 Argon is Google's new flagship AI model, unveiled on Sept. 30, 2026\. It is Google's first proprietary model above the Flash class in more than seven months, after the company announced Gemini 3.5 Pro in May, promised it for June and then abandoned it.

How does Argon compare with OpenAI and Anthropic models?

On the Artificial Analysis Intelligence Index, Argon scored 53, matching GPT-6 Astra and Claude Fable 5.1, while Claude Opus 5.5 scored 58 and Claude Sonnet 5.5 scored 56\. On Terminal Bench 4 it scored 57%, behind Sonnet 5.5, Opus 5.5 and GPT-6 Astra. It leads Text Arena with 1,525 points and the Vals Index at 68.9%.

What are Google employees saying about Argon?

Anonymous employees with direct access said it struggled with certain coding tasks when they put it to work, and one singled out front-end design. Some believe Anthropic's Fable and OpenAI's Astra models are improving faster and that Gemini 4 will still lag behind them in some areas. Others think Google has caught up.

What is benchmaxxing?

Experts describe benchmaxxing as concentrating engineering effort on getting a high test score, with less attention to whether the product does the user's job well. They say labs do it because customers often judge models by benchmark rankings. Two people familiar with Argon said it appeared affected.

When can developers use Gemini 4 Argon?

Initial access runs through Google's Fairwind cybersecurity program, and Google is participating in the U.S. government's voluntary pre-release model access process. Paid API customers and Google AI Ultra subscribers are next in line, but Google has given no release date beyond "as soon as possible."

AI-generated summary, reviewed by an editor. [More on our AI guidelines](https://www.implicator.ai/about/).

## Related stories

[Hassabis Steps Aside at Google DeepMind, Jeff Dean ExitsDemis Hassabis is handing over day-to-day control of Google DeepMind, and Jeff Dean is leaving after 27 years to run a startup Google is investing in. Alphabet shares fell more than 5%, a reading that sits awkwardly beside the company's own account of the day.Implicator.ai![](https://www.implicator.ai/content/images/size/w1200/2026/08/2026-08-05-10.29.44-google-deepmind-hassabis-steps-aside-jeff-dean-exit@2x.webp)](https://www.implicator.ai/google-deepmind-hassabis-steps-aside-jeff-dean-exit/)

[Gemini 3.8 Flash Scores 59, Trails Fable and SolGoogle’s Gemini 3.8 Flash sits two index points behind GPT-5.6 Sol and seven behind Claude Fable 5.1, but costs less per measured task. Open-weights GLM is cheaper, and the Cyber model’s external cost is still missing.Implicator.ai![](https://www.implicator.ai/content/images/2026/09/20260902-143849-gemini_flash_bench2.webp)](https://www.implicator.ai/gemini-3-8-flash-scores-59-behind-fable-and-sol/)

[Gemini Hits 950 Million Users as Alphabet Burns CashAlphabet said the Gemini app reached 950 million monthly active users, up from about 650 million last October. The same quarter brought its first negative free cash flow since 2004, a capital-spending forecast of up to $205 billion, and a CEO defending a delayed flagship model.Implicator.ai![](https://www.implicator.ai/content/images/2026/07/2026-07-22-19.32.52-google-gemini-950-million-users-negative-cash-flow@2x.webp)](https://www.implicator.ai/google-says-gemini-reached-950-million-monthly-users-as-cash-flow-turned-negative/)