Koray Kavukcuoglu, a longtime lieutenant of Demis Hassabis who took over Google DeepMind's day-to-day operations in August when Hassabis became chairman, told a conference audience last week he was encouraged by the model's performance. "I have the utmost trust in the team," he said. "In my mind, it's a certainty that we are always gonna be at the frontier."
Some Google employees who have used the model say it does less well on real work, coding in particular, than its benchmark scores suggest, Bloomberg reported.
Google unveiled Gemini 4 Argon on Wednesday, initially for selected cybersecurity partners, with leading scores on several software engineering and professional-work benchmarks. Argon scored 53 on the Artificial Analysis Intelligence Index, level with OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 and trailing Anthropic's newest models.
The Breakdown
- Google unveiled Gemini 4 Argon on Sept. 30, first only to cybersecurity partners in its Fairwind program; paid API customers and Google AI Ultra subscribers are next, with no date.
- Argon scored 53 on the Artificial Analysis Intelligence Index, level with GPT-6 Astra and Claude Fable 5.1 and behind Claude Opus 5.5 (58) and Claude Sonnet 5.5 (56).
- Some Google employees with direct access say it struggles with certain coding tasks, including front-end design, and two people said it appears affected by "benchmaxxing."
- Google said it would be inaccurate to say the model underperforms in coding; one employee cited "large consensus" that Argon is at the frontier.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
The doubts inside Google
Employees with direct access to Argon said it struggled with certain coding tasks when they put it to work. One person singled out front-end design, the work that determines how an app or website looks and feels. Some employees believed Anthropic's Fable and OpenAI's Astra models were improving faster than Gemini and that Gemini 4, even at its best, would still lag behind those models in some areas. Others thought Google had caught up.
Alphabet shares rose about 1.7% in extended trading after the announcement; late in regular trading they had pared the day's earlier gains after the report on employee skepticism appeared.
Two people familiar with Argon said it appeared affected by "benchmaxxing." Experts describe this as concentrating engineering effort on getting a high test score, with less attention to whether the resulting product does the user's job well. They say labs prioritize those scores because customers often judge models by benchmark rankings.
Edwin Chen, founder of the AI startup Surge AI, described how reliance on benchmarks generally could lead a lab to produce code in a particular programming language while giving less attention to whether the finished app was easy to use or well designed.
"An analogy would be, 'Oh yeah, my kid got a really good score on the SAT', but the SAT doesn't translate into real-world performance," Chen said. "It's an incredibly pernicious problem."
FREE · ABOUT FIVE MINUTES
Track the frontier model race as it shifts.
Implicator cuts the flood to the developments worth your attention, with concise reporting on what changed and why it matters. Delivered every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.
From San Francisco. No spam. Unsubscribe anytime.
Google's case
The employees describing Argon's shortcomings were anonymous. Google disputed their account, saying it would be inaccurate to say the model underperformed in areas such as coding. An employee familiar with model development said there was "large consensus" internally that Argon was at the frontier, cited rigorous testing and denied difficulties with messy, real-world coding tasks.
Tulsee Doshi, who leads Gemini products at DeepMind, said, "We've seen strong performance up close as Googlers have put the model through its paces in recent weeks, with many relying on it for their hardest coding and research problems."
Google's launch announcement offered an example from its libgav1 video decoder. Argon agents replaced 32,000 lines of SIMD code in an existing Rust port. Google said the resulting decoder ran 2.7 times faster than that earlier Rust version with identical video output.
A person familiar with the matter said the model stood out at making sense of inputs beyond text, such as extracting metadata from video, and cited its safety, cybersecurity and ability to communicate clearly and naturally.
What the independent tests show
On the Artificial Analysis Intelligence Index, Argon at its highest available reasoning setting scored 53 points, matching GPT-6 Astra and Claude Fable 5.1. Claude Opus 5.5 scored 58 and Claude Sonnet 5.5 scored 56. Google's previous non-Flash model, Gemini 3.1 Pro Preview, scored 30 on the same index. Artificial Analysis described Google as "one of the top three labs in intelligence achieved."
Coding rankings varied by test. In Google's launch comparison, Argon scored 77.9% on DeepSWE v1.1, which measures long-horizon software engineering, against Claude Opus 5.5's 74.2% and GPT-6 Astra's 74.1%. Across the 18 benchmarks Google disclosed, Argon led outright on 12 and tied for first on another.
On Artificial Analysis's Terminal Bench 4 evaluation, Argon scored 57%, behind Claude Sonnet 5.5 at 64%, Claude Opus 5.5 at 60% and GPT-6 Astra at 59%.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
On Arena's web-development ranking, where people vote on outputs, Argon placed eighth with 1,679 points, up from 29th for Gemini 3.8 Flash.
The independent results also gave Google category leads. Argon topped Text Arena's human preference ranking with 1,525 points, 20 ahead of Claude Opus 4.6. It placed first on the Vals Index at 68.9%, becoming the first Gemini model to lead that measure of professional work. The Vals ranking also puts Anthropic's mid-tier Claude Sonnet 5.5 above its top model, Claude Opus 5.5, a quirk to treat with caution.
A year of catching up
Last November, Google debuted Gemini 3, a well-received model widely seen as a turning point in its efforts to keep up with OpenAI and Anthropic. Google announced Gemini 3.5 Pro at I/O in May, promised it for June and then abandoned it. Argon is its first proprietary model above the Flash class in more than seven months. Training runs to build such a model can cost as much as $400 million, according to Mandeep Singh, an analyst at Bloomberg Intelligence.
Google's worries about AI coding predate Argon. In April, three former employees said some DeepMind teams, including teams working on Gemini, used Anthropic's Claude Code. Kathy Korevec, who oversaw Google's Jules coding tool before leaving for OpenAI that month, described fragmented, overlapping developer tools in a post on X.
"That's not a talent problem. It's a systems problem," she wrote.
Who can use it
Most developers cannot yet test Argon against their own work. Initial access is through Google's Fairwind cybersecurity program, and the company is participating in the U.S. government's voluntary pre-release model access process.
Paid API customers and Google AI Ultra subscribers are next in line. Google has given them no release date, saying access will come "as soon as possible."
Frequently Asked Questions
What is Gemini 4 Argon?
Gemini 4 Argon is Google's new flagship AI model, unveiled on Sept. 30, 2026. It is Google's first proprietary model above the Flash class in more than seven months, after the company announced Gemini 3.5 Pro in May, promised it for June and then abandoned it.
How does Argon compare with OpenAI and Anthropic models?
On the Artificial Analysis Intelligence Index, Argon scored 53, matching GPT-6 Astra and Claude Fable 5.1, while Claude Opus 5.5 scored 58 and Claude Sonnet 5.5 scored 56. On Terminal Bench 4 it scored 57%, behind Sonnet 5.5, Opus 5.5 and GPT-6 Astra. It leads Text Arena with 1,525 points and the Vals Index at 68.9%.
What are Google employees saying about Argon?
Anonymous employees with direct access said it struggled with certain coding tasks when they put it to work, and one singled out front-end design. Some believe Anthropic's Fable and OpenAI's Astra models are improving faster and that Gemini 4 will still lag behind them in some areas. Others think Google has caught up.
What is benchmaxxing?
Experts describe benchmaxxing as concentrating engineering effort on getting a high test score, with less attention to whether the product does the user's job well. They say labs do it because customers often judge models by benchmark rankings. Two people familiar with Argon said it appeared affected.
When can developers use Gemini 4 Argon?
Initial access runs through Google's Fairwind cybersecurity program, and Google is participating in the U.S. government's voluntary pre-release model access process. Paid API customers and Google AI Ultra subscribers are next in line, but Google has given no release date beyond "as soon as possible."
AI-generated summary, reviewed by an editor. More on our AI guidelines.
Related stories



Free AI briefing · Weekdays
The AI stories that matter, sourced and explained.
Join the Morning Briefing. It goes out every weekday at 4:45 a.m. Pacific, with later sends for the East Coast, Berlin and Tokyo.
Free when you sign up: The Paperclip Compendium, our tested guide to running AI agents.
Free. Unsubscribe in one click.
IMPLICATOR