SpaceXAI released Grok 4.6 for long-running coding and knowledge-work agents, and independent testing scored it 61 on the Artificial Analysis Intelligence Index. Its output tokens cost one-fifth as much as GPT-5.6 Sol’s in the cited standard API modes. That gap matters for long-running agents because their work can span many prompts and tool calls before completion.

The same testing measured output at about 85.8 tokens per second, above the 71.2-token median for comparable reasoning models in its price tier. Its 32.30-second time to first token was far slower than the 2.88-second median.

What Changed

AI-generated summary, reviewed by an editor. More on our AI guidelines.

A frontier score

Artificial Analysis ran Grok 4.6 high through version 4.1.1 of its index, which combines nine evaluations of reasoning, knowledge, mathematics and coding. The August 12 results put Grok level with GPT-5.6 Sol max, below Claude Opus 5 max at 63 and Fable 5 max with fallback at 62.

Grok 4.5 scored 56 on the same index, giving its successor a five-point gain over one model generation.

The cost per task

Below 200,000 prompt tokens, the SpaceXAI API price is $2 per million input tokens, $0.50 per million cached input tokens and $6 per million output tokens. The captured standard-mode price for GPT-5.6 Sol is $5 for input and $30 for output. The one-fifth comparison applies to those output rates, not every model variant or provider.

Artificial Analysis measured an average cost of $0.84 for each Intelligence Index task and placed Grok 4.6 on its intelligence-versus-cost Pareto frontier. The token meter also depends on how many steps an agent takes. An agent keeps working across a sequence of prompts, tools, files and checks instead of returning one answer. Each turn can feed more text into the model and generate another response, adding tokens to the run before the task is finished. Fewer turns can therefore reduce the amount of input and output billed for the completed workload.

On the AA-Briefcase knowledge-work test, Grok completed workloads in about 53 turns and used about 0.5 billion input tokens on average. Claude Opus 5 max took about 103 turns and 2 billion input tokens. Grok’s Elo score was 1577 on that test.

The comparison has limits

Grok did not lead every test. In SpaceXAI’s August 12 comparison table, it scored 26% on Terminal-Bench v3.0, behind GPT-5.6 Sol max at 34.6% and Fable 5 max at 34.1%.

That table combines competitors’ best self-reported or publicly available scores. Differences in settings, harnesses and tool access mean it is not a controlled four-model test. Artificial Analysis’s Intelligence Index is separate and independently run.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

Long prompts also change the bill. Grok has a 500,000-token context window, but once a prompt reaches 200,000 tokens, input, cached-input and output rates double to $4, $1 and $12 per million tokens. The higher band applies to every token in the request.

The production test

SpaceXAI said the model received longer supplemental training than Grok 4.5, using curated model-generated reasoning and engineering data, regenerated supervised fine-tuning trajectories and reinforcement learning for coding and knowledge work.

Grok 4.6 became available August 12 through Cursor, Grok Build, the SpaceXAI API, OpenRouter, Vercel and Cloudflare. Cursor and Grok Build include twice the usual usage during the first week.

Controlled tests cannot establish the same savings in production. Harness design, prompts, tool calls, caching, retries and task choice all change cost. Will the measured turn-and-token advantage survive once those factors enter the loop?

Frequently Asked Questions

What is Grok 4.6?

Grok 4.6 is SpaceXAI's model for long-running coding and knowledge-work agents. It became available on August 12 through Cursor, Grok Build, the SpaceXAI API and several partners.

How did Grok 4.6 score against GPT-5.6 Sol?

Artificial Analysis scored Grok 4.6 high at 61 on its Intelligence Index, level with GPT-5.6 Sol max. Grok remained behind rival models on some individual tests, including Terminal-Bench v3.0.

How much does the Grok 4.6 API cost?

Below 200,000 prompt tokens, pricing is $2 per million input tokens, $0.50 per million cached-input tokens and $6 per million output tokens.

What happens to Grok 4.6 pricing on long prompts?

Once a prompt reaches 200,000 tokens, input, cached-input and output prices double to $4, $1 and $12 per million tokens. The higher rate applies to every token in the request.

Do the benchmark savings guarantee lower production costs?

No. Harness design, prompts, tool calls, caching, retries and task choice can change the number of tokens and turns an agent uses in production.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Nvidia Says NeMo Switchyard Cuts AI Agent Benchmark Costs by About 60%
Nvidia released Nemotron 3.5 Lightning and open-source NeMo Switchyard on August 11, 2026, saying its internal benchmark cut AI agent task costs by about 60%. The library selects a model for each work
Meta Opens Muse Glimmer as Zuckerberg Presses U.S. AI Policy Shift
On Aug. 10, 2026, Meta released Muse Glimmer, a 30-billion-parameter model whose weights developers can download. It was built for agents that run on a Mac or PC. The release carried a policy argumen
DeepSeek's Retrained V4 Flash Scores 50 on Independent Intelligence Index
DeepSeek released V4-Flash-0731 and placed its official API in public beta on July 31, while Artificial Analysis scored the model 50 on its Intelligence Index, 10 points above the April preview. The m
AI News

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai