SpaceXAI released Grok 4.6 for long-running coding and knowledge-work agents, and independent testing scored it 61 on the Artificial Analysis Intelligence Index. Its output tokens cost one-fifth as much as GPT-5.6 Sol’s in the cited standard API modes. That gap matters for long-running agents because their work can span many prompts and tool calls before completion.
The same testing measured output at about 85.8 tokens per second, above the 71.2-token median for comparable reasoning models in its price tier. Its 32.30-second time to first token was far slower than the 2.88-second median.
What Changed
- Grok 4.6 scored 61 on the independently run Artificial Analysis Intelligence Index, level with GPT-5.6 Sol max.
- Below 200,000 prompt tokens, Grok 4.6 costs $2 per million input tokens and $6 per million output tokens.
- Artificial Analysis measured $0.84 per Intelligence Index task and about 53 turns per AA-Briefcase workload.
- At 200,000 prompt tokens, Grok 4.6's input, cached-input and output rates double for the entire request.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
A frontier score
Artificial Analysis ran Grok 4.6 high through version 4.1.1 of its index, which combines nine evaluations of reasoning, knowledge, mathematics and coding. The August 12 results put Grok level with GPT-5.6 Sol max, below Claude Opus 5 max at 63 and Fable 5 max with fallback at 62.
Grok 4.5 scored 56 on the same index, giving its successor a five-point gain over one model generation.
The cost per task
Below 200,000 prompt tokens, the SpaceXAI API price is $2 per million input tokens, $0.50 per million cached input tokens and $6 per million output tokens. The captured standard-mode price for GPT-5.6 Sol is $5 for input and $30 for output. The one-fifth comparison applies to those output rates, not every model variant or provider.
Artificial Analysis measured an average cost of $0.84 for each Intelligence Index task and placed Grok 4.6 on its intelligence-versus-cost Pareto frontier. The token meter also depends on how many steps an agent takes. An agent keeps working across a sequence of prompts, tools, files and checks instead of returning one answer. Each turn can feed more text into the model and generate another response, adding tokens to the run before the task is finished. Fewer turns can therefore reduce the amount of input and output billed for the completed workload.
On the AA-Briefcase knowledge-work test, Grok completed workloads in about 53 turns and used about 0.5 billion input tokens on average. Claude Opus 5 max took about 103 turns and 2 billion input tokens. Grok’s Elo score was 1577 on that test.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
The comparison has limits
Grok did not lead every test. In SpaceXAI’s August 12 comparison table, it scored 26% on Terminal-Bench v3.0, behind GPT-5.6 Sol max at 34.6% and Fable 5 max at 34.1%.
That table combines competitors’ best self-reported or publicly available scores. Differences in settings, harnesses and tool access mean it is not a controlled four-model test. Artificial Analysis’s Intelligence Index is separate and independently run.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Long prompts also change the bill. Grok has a 500,000-token context window, but once a prompt reaches 200,000 tokens, input, cached-input and output rates double to $4, $1 and $12 per million tokens. The higher band applies to every token in the request.
The production test
SpaceXAI said the model received longer supplemental training than Grok 4.5, using curated model-generated reasoning and engineering data, regenerated supervised fine-tuning trajectories and reinforcement learning for coding and knowledge work.
Grok 4.6 became available August 12 through Cursor, Grok Build, the SpaceXAI API, OpenRouter, Vercel and Cloudflare. Cursor and Grok Build include twice the usual usage during the first week.
Controlled tests cannot establish the same savings in production. Harness design, prompts, tool calls, caching, retries and task choice all change cost. Will the measured turn-and-token advantage survive once those factors enter the loop?
Frequently Asked Questions
What is Grok 4.6?
Grok 4.6 is SpaceXAI's model for long-running coding and knowledge-work agents. It became available on August 12 through Cursor, Grok Build, the SpaceXAI API and several partners.
How did Grok 4.6 score against GPT-5.6 Sol?
Artificial Analysis scored Grok 4.6 high at 61 on its Intelligence Index, level with GPT-5.6 Sol max. Grok remained behind rival models on some individual tests, including Terminal-Bench v3.0.
How much does the Grok 4.6 API cost?
Below 200,000 prompt tokens, pricing is $2 per million input tokens, $0.50 per million cached-input tokens and $6 per million output tokens.
What happens to Grok 4.6 pricing on long prompts?
Once a prompt reaches 200,000 tokens, input, cached-input and output prices double to $4, $1 and $12 per million tokens. The higher rate applies to every token in the request.
Do the benchmark savings guarantee lower production costs?
No. Harness design, prompts, tool calls, caching, retries and task choice can change the number of tokens and turns an agent uses in production.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR