In April 2026, Uber chief technology officer Praveen Neppalli Naga sat down to demonstrate the company’s AI coding tools. Over the next two hours, he used $1,200 worth of tokens, the units providers charge for when their models process and generate text.

By then, Uber had already exhausted its entire 2026 AI budget, about four months into the year.

AI work is billed by the token, and a model’s rate card is a poor guide to its bill because token consumption varies between models and even between runs of the same model. Gemini 3 Flash was listed 80% cheaper than GPT-5.4 at May 1, 2026 prices, but cost 38% more across the study’s tasks. Only 11% of 396 enterprises surveyed in April and May 2026 could forecast AI costs within 10%, down from 15% in 2025, in a report released July 29 by Benchmarkit and Mavvrik, which sells AI cost-management software; the figures are self-reported.

“I’m back to the drawing board, because the budget I thought I would need is blown away already,” Naga said.

The Breakdown

AI-generated summary, reviewed by an editor. More on our AI guidelines.

The study

Lingjiao Chen, a researcher at Stanford University and Microsoft Research, tested whether listed API prices predict what a model actually costs to run, with co-authors from Carnegie Mellon, UC Berkeley and Microsoft Research. Their paper, “The Price Reversal Phenomenon,” first appeared on March 25, 2026, and was revised May 28.

The revised study tested eight frontier reasoning models across 12 tasks. Of 336 pairwise cost comparisons, 106, or 32%, showed the model with the lower listed price costing more in total.

No model in the study was consistently the cheapest or the most expensive across all of its benchmarks; the ranking changed from task to task. Listed price was defined as input plus output rates.

Thinking tokens are the hidden reasoning a model writes before answering. Their volume helped explain the reversals. On one MMLU-Pro problem in the study, Gemini 3 Flash consumed more than 60,000 thinking tokens; GPT-5.4 solved the same problem with 25.

For tasks requiring a model to interact repeatedly with tools or a computer environment, the number of turns also drove costs. Each turn can include earlier conversation history as input, adding another charge as the model continues working.

“The practical takeaway is clear,” Chen said. “Price alone should not be used to infer which model is actually cheaper.”

On one prompt in the researchers’ data, the cheaper model cost 14 times as much and still failed. Gemini 3.1 Pro finished in 85 steps for about $1. Gemini 3 Flash went through nearly 1,000 steps, accumulated $14 in token charges and failed.

The researchers published their per-run cost data and code so companies could repeat the comparison on their own workloads.

Same prompt, different bill

Choosing a model that used fewer tokens in one test did not guarantee a repeatable bill. In the May revision, repeated runs of an identical query on the same model varied by up to 9.7 times between the cheapest and most expensive run.

The paper describes an irreducible noise floor, a baseline of randomness that no forecaster can get below: models can follow different reasoning paths even when their inputs stay fixed. That makes predicting the cost of an individual query difficult. The variation cannot be removed by re-prompting.

A follow-up analysis of the researchers’ data on two selected programming prompts showed variation across models from Anthropic, Google and OpenAI. Anthropic models were added through additional data collection after the original study, which had not included them for this measure. Each model received the same prompt five times, with a different cost on each run. Unsuccessful runs consumed tokens and incurred charges too.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

That uncertainty has reached customers building their own software. Mazda Marvasti, co-founder and chief executive of Amberd.ai, said some customers abandoned internally built automation tools because they could not forecast or justify the costs. Amberd.ai builds on private, open-source models run on bare-metal servers using QumulusAI hardware; his remarks come from a Futurum report sponsored by QumulusAI.

“When they start deploying it throughout the organization, the cost starts skyrocketing because it’s a useful tool that somebody built, but it’s now priced on a variable basis,” he said.

The other side

“Some prompt-level fluctuation is inherent to AI, and our testing shows this averages out across a high volume of real-world, diverse workloads,” a Google spokeswoman said.

“Total costs depend on many factors for a given task,” she said, “which can make it hard to forecast new and evolving technology with precision.” Google offers spending caps and flexible pricing. Google, Anthropic and OpenAI have also released newer models that perform better on industry benchmarks since the models in the study were tested.

The paper uses a single pricing snapshot from May 1, 2026, runs each model at one reasoning setting and measures cost separately from answer quality. The cost comparison does not account for whether the answer was right, so a cheap model that fails and an expensive one that succeeds are compared on cost alone.

Budget limits

By June 2026, Uber had capped agentic coding tools at $1,500 per employee per month, per tool.

Inside Uber, chief operating officer Andrew Macdonald described the difficulty of connecting usage measures to what customers receive. On the Rapid Response podcast, he said: “It's very hard to draw a line between one of those stats and 'OK, now we're actually producing like 25% more useful consumer features.'”

Frequently Asked Questions

What is the price reversal phenomenon?

It is the finding that a reasoning model with a lower listed API price can cost more in total to run than a pricier model. In the study by Lingjiao Chen and co-authors, this happened in 106 of 336 pairwise comparisons, or 32%, across eight frontier models and 12 tasks.

Why can a cheaper AI model end up costing more?

Models consume very different numbers of tokens on the same work. On one MMLU-Pro problem, Gemini 3 Flash used more than 60,000 thinking tokens while GPT-5.4 used 25. In tasks with tools, the number of interaction turns also drove costs.

Does the same model cost the same every time?

No. Repeated runs of an identical query on the same model varied by up to 9.7 times in cost. The paper says this variation cannot be removed by re-prompting, which makes the cost of an individual query difficult to predict.

What did Google say about the findings?

A Google spokeswoman said some prompt-level fluctuation is inherent to AI and that Google's testing shows it averages out across a high volume of diverse workloads. She said Google offers spending caps and flexible pricing.

What are the study's limits?

It uses a single pricing snapshot from May 1, 2026, runs each model at one reasoning setting, and measures cost separately from answer quality, so a cheap model that fails and an expensive one that succeeds are compared on cost alone.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

OpenAI’s Sol costs half as much as Opus 5.5; Trump renames AI super intelligence
IMPLICATOR .ai Morning Briefing · From San Francisco   Wednesday, September 23, 2026 10 stops From San Francisco 1 The Editorial   Morning, humans. Today’s theme: w
Palo Alto Networks CEO Says AI Token Costs Must Fall Up to 90%
Palo Alto Networks CEO Nikesh Arora said on CNBC on Thursday that AI token costs need to fall as much as 90% to support large-scale enterprise adoption. He called OpenAI CEO Sam Altman’s claim that th
Anthropic shifts enterprise billing to per-token pricing. The flat-fee era is over.
Anthropic has restructured its enterprise plan to bill Claude, Claude Code, and Cowork usage separately from seat fees, moving its largest business customers to per-token pricing at standard API rates
AI News

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai