From San Francisco
1 |
The Editorial |
Marcus here.
The price on an AI label gives you one number, while the bill has room for a surprise.
Stanford's study puts eight reasoning models through twelve tasks at May 1, 2026 prices. The shopping decision looks easy until the meter starts counting what the label leaves out.
SemiAnalysis’s October 5, 2026 study tests paid Claude and ChatGPT accounts. OpenAI's mid-tier plans lose the comparison, with War and Peace supplying rather more than bedtime reading.
Our PRO piece asks which AI memory tool belongs in your coding agent. Five plugins want the job, and Claude Code already brings memory of its own.
Stay curious,
Marcus Schuler
192 prompts | 16 categories | 384 test runs |
2 |
The Big Story |
A Stanford-led study found that the model with the lower listed price cost more in 106 of 336 pairwise comparisons across eight models, or 32%.
Gemini 3 Flash was listed 80% cheaper than GPT-5.4 at May 1, 2026 prices but cost 38% more across the tasks. Thinking tokens helped explain the reversals: on one MMLU-Pro problem Gemini 3 Flash used more than 60,000 of them and GPT-5.4 used 25.
Uber exhausted its 2026 AI budget by April. CTO Praveen Neppalli Naga spent $1,200 in a two-hour demo. “I'm back to the drawing board, because the budget I thought I would need is blown away already,” Naga said.
Why This Matters:
- Teams buying AI by listed token rates can underestimate spending when a cheaper model uses more tokens to complete their tasks.
- Procurement decisions need repeated tests on actual workloads, because cost rankings change by task and identical queries can produce different bills.
3 |
Also Today |
SemiAnalysis's October 5, 2026 test compares Opus 5.5 with GPT-6.1 Sol on same-price plans, measuring API-equivalent dollars at list prices.
OpenAI halved its $200 Pro allowance; Thibault Sottiaux, who leads Codex there, wrote that it “will net out at half the dollar in API spend compared to the old Pro $200 plan.” The result covers mid-tier models; finished tasks are unmeasured. Subscribers need task costs before choosing.
4 |
Worth the Paywall |
Choose a memory plugin for your coding agent with its costs and data storage in view.
Claude Code already loads up to 200 lines or 25KB of memory per session. Tuesday’s PRO guide helps coding-agent users decide whether a plugin earns its place.
5 |
The Outside Read |
Simon Willison argues that usage-based services need hard monthly budget caps by default because coding agents make spending easy.
“Soft caps, 'after $X/month, send me a warning email', will not cut it,” he writes. Google Cloud launched Spend Caps in July, and AWS began offering project spend limits to a limited number of customers on September 16, 2026.
6 |
The One Number |
7 |
Today's Headlines |
- OpenAI will watermark ChatGPT and Codex text in the EU over the coming weeks, with optional API watermarking available from October 5, 2026.
- Reflection AI unveils Beam, a 501-billion-parameter open-weight model, claiming GLM-5.2 reasoning performance at three to four times less compute, with no independent evaluation yet.
- Nokia CEO Justin Hotard says customers would probably build twice as fast if supply allowed, amid memory-chip and energy shortages.
- New York City Council examines an AI shutdown requirement that puts validation duties on companies, while a civic group identifies a practical gap involving open-weight models.
- OpenAI is in talks for a $30 billion funding round with UAE funds and BlackRock, Bloomberg reported on October 5, 2026.
Tue 10/6 |
Economy: the BEA releases the August U.S. trade report on goods and services at 8:30 a.m. Eastern. |
Tue 10/6 |
Tech: AIME-Con, the conference on AI in measurement and education, holds keynotes and sessions in Pittsburgh on Oct. 6 and 7. |
Wed 10/7 |
Fed: the central bank publishes the minutes of its September 15-16 meeting at 2 p.m. Eastern. |
Thu 10/8 |
Tech: the IAPP privacy, security, risk and AI governance conference opens in Seattle and runs through Oct. 9. |
8 |
The 5-Minute Skill |
A polished proposal can hide the assumptions carrying its conclusion. This move turns the document into a short test list before approval.
Your raw input:
Paste the proposal, business case, forecast, or strategy memo, plus the decision deadline and any source notes attached to it.
The prompt:
Why this works: The model must separate explicit claims from inferred ones and tie each assumption to evidence in the document. Ranking dependence and naming the fastest check converts critique into work an executive can assign today.
What to use: Use your strongest reasoning model for a long or numerical proposal. A standard general-purpose model is enough for a short memo.
9 |
AI Toolbox |
Edyt is a desktop text assistant for macOS and Windows that rewrites selected text in place inside email, documents, chat and other editable fields.
It avoids the copy-paste loop of a separate chatbot: its free plan costs $0, while Pro costs $12 a month or $8.25 a month with annual billing, all checked on September 26, 2026.
How to use it:
- Download the stable Edyt build for your operating system and install it.
- Open Edyt and, on a Mac, allow Accessibility when the app asks.
- Select the text you want to revise in any editable field.
- Press
Ctrl+Shift+C; Edyt replaces the selection with its rewrite in place.
Where it falls short: Edyt works on demand and does not provide always-on proofreading while you type.
10 |
AI Image of the Day |
11 |
The Rausschmeisser* |
In an interview with POLITICO’s Decoded previewed October 4, 2026, Sam Altman said, “we believe that the world should accept some bad things happening for the benefits of this technology and people having the agency.”
Our take: Altman has volunteered the world for a tolerance exercise. Agents built on OpenAI’s platform may already have conducted unauthorized intrusions or caused harm at more than 100 organizations, so the participants appear to have been enrolled ahead of the interview. Consent takes longer than a demo.
The proposed bargain comes from a CEO who rejects a trade promising “no major hacks” and “zero scams.” Anyone whose organization ends up on the damage list gets to appreciate the benefits from close range. Perhaps the next subscription tier can include a receipt for the bad things. Accounting departments do like supporting documents.
FREE WEEKDAY MORNING BRIEFING
Free AI briefing · Weekdays
The AI stories that matter, sourced and explained.
Join the Morning Briefing. It goes out every weekday at 4:45 a.m. Pacific, with later sends for the East Coast, Berlin and Tokyo.
Free when you sign up: The Paperclip Compendium, our tested guide to running AI agents.
Free. Unsubscribe in one click.
IMPLICATOR