From San Francisco
1 |
The Editorial |
Marcus here.
Three times today the stated number and the real one part company.
Apple's new Mac Studio processes prompts up to four times faster, Apple says. That multiplier covers half a request, and the other half decides how fast your answer arrives.
In Keelung, prosecutors charged nine people after 74 Nvidia B300 servers reached Chinese buyers. An inspector had found the declared site could not run the order. An Nvidia manager cleared it anyway.
And Ramp's card data has Fable 5 at 6% of Anthropic tokens in July but 11.4% of the spend. It costs about double per token.
Stay curious,
Marcus Schuler
2 |
The Big Story |
Apple's new Mac Studio puts matrix accelerators in every GPU core and scales to 512GB of unified memory, aimed at the pause before a model starts answering.
Pre-orders opened Aug. 25 and deliveries begin Sept. 22. The M5 Max starts at $2,499 and the M5 Ultra at $5,499, the highest starting prices the line has carried.
Apple's own July tests put M5 Ultra prompt processing at up to four times the M3 Ultra result. That covers prefill, the compute-bound phase that decides how long a long prompt sits before the first token appears. Generation afterward runs on memory bandwidth, which rose by what Apple describes as 50 percent.
Why This Matters:
- Anyone sizing a local-inference machine now has to ask which half of the request their own workload actually spends its time in.
- The 512GB configuration cannot be ordered yet and arrives in late October, which makes the top tier a fourth-quarter decision.
3 |
Also Today |
Keelung prosecutors indicted nine people over 74 Nvidia B300 servers that reached Chinese buyers on false end-user documents.
A site inspection had already found that the declared facility could not run an order that size. An Nvidia manager approved the shipment anyway. Export controls that rest on paperwork and a vendor's own sign-off leave the enforcement burden with the company earning the sale.
4 |
The Outside Read |
Ramp Economics Lab uses company payment data to show where buyers have started rejecting frontier-model prices.
In July 2026, Fable 5 accounted for 6% of Anthropic tokens and 11.4% of Anthropic model spend in Ramp's sample, while GPT-5.6 Sol accounted for 25% of OpenAI tokens and 23% of OpenAI model spend. The data puts a buyer-side price ceiling beneath the model race.
5 |
The One Number |
6 |
Today's Headlines |
- Broadcom is in talks with lenders to raise more than $60 billion in debt to finance AI chips for Anthropic and others, with Apollo and Blackstone discussing a role.
- Fasset reached a $1 billion valuation on a $68 million Series C led by SBI Group, and says it now processes more than $40 billion in annualized volume.
- Nvidia detailed its 88-core Vera CPU at Hot Chips on Aug. 24, where the published SPECrate 2026 integer score came in 3% above AMD's EPYC 9755.
- Ten local MCP servers turn a laptop into a workbench with sources of truth beyond files and Git, provided each runs behind a narrow identity.
- Supacode turns Git worktrees into a command center for coding agents, though its early traction figures leave the adoption question open.
Wed 8/26 |
Earnings: Nvidia reports fiscal second-quarter 2027 results after the close, with the call at 5 p.m. ET. |
Wed 8/26 |
Economy: the Bureau of Economic Analysis releases its second estimate of second-quarter GDP and July PCE inflation, both at 8:30 a.m. ET. |
Thu 8/27 |
Policy: the Jackson Hole symposium opens, on financial innovation and its implications for payments and policy. |
Thru 8/30 |
Fairs: Gamescom runs in Cologne through Sunday, the largest games trade fair of the year. |
7 |
The 5-Minute Skill |
Extract the decisions a meeting did not make. After a crowded meeting, people can leave with different ideas about what was settled. This turns the record into a decision check.
Your raw input: Gather the agenda, transcript or notes, and any decision list already circulated. Keep speaker names and timestamps.
The prompt:
Why this works: The model must tie every status to the record, which limits confident invention. Fixed labels separate actual commitments from ideas that merely received airtime.
What to use: A long-context model such as GPT-5.6 Sol for a full transcript. Any strong general model is fine for short notes.
8 |
AI Profile |
Convex sells a managed backend for web applications built by developers and coding agents. Its arguable position is that AI-written software needs a database and runtime that enforce correctness because fewer engineers will inspect every generated line.
Founders: Jamie Turner, James Cowling and Sujay Jayakar founded Convex in 2020 after working together at Dropbox. Turner led storage and database engineering, Cowling was a senior principal engineer on multi-exabyte storage, and Jayakar was a principal engineer who led a rewrite of Dropbox's sync engine.
Product: Convex combines a reactive database, serverless TypeScript functions, synchronization, search, file storage and workflows. Developer teams pay for the hosted service; its Business plan starts at $2,500 a month. Convex said it had nearly 10,000 paying teams in April 2026 and dozens of enterprise-plan customers by August 4. Its site names OpenAI, Tripadvisor, Zapier and Reducto among users.
Financing: Convex has raised $110.5 million, including a $57 million Series B led by Insight Partners on August 4, 2026. Etna Labs, Spark Capital, Andreessen Horowitz and Justin Kan also participated. The company did not disclose a valuation.
The risk: Supabase and Google's Firebase have larger developer ecosystems, while Supabase offers portable PostgreSQL. Convex's non-SQL model can make adoption a rewrite rather than a simple migration, weakening its pitch for established enterprise systems.
9 |
AI Image of the Day |
10 |
The Rausschmeisser* |
Ox Alpha turned up as a free, unattributed model with a million-token context window and a claim of 100 trillion tokens of daily capacity. Researchers then ran 95 vocabulary probes against Zhipu's released GLM-5 tokenizer. It matched on 95 of 95, and separate serving-layer evidence points toward Z.ai. (Implicator, August 23, 2026)
Our take: Ninety-five out of ninety-five is not a clue. It is a signed confession with the notary's stamp still wet. A tokenizer is the one part of a model you cannot casually redraw, which is why fingerprinting it works, and which is why anyone shipping a model they want kept anonymous should have started there.
The stealth launch has become a genre with rules nobody follows. Pick a name from no known language, withhold the lab, claim a capacity figure too large to check, then wait for the leaderboard chatter. And leave the vocabulary file exactly as it came. The mystery held for as long as it took one researcher to run a script.
IMPLICATOR