IMPLICATOR​.ai
Morning Briefing · From San Francisco

 

Friday, September 4, 2026
11 stops

From San Francisco

1
The Editorial
 

Morning, humans.

Today's three picks come down to one question: who ran the test?

OpenAI put 99.9% on the front of the Astra launch. ARC Prize built that benchmark, ran the model on its own neutral harness, got 62.7%, and is not claiming AGI.

All 20 G-20 members signed the US-drafted Carolina Principles in Chapel Hill, which ask governments to save new rules for novel cases. Brussels spent it mailing questions to 30-plus AI companies.

And Heretic, a Python tool with 30,276 stars, took Gemma 3 12B from refusing 97 of 100 prompts down to three. The refusal count turns out to be a setting.

Stay curious,

Marcus Schuler

Good briefing? Pass it on
X · LinkedIn · Bluesky · Email
 
2
The Big Story
OpenAI's headline Astra score came from a harness ARC Prize did not run.

OpenAI released GPT-6 Astra on September 3 and led its launch case with a 99.9% score on ARC-AGI-3, a benchmark it did not build.

ARC Prize, which did build it, ran the same model on its own provider-neutral harness and recorded 62.7%. It published the evaluation the same day and wrote that it is not claiming Astra is AGI.

Standard API pricing is $10 per million input tokens and $50 per million output, double GPT-5.6 Sol's input rate. Artificial Analysis put Astra at 61 on its Intelligence Index, level with Sol. Access is off by default and enterprise administrators must switch it on.

Why This Matters:

Reality Check
What's confirmed: OpenAI released Astra on September 3 at $10 per million input tokens and $50 per million output. ARC Prize scored it 62.7% on its Standard harness at maximum effort.
What's implied (not proven): That a 99.9% result marks the arrival of general intelligence. ARC Prize wrote that saturating the benchmark would not be proof of AGI.
What could go wrong: A team budgets a rollout against the 99.9% figure, then meets 62.7%-shaped behavior on its own unfamiliar internal tools.
What to watch next: Whether GDPval, OpenAI's own benchmark for economically valuable work, turns up in follow-up materials. It is absent from the launch.
Read the full story →
 
3
Also Today
The G-20 agreed to hold back new AI rules while Brussels sent out 30 letters.

All 20 G-20 members adopted the US-drafted Carolina Principles at the Chapel Hill innovation ministerial on September 2.

The framework asks governments to reserve new AI regulation for "novel considerations," which the announcement never defines. Commerce Secretary Howard Lutnick said China signed with the rest. The European Commission spent the same meeting sending information requests to more than 30 AI companies under the AI Act, so the consensus lands with Brussels already moving.

Read our coverage →
 
4
Repo Spotlight
 

Worktrunk is a Rust CLI that addresses git worktrees by branch name, so one command creates the tree, moves into it and starts the agent. It adds hooks on create and merge, a unique dev-server port per tree, build-cache copying between trees, and a list view carrying CI status per branch.

It earns its keep once a second coding agent is running, where the cost stops being the agent and starts being the branch juggling around it.

wt switch -c -x claude feat

Four more from this week's Repo Radar, including JetBrains shipping its Go idiom rules as an agent skill.

Worktrunk on GitHub →
 
5
The Outside Read
 

CNBC maps the Chinese supply-chain dependencies hiding beneath America's AI data-center buildout.

Chinese suppliers account for nearly 30% of certain United States transformer and switchgear categories and roughly two-thirds of global optical-transceiver units. Wood Mackenzie estimates that 2026 shortages already equal 15% of power-transformer demand and 8% of substation demand, making new restrictions a near-term cost risk.

Read it at CNBC →
 
6
The One Number
 
$26,098
What OpenAI's headline benchmark cost on a harness the vendor did not shape. ARC Prize ran Astra at maximum effort on its provider-neutral setup and recorded 62.7% for that money. The vendor-adapted run scored 99.9% and cost $18,817, so the neutral test came in 39% more expensive and 37 points worse.
Source: ARC Prize, September 3, 2026
 
7
Today's Headlines
 
Tuesdays go deeper.  Sign up for Implicator PRO for the weekly Tuesday deep dive on deploying AI where it pays. $8 a month, $89 a year.
 
8
The 5-Minute Skill
 

Reference calls drift into polite praise when the questions are broad. Turn the candidate's strongest claim into a short verification sequence.

Your raw input: the candidate's résumé and your interview notes, plus the job's most important outcome and the reference's role if you know it.

The prompt:

Use the material above to identify the candidate claim that matters most to the role. Create a ten-minute reference-call script that tests this claim without revealing the answer I hope to hear. Begin with an open question asking the reference to recall a specific episode. Follow with questions that establish the candidate's personal contribution and the observable result. Include one question about where the claim may overstate the candidate's role. End with a neutral sentence I can use to check my interpretation with the reference. Do not invent details.

Why this works: anchoring the call to one claim blocks generic praise from becoming evidence. An open recollection question reduces priming, and the interpretation check catches misunderstandings before they enter the hiring record.

What to use: ChatGPT handles the concise call script well. Claude is a reliable fallback when the interview notes run long.

 
9
What To Watch Next
 
SAT 9/5
Fairs: IFA Berlin begins its first full public day at Messe Berlin, where organizers host more than 1,900 consumer-tech, robotics and AI brands from 10 a.m. to 6 p.m. CEST.
MON 9/7
Finance: The NYSE and Nasdaq close for the Labor Day holiday.
MON 9/7
AI: ECML PKDD organizers open the five-day European machine-learning and data-mining conference in Naples.
THU 9/10
Policy: The ECB announces its monetary policy decision in Berlin, followed by President Christine Lagarde's press conference at 2:45 p.m. CEST.
FRI 9/11
Finance: The Bureau of Labor Statistics reports August consumer-price data at 8:30 a.m. ET.
 
10
AI Image of the Day
 
Portrait of a man with pale violet reptilian scales, ridged spines running over his brow and shoulders, and a large spiral horn curling behind one ear, against a dark teal background
Credit: Ideogram
Prompt: Change to man
 
11
The Rausschmeisser*
A Python script took a model's refusals from 97 in 100 down to three.

Heretic, a Python tool now at 30,276 stars, strips safety alignment out of an open-weights model in a single command. On Gemma 3 12B it took refusals from 97 of 100 prompts down to three (The Implicator, September 3, 2026).

Our take: The method is the part worth sitting with. Heretic runs directional ablation under an Optuna search that co-minimizes refusals against KL divergence from the original, which is a careful way of saying it finds the cheapest edit that removes the conscience without damaging the brain.

Every lab shipping open weights has now been handed a receipt. Ninety-seven down to three, one command, no fine-tuning budget, and a parameter search that did the thinking. The alignment was real work. It also turned out to be a setting.

*German for the last song of the night, the one that clears the room.

FREE · ABOUT FIVE MINUTES

Stop chasing the AI news cycle.
Implicator cuts the flood to the developments worth your attention, with concise reporting on what changed and why it matters. Delivered every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.
Get the free briefing →
From San Francisco. No spam. Unsubscribe anytime.
Morning Briefing

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai