> ## Content Index
> Fetch the complete content index at: https://www.implicator.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Astra's AGI Benchmark Reads 99.9% or 62.7%; G-20 Backs Lighter AI Rules
- URL: https://www.implicator.ai/astras-agi-benchmark-reads-99-9-or-62-7-g-20-backs-lighter-ai-rules/
- Published: 2026-09-04T11:45:41.000Z
- Updated: 2026-09-04T11:45:41.000Z
- Description: Heretic took a model's refusals from 97 in 100 down to three. One command, no fine-tuning budget.
- Author: Marcus Schuler
- Tags: Morning Briefing

**IMPLICATOR** **​.ai**

Morning Briefing · From San Francisco

Friday, September 4, 2026

11 stops

From San Francisco

| 1 | The Editorial |
| - | ------------- |

*Morning, humans.*

*Today's three picks come down to one question: who ran the test?*

*OpenAI put 99.9% on the front of the Astra launch. ARC Prize built that benchmark, ran the model on its own neutral harness, got 62.7%, and is not claiming AGI.*

*All 20 G-20 members signed the US-drafted Carolina Principles in Chapel Hill, which ask governments to save new rules for novel cases. Brussels spent it mailing questions to 30-plus AI companies.*

*And Heretic, a Python tool with 30,276 stars, took Gemma 3 12B from refusing 97 of 100 prompts down to three. The refusal count turns out to be a setting.*

*Stay curious,*

*Marcus Schuler*

Good briefing? Pass it on

[X](https://twitter.com/intent/tweet?text=Strategic%20AI%20news%20from%20San%20Francisco.%20No%20hype%2C%20just%20what%20moved%2C%20who%20won%2C%20and%20why%20it%20matters.&url=https%3A%2F%2Fwww.implicator.ai%2F%3Futm%5Fsource%3Dnewsletter%26utm%5Fmedium%3Dshare%26utm%5Fcampaign%3Dreader%5Fshare%26utm%5Fcontent%3Dtwitter&ref=implicator.ai) · [LinkedIn](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fwww.implicator.ai%2F%3Futm%5Fsource%3Dnewsletter%26utm%5Fmedium%3Dshare%26utm%5Fcampaign%3Dreader%5Fshare%26utm%5Fcontent%3Dlinkedin&ref=implicator.ai) · [Bluesky](https://bsky.app/intent/compose?text=Strategic%20AI%20news%20from%20San%20Francisco.%20No%20hype%2C%20just%20what%20moved%2C%20who%20won%2C%20and%20why%20it%20matters.%20https%3A%2F%2Fwww.implicator.ai%2F%3Futm%5Fsource%3Dnewsletter%26utm%5Fmedium%3Dshare%26utm%5Fcampaign%3Dreader%5Fshare%26utm%5Fcontent%3Dbluesky&ref=implicator.ai) · [Email](mailto:?subject=A%20newsletter%20I%20think%20you%27d%20like&body=This%20is%20one%20of%20the%20few%20AI%20newsletters%20I%20actually%20read.%20Strategic%2C%20no%20hype.%0A%0ASubscribe%20free%3A%20https%3A%2F%2Fwww.implicator.ai%2Fsubscribe%2F%3Futm%5Fsource%3Dnewsletter%26utm%5Fmedium%3Dshare%26utm%5Fcampaign%3Dreader%5Fshare%26utm%5Fcontent%3Demail)

| 2 | The Big Story |
| - | ------------- |

OpenAI's headline Astra score came from a harness ARC Prize did not run.

**OpenAI released GPT-6 Astra on September 3 and led its launch case with a 99.9% score on ARC-AGI-3, a benchmark it did not build.**

ARC Prize, which did build it, ran the same model on its own provider-neutral harness and recorded 62.7%. It published the evaluation the same day and wrote that it is not claiming Astra is AGI.

Standard API pricing is $10 per million input tokens and $50 per million output, double GPT-5.6 Sol's input rate. Artificial Analysis put Astra at 61 on its Intelligence Index, level with Sol. Access is off by default and enterprise administrators must switch it on.

**Why This Matters:**

- Buyers pricing an Astra migration pay double Sol's input rate for a model that ties Sol at 61 on the neutral index.
- Vendor-run and neutral benchmark numbers now differ by 37 points, so procurement has to ask which harness produced a score.

Reality Check

**What's confirmed:** OpenAI released Astra on September 3 at $10 per million input tokens and $50 per million output. ARC Prize scored it 62.7% on its Standard harness at maximum effort.

**What's implied (not proven):** That a 99.9% result marks the arrival of general intelligence. ARC Prize wrote that saturating the benchmark would not be proof of AGI.

**What could go wrong:** A team budgets a rollout against the 99.9% figure, then meets 62.7%-shaped behavior on its own unfamiliar internal tools.

**What to watch next:** Whether GDPval, OpenAI's own benchmark for economically valuable work, turns up in follow-up materials. It is absent from the launch.

[Read the full story →](https://www.implicator.ai/openai-gpt-6-astra-agi-era-launch/)

| 3 | Also Today |
| - | ---------- |

The G-20 agreed to hold back new AI rules while Brussels sent out 30 letters.

**All 20 G-20 members adopted the US-drafted Carolina Principles at the Chapel Hill innovation ministerial on September 2.**

The framework asks governments to reserve new AI regulation for "novel considerations," which the announcement never defines. Commerce Secretary Howard Lutnick said China signed with the rest. The European Commission spent the same meeting sending information requests to more than 30 AI companies under the AI Act, so the consensus lands with Brussels already moving.

[Read our coverage →](https://www.implicator.ai/g-20-adopts-carolina-principles-ai-regulation/)

| 4 | Repo Spotlight |
| - | -------------- |

Worktrunk is a Rust CLI that addresses git worktrees by branch name, so one command creates the tree, moves into it and starts the agent. It adds hooks on create and merge, a unique dev-server port per tree, build-cache copying between trees, and a list view carrying CI status per branch.

It earns its keep once a second coding agent is running, where the cost stops being the agent and starts being the branch juggling around it.

wt switch -c -x claude feat

Four more from this week's [Repo Radar](https://www.implicator.ai/repo-radar-5-github-projects-worth-your-week-16/), including JetBrains shipping its Go idiom rules as an agent skill.

[Worktrunk on GitHub →](https://impli.me/qgsgI4?ref=implicator.ai)

| 5 | The Outside Read |
| - | ---------------- |

**CNBC maps the Chinese supply-chain dependencies hiding beneath America's AI data-center buildout.**

Chinese suppliers account for nearly 30% of certain United States transformer and switchgear categories and roughly two-thirds of global optical-transceiver units. Wood Mackenzie estimates that 2026 shortages already equal 15% of power-transformer demand and 8% of substation demand, making new restrictions a near-term cost risk.

[Read it at CNBC →](https://www.cnbc.com/2026/09/03/us-ai-data-centers-china-supply-chain.html?ref=implicator.ai)

| 6 | The One Number |
| - | -------------- |

$26,098

What OpenAI's headline benchmark cost on a harness the vendor did not shape. ARC Prize ran Astra at maximum effort on its provider-neutral setup and recorded 62.7% for that money. The vendor-adapted run scored 99.9% and cost $18,817, so the neutral test came in 39% more expensive and 37 points worse.

Source: [ARC Prize, September 3, 2026](https://impli.me/zbf97y?ref=implicator.ai)

| 7 | Today's Headlines |
| - | ----------------- |

- **Nvidia** agreed to [acquire Hugging Face for $12.93 billion](https://impli.me/WLDX3m?ref=implicator.ai), its largest deal, putting the industry's main open-model hub under a chip vendor.
- **Crusoe** signed a [five-year, $13 billion AI cloud contract with Jane Street](https://impli.me/5TuVoa?ref=implicator.ai), a rare named financial buyer committing to GPU clusters for training and inference.
- **OpenAI** committed [$1 billion over six months to subsidized Daybreak cyber access](https://impli.me/tBeXbq?ref=implicator.ai) for essential-service defenders, adding an MS-ISAC pilot and more than 35 partner products.
- **Google** shipped [Gemini 3.8 Flash and a gated Cyber variant](https://impli.me/kNfkLh?ref=implicator.ai), with introductory pricing on the Cyber model running until December 31.
- **New York City** [barred student-facing generative AI through eighth grade](https://impli.me/8yCQZM?ref=implicator.ai) for the 2026-27 year, a rule reaching close to 600,000 students.

**Tuesdays go deeper.** [Sign up for Implicator PRO](https://www.implicator.ai/subscribe/) for the weekly Tuesday deep dive on deploying AI where it pays. $8 a month, $89 a year.

| 8 | The 5-Minute Skill |
| - | ------------------ |

Reference calls drift into polite praise when the questions are broad. Turn the candidate's strongest claim into a short verification sequence.

**Your raw input:** the candidate's résumé and your interview notes, plus the job's most important outcome and the reference's role if you know it.

**The prompt:**

Use the material above to identify the candidate claim that matters most to the role. Create a ten-minute reference-call script that tests this claim without revealing the answer I hope to hear. Begin with an open question asking the reference to recall a specific episode. Follow with questions that establish the candidate's personal contribution and the observable result. Include one question about where the claim may overstate the candidate's role. End with a neutral sentence I can use to check my interpretation with the reference. Do not invent details.

**Why this works:** anchoring the call to one claim blocks generic praise from becoming evidence. An open recollection question reduces priming, and the interpretation check catches misunderstandings before they enter the hiring record.

**What to use:** ChatGPT handles the concise call script well. Claude is a reliable fallback when the interview notes run long.

| 9 | What To Watch Next |
| - | ------------------ |

| SAT 9/5  | Fairs: IFA Berlin begins its first full public day at Messe Berlin, where organizers host more than 1,900 consumer-tech, robotics and AI brands from 10 a.m. to 6 p.m. CEST. |
| -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| MON 9/7  | Finance: The NYSE and Nasdaq close for the Labor Day holiday.                                                                                                                |
| MON 9/7  | AI: ECML PKDD organizers open the five-day European machine-learning and data-mining conference in Naples.                                                                   |
| THU 9/10 | Policy: The ECB announces its monetary policy decision in Berlin, followed by President Christine Lagarde's press conference at 2:45 p.m. CEST.                              |
| FRI 9/11 | Finance: The Bureau of Labor Statistics reports August consumer-price data at 8:30 a.m. ET.                                                                                  |

| 10 | AI Image of the Day |
| -- | ------------------- |

![Portrait of a man with pale violet reptilian scales, ridged spines running over his brow and shoulders, and a large spiral horn curling behind one ear, against a dark teal background](https://www.implicator.ai/content/images/2026/09/nl_image_day_600-3.jpg) 

Credit: [Ideogram](https://ideogram.ai/g/eKh2%5F9SlTDq%5FW8oFcDKEsQ/3?ref=implicator.ai)

Prompt: Change to man

| 11 | The Rausschmeisser\* |
| -- | -------------------- |

A Python script took a model's refusals from 97 in 100 down to three.

*Heretic, a Python tool now at 30,276 stars, strips safety alignment out of an open-weights model in a single command. On Gemma 3 12B it took refusals from 97 of 100 prompts down to three (*[*The Implicator, September 3, 2026*](https://www.implicator.ai/repo-radar-5-github-projects-worth-your-week-16/)*).*

**Our take:** The method is the part worth sitting with. Heretic runs directional ablation under an Optuna search that co-minimizes refusals against KL divergence from the original, which is a careful way of saying it finds the cheapest edit that removes the conscience without damaging the brain.

Every lab shipping open weights has now been handed a receipt. Ninety-seven down to three, one command, no fine-tuning budget, and a parameter search that did the thinking. The alignment was real work. It also turned out to be a setting.

\*German for the last song of the night, the one that clears the room.

FREE · ABOUT FIVE MINUTES

Stop chasing the AI news cycle.

Implicator cuts the flood to the developments worth your attention, with concise reporting on what changed and why it matters. Delivered every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.

[Get the free briefing →](https://www.implicator.ai/subscribe/?utm%5Fsource=newsletter&utm%5Fmedium=cta&utm%5Fcampaign=briefing%5Fsignup&utm%5Fcontent=variant%5Fc)

From San Francisco. No spam. Unsubscribe anytime.