From San Francisco
1 |
The Editorial |
Good morning.
Today's three picks were all graded by the company that shipped them.
Meta's Muse Spark 1.3 tied GPT-5.6 Sol at 61 on the Artificial Analysis index and runs $0.55 a task against $0.95. Its cost climbed from $0.40 in version 1.2, and two evaluations went backward.
OpenAI says Astra is the first model in any risk domain to reach Critical, on a 100% ExploitBench score. No outside party has seen the measurement.
OpenClaw 2.0 shipped shared sessions that let a second person take over a running agent. The project's own documentation says those controls are not a security boundary.
Stay curious,
Marcus Schuler
2 |
The Big Story |
Meta released Muse Spark 1.3 on September 2 with an Artificial Analysis index score of 61, level with GPT-5.6 Sol, at $0.55 per task against Sol's $0.95.
Token prices are unchanged from version 1.2: $1.25 per million input, $4.25 output, $0.15 for cached input. The gap to Sol comes from those rates, not from shorter runs.
Against its own predecessor the model got more expensive. Version 1.2 cost $0.40 a task and scored 57, and 1.3 uses 57% more input tokens. AA-LCR fell from 83% to 79%, and Omniscience accuracy dropped three points at xhigh. Max mode scores 62, limited to Meta partners with no published price.
Why This Matters:
- The price advantage sits in the token rate, so it survives only while Sol and Grok hold their current list prices.
- Buyers comparing 1.3 against 1.2 rather than against Sol will find a model that costs 38% more per task.
3 |
Also Today |
OpenAI said on September 1 that Astra is the first model to reach Critical under its Preparedness Framework, on a 100% ExploitBench score.
Critical means finding and developing working zero-day exploits against hardened systems unaided. OpenAI reports a 91.5% cyber-jailbreak refusal rate against Sol's 59%, and zero attempts on prohibited targets in a honeypot where Sol tried 56% of the time. Every figure is OpenAI's own measurement of an unreleased model.
4 |
The Outside Read |
Engineering at Meta explains a production agent that converts expert corrections into auditable updates without retraining its model.
The architecture stores institutional knowledge in explicit source files and keeps reasoning procedures in separate recipes. Expert feedback becomes a proposed diff that must pass blind replay against the original failure before a human specialist approves it; Meta says this reduced individual assessments from days to minutes.
5 |
The One Number |
6 |
Today's Headlines |
- Anthropic signed a six-year, $35 billion cloud agreement with Lambda for 350 megawatts of Nvidia capacity in Texas, due to begin energizing in 2027.
- A federal judge spared Google an ad-tech breakup and ordered it to interoperate with rival ad servers and exchanges instead.
- The Justice Department backed OpenAI's training-use defense in The New York Times copyright case, filing its position while the suit is live.
- The European Commission sent AI Act information requests to more than 30 companies on safety and copyright compliance.
- Palo Alto Networks paid a reported $500 million for Console, whose help-desk agents go into the Cortex security platform.
- Google's Gemini 3.8 Flash scored 59 on the Artificial Analysis index, two behind GPT-5.6 Sol and seven behind Claude Fable 5.1, at $0.58 per task.
Thu 9/3 |
Economy: the Bureau of Labor Statistics releases revised second-quarter productivity and costs at 8:30 a.m. Eastern. |
Thu 9/3 |
Chips: SEMICON Taiwan continues in Taipei with forums on advanced testing, chip design and materials. |
Fri 9/4 |
Jobs: the Bureau of Labor Statistics releases the August employment report at 8:30 a.m. Eastern. |
Fri 9/4 |
Consumer tech: IFA Berlin opens its five-day show at Messe Berlin. |
7 |
The 5-Minute Skill |
A vendor proposal often hides its biggest risk inside an untested claim. Use the model to turn that claim into a proof request before approval.
Your raw input:
The prompt:
Why this works: The prompt narrows the analysis to one load-bearing claim. Requiring quoted language keeps the analysis tied to the source, while the decision rule turns uncertainty into an action.
What to use: Claude handles long proposals well. ChatGPT is a strong fallback when the source packet is shorter.
8 |
Repo Spotlight |
Tencent's AI-Infra-Guard scans MCP servers and agent skill packages against a vulnerability library covering 146 AI components and more than 2,000 CVE rules. It also runs jailbreak tests using Many-Shot, PAIR, GOAT and ActorAttack methods.
It is built for security teams who have to approve an MCP server or a skill package before it reaches an agent in production.
Version 4.6.0, Apache-2.0, about 6,100 stars, maintained by Tencent's Zhuque Lab.
9 |
AI Image of the Day |
10 |
The Rausschmeisser* |
OpenClaw 2.0 landed on August 30 after a seven-week pause, carrying 16,000 pull requests from 933 contributors and a headline feature that lets a second person join a running agent or take it over outright. The documentation says the permission controls "are not tenant isolation and not a security boundary" (The Implicator, August 31, 2026).
Our take: Read that again. The feature is that a stranger can pick up your agent while it holds your credentials and runs your commands. The permission modes run from read-only to contribute directly, which sounds like a security model right up to the sentence where the project says it is not one. Sandboxing ships off by default.
The honest version is in the docs: one Gateway is one trust domain, and anyone who needs real separation should run a second Gateway. The multiplayer feature works fine as long as everyone in the session is someone you would hand your laptop and your password to. At that headcount it stops being multiplayer and starts being the person at the next desk.
FREE WEEKDAY MORNING BRIEFING
IMPLICATOR