From San Francisco
1 |
The Editorial |
Marcus here.
OpenAI released GPT-5.6-Cyber to a vetted list of defenders. It completes 95% of exploit-development requests; the consumer model completes 1.5%. The release follows OpenAI saying it could not rule out a Critical cyber capability in Astra. Zuckerberg asked Washington to loosen the rules on training data and distillation. Both permission slips were written in-house.
Below: the cyber model that ships with a customer list, Meta's open-weight argument arriving with an actual product, and the valuation OpenAI just declined to move in a $7 billion buyback.
Stay curious,
Marcus Schuler
2 |
The Big Story |
OpenAI released GPT-5.6-Cyber to approved security researchers through a new Daybreak Red tier, days after it said preliminary evaluations left it unable to rule out a Critical cyber capability for Astra.
On the company's internal completion test, the model finished 95.0% of requests covering exploit-chain development, authentication bypass and privilege escalation. GPT-5.6 Sol with its safeguards on finished 1.5%.
Accenture, IBM, CrowdStrike, Cisco and Palo Alto Networks may put the model into security products. Entry requires identity verification, account monitoring and legal attestations, and hardware security keys become mandatory for individual Daybreak accounts on September 1. The completion rates are OpenAI's own measurements, and no system card has been published.
Why This Matters:
- Security teams at the five named partners can buy exploit-development capability as a product feature instead of building it in house.
- The vetting list is OpenAI's to write, so access to offensive capability now runs through a private approval process.
3 |
Also Today |
Meta released Muse Glimmer, a 30-billion-parameter open-weight model built to run agents on a single consumer graphics card.
Zuckerberg paired it with a 6,500-word essay calling for fewer U.S. restrictions on training data and distillation, the method Meta used to build Glimmer from its larger Muse Spark. The company kept Spark itself behind a paid cloud interface. The open release lands while Meta plans up to $145 billion in capital spending this year.
4 |
The Outside Read |
SemiAnalysis tests AMD's improving software stack against Nvidia's CUDA moat and names the operational failures that could still block AMD's advance.
Anthropic's announced deployment totaled 2 gigawatts of AMD chips as of July 24, 2026, while AMD's February benchmarks put AITER optimizations at 1.08 to 1.2 times baseline framework throughput. SemiAnalysis also finds that unstable internal GPU clusters are impeding AMD's own software development and automated testing.
5 |
The One Number |
6 |
Today's Headlines |
- Apple published a Chinese-language Mac guide for connecting Siri to Alibaba's Qwen, then pulled the page a day later while its support staff said the feature had not launched.
- A federal appeals court refused to pause the trial brought by 29 state attorneys general over Meta's design choices and youth mental health, with jury selection due Wednesday in Oakland.
- Applied Compute is in talks to raise hundreds of millions at roughly a $3 billion valuation led by Elad Gil, double the mark it took in April.
- OpenAI's head of ethics, Chloé Bakalar, left after less than a year, following exits by the head of safety systems and a former head of mission alignment.
- Anthropic put an unreleased Claude model on the Riemann hypothesis; it did not solve it but raised a lower bound on the zeta function's zeros from 41.6% to 67.2%.
- Claude passed ChatGPT in our enterprise scorecard after the Black Hat disclosure round.
Wed 8/12 |
Economy: the Bureau of Labor Statistics releases July CPI at 8:30 a.m. Eastern. |
Wed 8/12 |
Courts: jury selection begins in Oakland in the 29-state case against Meta, with opening arguments set for Aug. 18. |
Wed 8/12 |
Tech: Made by Google in New York at 6 p.m. Eastern, with the Pixel 11 line and Pixel Watch 5; pre-orders open the same day. |
Thu 8/13 |
Economy: July PPI follows at 8:30 a.m. Eastern. |
7 |
The 5-Minute Skill |
A polished candidate submission can hide weak judgment, and a plain one can contain the better decision. Make the model score observable evidence instead of writing style.
Your raw input:
The prompt:
Why this works: Weighted criteria anchor the evaluation to the job. Evidence quotes make the scoring inspectable and cut the chance that confident prose stands in for judgment.
What to use: GPT-5.6 for disciplined rubric work. Claude Opus 4.7 is a good fallback when the submission runs long.
8 |
AI Toolbox |
StepShot turns a Mac workflow into a numbered, illustrated SOP while you perform it. It captures pre-click frames and interface labels automatically instead of leaving you to shoot and label screenshots by hand. The free tier exports the first 10 steps; Pro costs $129 once, per StepShot's pricing page checked August 9, 2026.
How to use it:
- Download StepShot for a Mac running macOS 14 or later and open the app.
- Complete the first-run wizard, grant Screen Recording, Accessibility and Input Monitoring permissions, then relaunch so macOS applies them.
- Choose Action-driven capture, press Option-Command-5, and perform the workflow you want to document.
- Press Option-Command-5 again, then reorder, merge or skip captured steps and add notes or section headings.
- Redact any sensitive pixels, preview the document and export it as Markdown or HTML; Pro is required for PDF or more than 10 exported steps.
Where it falls short: StepShot is Mac-only, needs macOS 14 or later, and caps free exports at the first 10 steps.
9 |
The Rausschmeisser* |
A Melbourne man asked his OpenClaw assistant, running on Claude, to book him into a popular gym class. It booked the class, left him fourth on the waitlist, and when he asked whether it could do better it found a flaw in the gym's booking API, worked around the restriction and cancelled a stranger's reservation to move him up one place. Australian authorities are treating it as the country's first known autonomous AI cyberattack (ABC News, August 10, 2026).
Our take: Nobody told it to do that. The instruction was "get me into the 6am class," the kind of request a person makes without thinking, and the agent found an API flaw the gym did not know it had. Somewhere in Melbourne a stranger opened an app and discovered they were no longer going to Pilates.
The gap between this and the enterprise pitch is smaller than anyone selling agents would like. Every deck promises an assistant that pursues your goal without being told each step. That is what happened here. The victim was a booking queue, so the cost is one cancelled class. Point the same helpfulness at a procurement system and the story stops being about Pilates.
IMPLICATOR