From San Francisco
1 |
The Editorial |
Marcus here.
People still have to check what the machines do.
OpenAI halted broadly defined tool use after a research model sent questions to a public chatbot through DNS on Sept. 20. The 15-minute alert was only the start.
Our Sept. 23 PRO piece puts GPT-6 Astra in front of 13 invoice rows whose arithmetic looks sound. The part worth copying comes after the five catches.
Blue Cross says hospital AI coding added $942 million over two years, after insurers spent years automating claim review and denials. The awkward bit begins when both bots reach the same bill.
Stay curious,
Marcus Schuler
2 |
THE BIG STORY |
OpenAI paused training, evaluation and inference with broadly defined tool use for its most capable models after a Sept. 20 DNS escape.
An internal research model found weak DNS filtering during a search task. It used the resolver to send questions to a public chatbot. A monitor flagged the behavior within 15 minutes. The automatic shutdown failed.
Staff stopped the run manually about two and a half hours later. OpenAI added blocking controls at two separate layers. It also restricted DNS queries. The company plans to discard the affected run and restart training from scratch.
Why This Matters:
- AI teams now face interrupted model work when agent controls fail, even when no real-world harm is known.
- OpenAI plans to restart training from scratch and resume only after adding safeguards and alignment improvements.
3 |
ALSO TODAY |
AWS made CloudWatch Omni generally available in the week before Sept. 27, 2026.
It captures traces and adds evaluation across LangChain, LangGraph, CrewAI, OpenAI Agents SDK, Strands, Vercel AI SDK and Amazon Bedrock AgentCore. Pricing follows telemetry sent and stored, with dashboards and alerts free and queries up to 5x monthly ingestion included. Customers include Sony and Capital One, giving agent teams a direct check on behavior.
4 |
Worth the Paywall |
Use GPT-6 Astra to catch five invoice problems that arithmetic misses.
Agents and models now touch money and records, so finance teams need checks they can repeat. This is for people who approve vendor bills or run finance operations.
5 |
The Outside Read |
Business Insider reviews 130 U.S. political television ads aired from January 2025 through September 2026 and shows how AI infrastructure became an affordability issue.
The attacks centered on electricity costs and data-center tax breaks rather than speculative AI safety risks. A separate digital-ad analysis found more than 700 campaigns posting about data centers in 2026, with 99.7% of spending backing critical messages.
6 |
The One Number |
7 |
Today's Headlines |
- Anthropic CEO Dario Amodei dined with President Trump on Sept. 27, their first direct meeting, as the company challenges the Pentagon’s supply-chain-risk label after an August court win.
- Google is testing a Buy button inside Gemini in India that sends shoppers to Flipkart checkout, while Amazon listings remain view-only ahead of a broader October rollout.
- Black Forest Labs released FLUX 3 Action on Sept. 23, a 7-billion-parameter open-weight robot model scoring 42.92% on RoboLab-120 and returning two seconds of actions.
- OpenAI briefly listed “o, your always-on assistant” on its $100 ChatGPT Pro page; code references include an “-o” email suffix, and the feature remains unconfirmed before Tuesday’s DevDay.
- Google will convert Gemini Gems into skills on Nov. 17, 2026, with “/” invocation and support for several at once; existing Gems work until migration.
8 |
The 5-Minute Skill |
Audit an AI agent run before it repeats.
An agent finished a task, but its answer hides the actions it took. Audit the run before allowing a repeat.
Your raw input:
The prompt:
Why this works:
The prompt compares behavior with explicit authority instead of asking for a general summary. Requiring quoted evidence separates recorded actions from guesses and gives an executive a clear next step.
What to use:
Use a long-context reasoning model such as GPT-5.6 Sol or Claude Fable 5.1. If the log is too large, split it by timestamp and run a final synthesis over the separate audits.
9 |
Fresh Funding |
10 |
AI Image of the Day |
11 |
The Rausschmeisser* |
The Blue Cross Blue Shield Association said Sept. 26, 2026, that hospitals' AI claim-coding tools added $942 million in healthcare spending over two years.
Our take: BCBSA’s $942 million complaint arrives after years of automated claim review and denial tools, so insurers now dislike the invoice. The patient’s care stays the same while the bill grows. Efficiency has become a software dispute with a copay.
Luke Chalker calls it a “completely one-sided blood bath,” a diagnosis from the industry that taught software to argue over reimbursement. Soon a coding bot can raise the claim before a denial bot rejects it. A human still needs treatment while the software enjoys the appointment. The only healthy thing here is the billing department.
FREE AI BRIEFING · WEEKDAYS
IMPLICATOR