Riley Brown, of Agent Native, had asked Claude Opus 5 to extract the branding from a website and turn it into an investor presentation, then watched the model keep working. In a video published July 25, he said, "I entered this prompt 41 minutes ago and it's still going. This is insane."
When the run finally ended, he called the result "genuinely one of the best PowerPoints I've ever seen."
Anthropic released Opus 5 on July 24, 2026, and measured tests put it near the top of software-engineering benchmarks. During its first five days, early users and published reviews repeatedly complained that it widened tasks and responded at length. Anthropic's system card also documented confidence that did not always hold up.
What Changed
- Anthropic released Claude Opus 5 on July 24, 2026. Measured tests put it near the top of software-engineering benchmarks, and its first five days drew repeated complaints that it widened tasks and answered at length.
- Anthropic's own prompting documentation says the model "can also expand the scope of a task, adding steps that weren't requested," and tells developers to strip out verification instructions that cause over-verification.
- The system card records that Opus 5 "often would state an answer as certain, when it was unsure," while its "Lack of overconfidence" benchmark is topping out the test Anthropic uses to check that behavior.
- Boris Cherny's team removed more than 80% of Claude Code's system prompt and found the model performed better. Opus 5 rewrote more than 100,000 lines of the Bun runtime from Zig to Rust in 11 days.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
The early complaints
Dan Shipper, of Every, had early access and spent a week testing the model across coding, writing and knowledge work. Shipper, quoted in Brown's video, called Opus 5 "very hard to love" and said it "argues with instructions, stopped before work was finished." His first reaction, as relayed by Brown, was: "What have they done to my boy?"
Brown reached a similar verdict when he compared the model with Fable, Anthropic's higher-priced frontier model. He said, "I can't really tell the difference between this model and Fable. It just feels kind of worse." Shipper went further, calling it "a poor man's Fable" with the personality quirks but less of what he regarded as the other model's ability.
CodeRabbit and Snorkel AI published independent measurements. CodeRabbit's review found lower coverage than its baseline and four times as many nitpicks, even as the model produced a cleaner stream of actionable comments. Snorkel's post analyzed 195 trajectories and ranked Opus 5 second on Senior SWE-bench, while placing it first for bug investigation.
Anthropic's warnings
Anthropic's prompting documentation describes much of the behavior its users encountered. The company wrote, "Claude Opus 5 can also expand the scope of a task, adding steps that weren't requested or applying its own judgment about what the task should be." Its guidance tells developers to remove explicit checks because "instructions like these cause over-verification on Claude Opus 5."
Turning down the effort setting does not necessarily solve the visible output problem. Anthropic says "lowering effort can reduce thinking volume without reliably shortening the visible response." The company also warns that Opus 5 delegates more readily than prior models, a habit that adds cost and time when the work is small.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
The system card recorded a separate limit. During Anthropic's testing, the model "often would state an answer as certain, when it was unsure." The model's score on the "Lack of overconfidence" benchmark was topping out the test Anthropic uses to check whether it knows when it is unsure. That left the test little room to show further improvement. Mythos 5, another Anthropic model, served as an external reviewer and objected that the draft understated how often Opus 5 made confident claims and then withdrew them. In the system card, Opus 5 scored 2.3 on misaligned behavior, below Sonnet 5's 3.35, the lowest among Anthropic's recent models.
As of July 29, the evidence was five days old and self-selected, drawn from early testers, one video, published reviews, Anthropic's documentation and its system card. No session-level measurement of how often Opus 5 retracts a claim has been published.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Removing the old scaffolding
Boris Cherny, Anthropic's creator of Claude Code, reported a different result. After his team removed more than 80% of Claude Code's existing system prompt, Opus 5 performed better, according to Cherny.
The model rewrote more than 100,000 lines of the Bun JavaScript runtime from Zig to Rust in 11 days, according to Cherny's account. Thousands of agents worked in parallel and checked the result against Bun's test suite. Anthropic also runs between 20 and 30 autonomous maintenance routines each day across its own codebases.
After the week of testing described in the video, Brown relayed Shipper's account that his team "deleted our existing skills and started from scratch." Anthropic's documentation tells developers to remove verification instructions that may be carried over from prompts written for earlier models.
Zvi Mowshowitz, an AI commentator who writes about model releases, cautioned against treating the early complaints as a property of the model alone. He said the harness and the way a person talks to Opus 5 can determine whether the trouble appears. The video's host, however, was left calculating the cost of that adjustment: "If every time a new model comes out, you have to redo all of your skills, it is a complete waste of time."
Frequently Asked Questions
What are developers complaining about in Claude Opus 5's first week?
That it widens tasks beyond what was asked and answers at length. Dan Shipper of Every, quoted in Riley Brown's video, called the model "very hard to love" and said it "argues with instructions, stopped before work was finished." Brown ran a single deck-building prompt for 41 minutes before it completed.
Does Anthropic acknowledge the scope problem?
Yes, in its own prompting documentation. The company writes that Opus 5 "can also expand the scope of a task, adding steps that weren't requested or applying its own judgment about what the task should be," and advises developers to constrain scope explicitly for narrow tasks.
Can you fix it by lowering the effort setting?
Not reliably. Anthropic states that "lowering effort can reduce thinking volume without reliably shortening the visible response." The company recommends prompting for length directly instead, and warns that Opus 5 also delegates to subagents more readily than earlier models.
What did independent testers measure?
CodeRabbit's review found lower coverage than its baseline and four times as many nitpicks, alongside a cleaner stream of actionable comments. Snorkel AI analyzed 195 trajectories, ranking Opus 5 second on Senior SWE-bench and first for bug investigation.
What is the recommended fix?
Deleting existing scaffolding. Anthropic's documentation tells developers to remove carried-over verification instructions. Boris Cherny's team cut more than 80% of Claude Code's system prompt. Shipper's team, as relayed by Brown, "deleted our existing skills and started from scratch."
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR