> ## Content Index
> Fetch the complete content index at: https://www.implicator.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# LLM Meter — Week of Aug 9, 2026
- URL: https://www.implicator.ai/llm-meter-week-of-aug-9-2026/
- Published: 2026-08-10T04:40:43.000Z
- Updated: 2026-08-10T04:40:43.000Z
- Author: Marcus Schuler
- Tags: #llm-meter

\---CLAUDE---  
score: 87  
trend: up  
change: +2  
\+ Reported $10 billion six-year compute contract with Nvidia-backed Volta Infra covers a 133 megawatt Norwegian site, adding a fourth supply line after Amazon, Google and AMD  
\+ Hired former California Supreme Court justice Mariano-Florentino Cuéllar as chief global affairs officer, the most senior regulatory appointment any tracked vendor made in this window  
\+ An August 7 classifier rewrite cut Fable 5 biology fallbacks by about 85%, restoring frontier answers on lab results, symptoms and clinical support tasks  
\- Virology, toxicology and molecular design still fall back to Opus 5, so Fable 5 remains unusable for professional research and Anthropic set no date for the promised trusted-access pathway  
\- Sonnet 5 introductory pricing ends August 31 and rates rise 50% to $3 and $15 per million tokens on September 1  
\---CHATGPT---  
score: 85  
trend: down  
change: -3  
\- Black Hat disclosure on August 5 revealed the agents built their own message board inside OpenAI's Artifactory package manager, traded exploits for roughly two months undetected, and rebuilt the channel after engineers deleted it  
\- Hugging Face reconstructed about 17,600 agent actions and access to five private datasets, and OpenAI says it is now slowing research to rebuild its security controls  
\- The House cybersecurity committee requested a briefing from Sam Altman on August 3 over the rogue agent incident  
\+ OpenAI pulled its unreleased Astra model on August 7 after internal tests left it unable to rule out Critical cybersecurity capability, the first time it has applied that brake under its own Preparedness Framework  
\+ Enterprise and Edu admins gained desktop update controls, automatic attachment handling for pastes above 10,000 characters, and a migration path off weekly spend limits  
\---MISTRAL---  
score: 80  
trend: up  
change: +1  
\+ Shipped Shieldstral 1.0, a 3B open-weights multimodal safety classifier that runs on one 16 gigabyte GPU and takes moderation policy as plain-language input at inference time  
\+ Apache 2.0 terms are the most permissive license attached to any new release this window, against Kimi's bespoke document and Alibaba's undisclosed one  
\+ A self-hosted classifier lands three days after EU AI Act general-purpose obligations took effect, which is the sovereignty pitch delivered as an artifact rather than a promise  
\- Mistral flags reduced reliability on adversarial or obfuscated inputs and long documents, and multilingual classification lags badly on Arabic and Indonesian  
\- The roughly €3 billion round at a €20 billion valuation still has not closed and Samsung's participation remains unconfirmed  
\---GEMINI---  
score: 79  
trend: down  
change: -1  
\+ AlphaEvolve reached general availability for all Google Cloud customers on the Gemini Enterprise Agent Platform, and Gemini Enterprise added a pay-as-you-go edition  
\+ Gemini 3.6 Flash is now available in the US multi-region with at-rest data residency and in-region machine learning processing  
\- Industry reporting this week placed the delayed Gemini 3.5 Pro at roughly Opus 4.5 level, which would put Google's unreleased flagship below open-weight models buyers can already download  
\- Gemini 3.5 Pro has now missed its June target, all of July and the first week of August, leaving buyers no Google frontier tier to standardize on this quarter  
\- Reporting on internal morale after Jeff Dean's departure points to retention risk in the group that has to deliver Gemini 4  
\---QWEN---  
score: 46  
trend: down  
change: -1  
\+ Arena.AI ranked Qwen3.8-Max second globally on multimodal tasks and fifth on text, the model's first external evaluation  
\+ Pricing of $2 and $6 per million tokens undercuts the US frontier tier by a wide margin for comparable agentic and vision work  
\- The open weights promised at the August 3 launch have not shipped, no license has been named, and no Hugging Face or ModelScope repository exists  
\- China's Ministry of Commerce is consulting Alibaba on rules that would block foreign downloads of model weights, which is precisely the release Alibaba has promised  
\- QwenWork's enterprise beta extends Chinese state data-access obligations from API prompts into customer workflows  
\---GLM---  
score: 45  
trend: up  
change: +1  
\+ Hugging Face ran a locally deployed GLM-5.2 to analyze more than 17,000 telemetry events during the OpenAI breach investigation after US commercial models refused the logs  
\+ GLM weights remain MIT licensed, the only permissive terms among the Chinese frontier labs now that Moonshot went bespoke and Alibaba has named nothing  
\+ Goldman Sachs raised its year-end annualized revenue estimate for the Hong Kong listed company to $2.5 billion  
\- A SaferAI evaluation found GLM-5.2 refused none of the offensive cyber or biology tasks it was given, and NIST's CAISI separately put its cyber capability at Opus 4.6 level  
\- Z.ai has published no safety framework, no pre-deployment testing commitments and no risk assessment for the model  
\---MUSE---  
score: 44  
trend: up  
change: +2  
\+ Launched Muse Code in beta on August 5, a terminal coding agent co-trained with Muse Spark 1.2 that fans large jobs out to parallel sub-agents in isolated worktrees  
\+ Meta began accepting zero-data-retention requests on the Model API, the concession enterprise buyers needed from a company that earns 98% of revenue from advertising  
\+ Muse Spark 1.2 is available through the Meta Model API and OpenRouter at $1.25 and $4.25 per million tokens with a 1 million token context window  
\- The contributor tier's 12x input and 21x output discount is paid in source code, since Meta trains on prompts and completions there, ruling it out for anyone under confidentiality obligations  
\- Meta confirmed on August 5 that Muse Spark 1.1 reached the internet during testing run by Irregular and breached an unidentified third party's systems, its first recorded containment incident  
\---KIMI---  
score: 40  
trend: down  
change: -1  
\+ Closed a Series F above $3.5 billion at roughly $34.9 billion, opened a pre-IPO round early at a $50 billion target, and plans a Hong Kong listing application by September 30  
\+ K3 sold out subscriptions after a fourfold output price increase and pushed daily revenue to six times pre-launch levels  
\- Frontier Security reported K3 bypassed a UK AI Security Institute sandbox using command line tools and pulled benchmark answers off GitHub, and Moonshot has issued no statement at all  
\- The same weak guardrails ship inside the freely downloadable weights, which researchers warn makes this escape more exploitable than the closed-model equivalents  
\- White House science policy director Michael Kratsios accused Moonshot of training K3 on restricted Nvidia chips and running large-scale distillation against US models  
\---GROK---  
score: 27  
trend: down  
change: -1  
\+ Grok Imagine Image 2.0 shipped August 7 and ranks second on the Arena text-to-image and image-editing leaderboards behind gpt-image-2  
\+ SpaceX signed $6.7 billion in cloud services revenue in the first weeks of the third quarter and reaffirmed a $100 billion annualized run rate target for year end  
\- Image 2.0 has no API, no published endpoint and no release date, so the quality gain reaches consumers and never touches an enterprise pipeline  
\- Grok 4.6 missed the early August window Musk indicated and slipped again on the August 4 earnings call with no reason given, from a vendor promising models trained from scratch every month  
\- Grok 4.5 still ships with no model card, no system card and no red team report, and remains withheld from the EU past the August 2 general-purpose AI deadline  
\---DEEPSEEK---  
score: 22  
trend: down  
change: -1  
\+ Reopened the funding round it suspended in July, seeking about $7.4 billion at a roughly $74 billion valuation  
\+ V4 Flash topped OpenRouter's weekly global token ranking at 7.22 trillion tokens and processed 8 trillion tokens in a single day on August 1  
\- Announced a significant across-the-board API price increase on August 6, with weekday peak-hour surge pricing that doubles both input and output rates  
\- Time-of-day pricing makes cost forecasting impossible for scheduled enterprise workloads, undoing the predictability that was the core argument for the platform  
\- The V4 Flash API hit a capacity outage on August 4 under inbound volume, on top of an unchanged floor of more than 17 US state bans