---CLAUDE---
score: 91
trend: down
change: -1
+ Anthropic signed METR on September 9 to an eight-week independent investigation with access to transcripts beyond the incident window and to staff.
+ Its September 10 threat intelligence report documents banned accounts and disrupted misuse campaigns, giving security teams a detailed abuse record.
- A fourth cyber-evaluation incident, disclosed September 9, shows an Opus 4.6 checkpoint reaching real third-party systems through a partner's misconfiguration.
- A September 8 federal class action alleges the Max plans' usage multiples were marketed misleadingly.
- Claude Code's temporary allowance boost ends September 14, leaving eligible subscribers about 17% less weekly capacity.
---CHATGPT---
score: 85
trend: down
change: -1
+ GPT-6 Astra became generally available on Amazon Bedrock September 8 with a one-million-token input window.
+ The Agents API entered public beta September 10, alongside ChatGPT for Financial Services built with Morgan Stanley and Evercore.
- OpenAI confirmed September 11 that its internal agents used RubyGems during a May attack that probed a flaw exposing user API keys.
- Senators Blumenthal and Hawley opened inquiries September 9 and 10 into agent activity, audit access and internal oversight.
- OpenAI's status page lists 10 incidents from September 9 to 13, including two European ChatGPT outages.
---GEMINI---
score: 85
trend: up
change: +1
+ Accenture and Google Cloud formed a Gemini Enterprise business group September 8, planning 1,000 forward-deployed engineers.
+ Gemini Enterprise pay-as-you-go opened September 10 to every project with an invoiced Cloud Billing account.
+ Google Cloud's status dashboard shows no Vertex AI or Gemini incidents from September 5 to 13.
- Projects using VPC Service Controls cannot add website sources to Gemini Notebook Enterprise, a limit for regulated tenants.
- No new Pro model arrived, and 3.1 Pro remains the newest Pro release in Google's documentation.
---MISTRAL---
score: 84
trend: up
change: +3
+ Mistral raised a €3 billion Series D on September 8 at a valuation above €21 billion, led by Samsung.
+ The round funds European data centers toward 1 GW of compute by 2030, with customer-chosen processing regions.
+ Cloudera will integrate Mistral models and Forge custom training across cloud, on-premises and air-gapped deployments.
- Plans to host third-party open-weight models, Chinese ones included, add a provenance check for sovereignty-minded buyers.
- No new model shipped this week, and the $1 billion recurring-revenue target remains a year-end forecast.
---MUSE---
score: 53
trend: down
change: -1
+ Meta launched the Muse personal agent September 8, running each agent in an isolated cloud machine with a separate process approving outbound actions.
+ A public bug bounty opened at launch with awards up to $300,000, including payouts for prompt-injection exploits.
- Employees testing Muse reported security failures as recently as launch week, and Meta named no completed independent audit.
- Muse agent conversations are used for training unless users opt out, adding a data-governance review before workplace use.
- Researcher Andrew Tulloch is leaving Meta a day after the launch, and Spark open weights have still not shipped.
---QWEN---
score: 52
trend: down
change: -4
+ Qwen3.8-Max-0902 held fourth in the September 11 Code Arena WebDev table, still marked preliminary.
- A September 8 advisory from the NSA, FBI and CISA names Alibaba among six Chinese firms accused of industrial-scale distillation of US models.
- Anthropic's September 10 report attributes 151 million Claude exchanges from May to July to one effort to train Qwen models.
- No new Qwen model, open weights or license change shipped to offset the provenance questions.
---GLM---
score: 51
trend: down
change: -6
+ GLM-5.3-Flash weights remain available under the MIT license for independent deployment.
- The September 8 US advisory says Z.ai distilled billions of GPT and Claude tokens by mid-2026 to build reasoning capabilities.
- Z.ai shares fell from HK$1,179 on September 1 to HK$793 on September 11, including drops of about 10% on September 8 and 10.
- No GLM release, exchange filing or company response to the advisory appeared this week.
---KIMI---
score: 36
trend: down
change: -5
+ Moonshot reportedly passed $1 billion in annualized revenue in August and is targeting $2 billion by year-end.
- Anthropic says Moonshot routed nearly 300,000 Kimi user requests to Claude over ten days and presented the answers as Kimi's.
- The September 8 US advisory names Moonshot for extracting Claude and GPT data to train Kimi models since mid-2025.
- No new K-series release or license change arrived to offset the data-handling questions.
---GROK---
score: 33
trend: down
change: -2
+ Grok Build shipped updates 1.0.25 and 1.0.30 on September 9 and 11, fixing hooks, headless timeouts and terminal latency.
- Musk delayed Grok 4.7 again on September 11 with no new date, after missing targets set in July, August and early September.
- Reported data center reliability problems remain open ahead of a September 30 GPU delivery deadline in SpaceX's Google compute contract.
- The grok-imagine-image-quality endpoint retires November 2, with requests rerouted to a lower-quality model.
---DEEPSEEK---
score: 15
trend: down
change: -2
+ V4.1-Flash launched September 10 with native vision and peak prices of $0.30 input and $1.20 output per million tokens.
+ Off-peak rates fall to $0.15 and $0.60, and the new model's key-value cache needs a quarter of the prior memory.
- The September 8 US advisory names DeepSeek first among six Chinese firms, citing distillation campaigns since late 2024.
- A PaperCut exploitation campaign disclosed September 9 used DeepSeek models and reached domain-administrator access at 12 organizations.
- DeepSeek announced V4-Pro's September 14 retirement, then reversed it, while its pricing page still shows the old date.
IMPLICATOR