---CLAUDE---
score: 92
trend: up
change: +1
+ Fable 5.1 launched September 1 across the major clouds, with cache reads cut 75% to $0.25 per million tokens.
+ Eligible customers can use Fable 5 and 5.1 with zero data retention while Enterprise Frontier Safeguards is being prepared.
- EFS rolls out later this fall, with customer storage costs and human review remaining the buyer's responsibility.
- Anthropic's August 31 security update still described its incident investigation and independent METR review as unfinished.
---CHATGPT---
score: 86
trend: up
change: +2
+ GPT-6 Astra began rolling out September 3 for coding, research, computer use and complex business tasks.
+ Astra led the September 5 Code Arena WebDev table, providing external evidence beyond OpenAI's launch benchmarks.
- Access remains phased, and the launch release notes explicitly say Astra is not yet generally available.
- OpenAI's system card records weaker reasoning monitorability despite improved compliance with safety restrictions.
---GEMINI---
score: 84
trend: up
change: +5
+ Gemini 3.8 Flash reached general availability September 2 with a one-million-token context window.
+ Introductory pricing is $0.75 per million input tokens and $3.75 per million output tokens through December 31.
- Standard token prices double January 1, and Google says complex tasks can consume more tokens than with 3.7 Flash.
- Flash Cyber is restricted to trusted defenders through Fairwind, so its capabilities do not belong in a general-access comparison.
---MISTRAL---
score: 81
trend: down
change: -1
+ Vibe Enterprise remains excluded from model training by default, with organization controls for administrators.
+ Standard Vibe and API customers can opt out of training without buying Enterprise.
- Vibe and API training controls are separate, adding a configuration check for businesses using both products.
- Zero data retention requires approval and covers eligible stateless API calls, excluding Vibe chats and stateful agents.
---GLM---
score: 57
trend: up
change: +3
+ First-half 2026 revenue reached RMB954 million, up 399.7%, with platform and API services contributing 86.5%.
+ GLM-5.3-Flash weights remain available under MIT, preserving a permissive route to independent deployment.
- First-half adjusted net loss widened 12.1% to RMB1.964 billion, despite the decline in reported net loss.
- Gross margin fell to 26.4% from 50.0%, limiting the financial benefit of rapid API growth.
---QWEN---
score: 56
trend: up
change: +4
+ Qwen3.8-Max-0902 shipped September 2 through Alibaba Cloud with a one-million-token context window.
+ The model placed fourth in the September 5 Code Arena WebDev table, with its result still marked preliminary.
+ The current Max weight license exempts qualifying internal use from its separate commercial-license requirement.
- Specified model-service and AI work-assistant businesses above $50 million in annual group revenue need a separate license.
---MUSE---
score: 54
trend: up
change: +6
+ Muse Spark 1.3 shipped September 2 in Muse Code and Meta Model API for longer coding and agent tasks.
+ Published token rates remain $1.25 per million input tokens and $4.25 per million output tokens.
- The live benchmark page differs from the September 2 snapshot, so the earlier 42% task-cost advantage is not a current comparison.
- Meta still lists Spark open weights as future work, leaving the new release dependent on hosted access.
---KIMI---
score: 41
trend: up
change: +2
+ Moonshot reportedly filed confidentially for a Hong Kong IPO targeting about $3 billion, improving its potential financing options.
+ K3's license exempts internal use and access through official products or certified inference partners from specified commercial conditions.
- The IPO proceeds and proposed cloud revenue-sharing deals remain uncompleted transactions.
- K3's published repository is 1.56TB, keeping independent deployment a substantial infrastructure commitment.
---GROK---
score: 35
trend: up
change: +4
+ Grok Bot for Enterprise launched September 3 with access, network and audit controls.
+ Grok and Cursor Enterprise customers received a two-week trial covering colleagues without existing seats.
+ SpaceXAI published a September 1 account of external biological safety testing, adding evidence for specialist procurement reviews.
- Trial access and narrow safety evaluations do not establish long-term operating cost or safety across business workflows.
---DEEPSEEK---
score: 17
trend: up
change: +1
+ DeepSeek's status page reported no incidents from August 29 through September 6, a limited positive reliability signal.
+ Current documentation supports Responses and Anthropic API formats, reducing integration work for existing agent tools.
- V4-Flash still costs $0.44 per million uncached input tokens and $1.32 per million output tokens during weekday peak hours.
- Off-peak rates are half those prices, leaving cost dependent on scheduling, while the vision endpoint remains experimental.
IMPLICATOR