Sunday, July 19, 2026 · San Francisco
Strategic AI Intelligence from San Francisco
Now
Alibaba Claims Qwen3.8 Is Second Only to Fable 5 Germany's Soofi S AI Model Tops All Open-Source Rivals on German Benchmarks Moonshot's Kimi K3 Won't Fit on a Single Nvidia DGX B200 Anthropic Resets Claude Limits for Third Week Without Explaining Why Moonshot Launches Kimi K3 With 2.8 Trillion Parameters and 1M Context Fudan's Open-Source Health Agent Posts 45.7% Accuracy on Its Own Synthetic Benchmark Alibaba Claims Qwen3.8 Is Second Only to Fable 5 Germany's Soofi S AI Model Tops All Open-Source Rivals on German Benchmarks Moonshot's Kimi K3 Won't Fit on a Single Nvidia DGX B200 Anthropic Resets Claude Limits for Third Week Without Explaining Why Moonshot Launches Kimi K3 With 2.8 Trillion Parameters and 1M Context Fudan's Open-Source Health Agent Posts 45.7% Accuracy on Its Own Synthetic Benchmark
AI News

Alibaba Claims Qwen3.8 Is Second Only to Fable 5

Alibaba Claims Qwen3.8 Is Second Only to Fable 5

Alibaba's Qwen team said Sunday that its 2.4-trillion-parameter Qwen3.8 trails only Anthropic's Claude Fable 5 among frontier models, and published no benchmark scores to support it. Two months ago the company shipped Qwen3.7-Max with a full benchmark table and an independent Artificial Analysis score of 56.6.

Read full story →
01 Latest Intelligence
Repo Radar: 5 GitHub Projects Worth Your Week
Tools & Workflows

Repo Radar: 5 GitHub Projects Worth Your Week

Repo Radar No. 13: graphify turns any folder of code, papers and screenshots into a queryable knowledge graph. Tencent's CubeSandbox boots a hardware-isolated agent sandbox in under 60ms. Together AI's hallmark stops coding agents shipping the same gradient hero. Microsoft's Flint compiles agent chart specs into Vega-Lite, ECharts or Chart.js. PentAGI runs autonomous security tests inside Docker. Five projects that narrow what an agent may see, run, render or spend.

Marcus Schuler · 11 min read ·
Repo Radar: 5 GitHub Projects Worth Your Week
Tools & Workflows

Repo Radar: 5 GitHub Projects Worth Your Week

This week's Repo Radar tracks five GitHub projects where AI agents move from chat into real production work: OpenMontage turns a coding assistant into a video studio, Google Labs' design.md gives agents a design-system spec, Strix runs autonomous penetration tests, Alibaba's page-agent drives live web interfaces in natural language, and MinerU converts messy PDFs into LLM-ready markdown. Difficulty scores, licenses, and push dates for each, plus why OpenMontage is Repo of the Week.

Marcus Schuler · 11 min read ·
Repo Radar: 5 GitHub Projects Worth Your Week
Tools & Workflows

Repo Radar: 5 GitHub Projects Worth Your Week

Repo Radar's eleventh issue tracks five GitHub projects builders attach to AI agents once a demo becomes a workload: Agent-Reach, a CLI giving agents live access to Twitter, Reddit, and YouTube; Flue, the Astro team's sandbox agent harness; cognee, a graph-based memory layer; hunk, a review-first diff viewer for agent-written code; and mistral.rs, a Rust engine for local inference. Each scored on stars, language, license, push date, and setup difficulty.

Marcus Schuler · 11 min read ·
Implicator PRO

The analysis your competitors are reading.

Weekly deep dives into the deals, strategies, and power shifts reshaping the AI industry.

Weekly long-form analysis Exclusive data briefings Early access to reports
Start PRO Subscription
Starting at $7.41/month · Cancel anytime

Weekly deep dives on AI power, strategy, and market shifts.

For founders, operators, investors, and decision-makers who need signal over noise.

Every Tuesday Morning

When Claude Fable 5 Is Worth Double, and When to Use Opus 4.8

Anthropic reports Claude Fable 5 falls back to Opus 4.8 on a fifth of Terminal-Bench trials, and after the July relaunch one tester measured three of four debugging tasks rerouted. The launch drew a researcher revolt over hidden limits; Andon Labs found an alignment slip. Double the price.

Every Tuesday Morning

Give Your AI a Memory Layer That Survives Sessions

Coding agents, support bots, and assistants all restart cold. A bigger context window does not fix it; a memory layer does. A build-along with Mem0 and Qdrant, the contradiction test most demos skip, and the real token math, costs, and controls a production memory store actually needs.

Thinking Machines’ Inkling Takes U.S. Open-Model Lead With 41 Score
Analysis 17 min read

Thinking Machines’ Inkling Takes U.S. Open-Model Lead With 41 Score

Thinking Machines’ Inkling leads U.S. open-weight releases with a 41 index score. Independent tests also found high pricing and a 63% hallucination rate, while the full checkpoint needs two terabytes of GPU memory. The open weights leave companies with a harder deployment decision.

Thinking Machines Lab’s first production model has taken the U.S. open-weight lead with a score of 41 on Artificial Analysis’s Intelligence Index. Inkling uses fewer output tokens than several Chinese rivals, yet testing found high prices and a 63% hallucination rate on one knowledge benchmark. Its weights are free, but the full checkpoint needs at least two terabytes of GPU memory. The test begins after the download: which companies can afford to turn open access into a working system?

Marcus Schuler
Marcus Schuler

AI moves fast. We move first.

Delivered to your inbox at 6 AM Pacific, every weekday.

ESC
The AI Briefing Join Free