Repo Radar #17 covers one repo: JustVugg/colibri, the Apache-2.0 pure-C engine that runs GLM-5.2, Kimi K3 and seven other open mixture-of-experts families by keeping dense layers in RAM and streaming routed experts off NVMe. It passed 33,000 GitHub stars in nine weeks. The community benchmark tables run from 0.05 tokens per second on a 25 GB laptop to 6.84 on six RTX 5090s, with a 4-bit quality cost attached. What teams with a workstation can actually run.
Marcus SchulerSeptember 15, 2026, 9:57 PM PST · 15 min read
Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco.
E-mail: editor@implicator.ai