Colibrì Runs 744B- to 2.8T-Parameter MoE Models on 25 GB Machines by Streaming Experts From Disk
Colibrì now runs MoE models as large as Kimi K3’s 2.8 trillion parameters on machines with just 25 GB of memory by streaming routed experts from disk, though cold GLM-5.2 inference falls to 0.05–0.1 tokens per second.