inference
2 projects
AI & ML
Colibrì runs 744B to 2.8T models off your SSD, in C
JustVugg/colibri
An engine treating VRAM, RAM and disk as one hierarchy, with a hard rule: run out of fast memory and it slows down, it never silently changes the model.
★ 24.9K
· C
· August 22, 2026
AI & ML
A 2.78-trillion-parameter model in 8 GB of RAM, in C99
FareedKhan-dev/kimi-k3-in-c
kimi-k3-in-c streams a 1.56 TB checkpoint off disk to run Kimi K3 on an ordinary laptop CPU. No BLAS, no framework, no GPU, 176 KB of engine.
★ 5.7K
· C
· August 17, 2026