
Colibrì
An inference engine in pure C with no engine dependencies that treats storage, RAM and VRAM as one hierarchy, so 744B to 2.8T-parameter mixture-of-experts models run on hardware people already own.
GitHub avatar of JustVugg, not the project’s own logo — taken from github.com on 2026-10-02.

What it is
Colibrì runs frontier mixture-of-experts models on consumer and heterogeneous hardware by treating VRAM, RAM and storage as a single memory hierarchy rather than assuming a model must fit in VRAM. Nine families run today — GLM-5.2/5.3 at 744B, GLM-5.3-Flash at 321B with vision, Inkling at 975B, Kimi K3 at 2.8T, DeepSeek V4 Flash at 284B and V4.1 Flash at 552B, Qwen3.8-Flash-Next at 125B, Qwen3.6 at 35B-A3B and OLMoE at 7B — each as one C file behind the same coli chat, coli serve and coli web front end.
Who built itThe repository accumulated 2,803 commits between July and October 2026, and the two largest shares belong to Vincenzo Fornaro (837) and the GitHub account JustVugg (593). Behind them the history is unusually wide for a project this young — ZacharyZcR 216, monotophic 141, Vincenzo 63, Steve Markgraf 60, woolcoxm 55 — in a contributor list that runs to well over a hundred names.
How it is put together
The parts · 5The engine is built around one inversion: instead of asking how to fit a model into VRAM, it treats storage, RAM and VRAM as one hierarchy and moves experts through it. That makes the model a set of files read on demand rather than a resident allocation, and it makes the interesting problems I/O placement, scheduling, kernels and CPU/GPU overlap rather than quantisation. Everything is pure C with no engine dependencies, so the whole of it is in one directory of files, one per supported model family, behind a shared front end.
- c/ — the engine
- 126 files and 6.3 MB: the shared runtime, the memory-tiering layer, and one C file per model family — GLM, Inkling, Kimi K3, DeepSeek V4 and V4.1, Qwen3.8 and Qwen3.6, OLMoE.
- c/tests/
- 390 files, 3.6 MB — the majority of the file count, and the mechanism behind the promise that an experiment may cost speed but never silently change precision or router semantics.
- The three front ends
coli chat,coli serveandcoli web, shared by every model. Which model is running is a configuration of the same program rather than a separate build.- docs/
- Per-topic and per-model documentation: environment and setup at 61 KB, the Metal backend at 47 KB, formats at 38 KB, Windows at 35 KB, benchmarks at 28 KB, and one document per model family — the largest being
qwen38.mdat 29 KB. - web/ and the project site
- The static site at
justvugg.github.io/colibriand theweb/directory behindcoli web, which is how a reader can see the engine without building it.
Choices, and what they beat
Treat storage, RAM and VRAM as one hierarchy over requiring the model to fit in VRAM
Stated as the central idea: models of 744B to 2.8T parameters run on consumer and heterogeneous hardware because experts are streamed from disk rather than held resident.
Guarantee semantics and refuse to guarantee speed over degrading precision when fast memory runs short
The README makes it a rule — no SLA on speed, a hard guarantee on semantics — and says why: insufficient fast memory may reduce speed, but it must not quietly redefine the model.
Pure C with zero engine dependencies over building on an existing inference runtime
One C file per model family behind a shared front end is only possible if the engine owns its own formats, I/O and kernels; the README frames the absence of dependencies as what makes the memory hierarchy testable.
Read fromREADME, the repository file tree and the per-topic documents under docs/ of JustVugg/colibri, read 2026-10-02.
Build log
4 stages- 01
A whole engine in three months, and who wrote it
The repository was created on 2026-07-01 and has been pushed to every day since; 2,803 commits landed in July (1,028), August (985) and September (790), across 857 files. The authorship is the part a reader should look at first: 591 commits carry a co-author trailer and roughly 575 of them name a Claude model — Fable 5 on 166, Opus 5 on 153, Opus 4.8 on 85, Opus 5 (1M context) on 57, Fable 5.1 on 41, Opus 4.8 (1M context) on 20, and smaller counts for Opus 5.5, Sonnet 4.6 and Sonnet 5, with Claude Code, Claude and Cursor on a handful more. Sixteen trailers name people, most of them the maintainers. The commits are not only attributed at the trailer level either: several appear under machine-generated author names of the form
codex 20260923-120015-1331522267. - 02
The claim is about where the model lives, not how fast it runs
The README states the design in two sentences: frontier MoE models of 744B to 2.8T parameters run on consumer hardware by treating storage, RAM and VRAM as one inference hierarchy, in pure C with zero engine dependencies. The project calls this AI memory multitiering and frames itself as a research platform first. The commitment that follows from that framing is stated as a rule rather than a hope: there is no SLA on speed, and a hard guarantee on semantics — experiments have to earn their place through reproducible end-to-end measurements, and the default policy never silently changes model precision or router semantics. Insufficient fast memory may cost speed; it may not quietly redefine the model.
- 03
Nine model families, one C file each
What runs today is listed in the README with parameter counts: GLM-5.2/5.3 at 744B, GLM-5.3-Flash at 321B with vision, Inkling at 975B, Kimi K3 at 2.8T, DeepSeek V4 Flash at 284B, DeepSeek V4.1 Flash at 552B with vision, Qwen3.8-Flash-Next at 125B plus a 51B n-gram, Qwen3.6 at 35B-A3B, and OLMoE at 7B. Each family is one C file, and all of them sit behind the same three front ends, so the model is a configuration of the engine rather than a fork of it. The documentation follows the same split:
docs/FORMATS.mdat 38 KB,docs/ENVIRONMENT.mdat 61 KB,docs/metal_implementation.mdat 47 KB,docs/windows.mdat 35 KB,docs/benchmarks.mdat 28 KB, and one document per model family —qwen38.mdat 29 KB,deepseek-v41.mdat 26 KB,deepseek-v4.mdat 25 KB. - 04
Tests are the largest thing in the repository
Of the 857 files, 390 sit in
c/testsand account for 3.6 MB against 6.3 MB of engine code inc/— a ratio that follows from the semantics guarantee, since a rule about never silently changing precision is only as good as the checks that would notice. The contributor list is also unusually broad for a project three months old: alongside the two maintainers there are more than a hundred named contributors, from twenty- and sixty-commit contributors down to a long tail of single commits.
Adjacent records
All records →No. 131
agent-memory
A long-term memory runtime for AI agents that keeps plain Markdown files as the single source of truth, ranks them locally without calling a model, answers recall with file paths the agent opens one level at a time, writes at conversation boundaries rather than on the agent’s initiative, and runs an independent sleep-time layer that may add and update on its own but can only ever file a deletion as a proposal — one store shared by Claude Code, Codex CLI and Hermes, with no API key.
No. 077
Engram
A learning engine that installs into a coding agent: a curriculum architect breaks a topic into a first-principles concept map, a tutor makes you predict, attempt and explain before it explains, a blind assessor grades your verbatim free recall and writes a receipt for every verdict, and a deterministic FSRS-4.5 core in one Python file decides when each concept comes back — with explorable HTML built only for the concepts whose content rewards manipulation.
No. 068
GSD Core
Git. Ship. Done. — a meta-prompting, context-engineering and spec-driven development framework that runs the same five-step loop on every milestone: discuss, plan, execute, verify and ship. The heavy work is pushed into fresh-context subagents so the main session stays lean, and every decision is written into Markdown and JSON under a planning directory instead of living in the conversation.