
Jeff
A 0.8B open System 1 model with swappable LoRA adapters that makes the quick decisions in front of a large local model, answering first and passing the query on only when it is unsure.
GitHub avatar of firelex, not the project’s own logo — taken from github.com on 2026-10-02.

What it is
Jeff is a 0.8B open model meant to sit in front of a larger one and take the fast decisions: it answers first, and only when it is unsure does the query continue to a model like Qwen3.8-27B. One base model is paired with swappable LoRA adapters — nine shipped with the v1.2 community preview on 2026-10-01 — covering guard, triage, support intents, tool choice, grounding, navigation, spam and legal clauses. The README reports 95.3% accuracy against 86.6% for the 27B deciding everything, at 0.25 s per decision instead of 8.1 s, for 1.96 GB of extra memory with all nine adapters loaded.
Who built itThirteen of the repository’s fourteen commits are his, the fourteenth from a second contributor. Nine of the fourteen carry a co-author trailer and all nine name Claude Opus 5.5 at 1M context.
How it is put together
The parts · 5One base model plus adapters, with the routing decision made by the small model itself. Jeff answers first and estimates whether it is sure; only the unsure cases continue to the large model. That inverts the usual cascade — the expensive model is not consulted by default and then filtered, it is consulted by exception — and it is what turns a 6.9% memory increase into a 38× latency reduction on the tasks the adapters cover.
- The v1.2 base
- A 0.8B open model, 1,706,027,688 bytes on disk, released as a community preview on 2026-10-01 with v1.3 announced as the long-term-support base.
- Nine adapters
- LoRA adapters over one base — guard, triage, support intents, tool choice, grounding, navigation, spam, legal clauses and one more in the headline table. The legal-clauses adapter is 41,459,776 bytes including its readout.
- The routing rule
- Jeff answers first and passes the query on only when unsure. The README reports the pass-through rate per task, from 0.0% on most of them to 1.7% on grounding.
- videos/ and assets/previews
- 24 video files at 58.8 MB and 2.9 MB of preview images — the larger part of a repository whose product is a 1.7 GB model, and the part that shows what the adapters do.
- docs/data-sources.md
- Where the training data comes from, at 7 KB, next to the published data guidelines that let somebody else’s dataset survive a base-version change.
Choices, and what they beat
Put the small model in front and pass on by exception over asking the large model every time
The README’s tables are the argument: on the same rows, 95.3% against 86.6%, at 0.25 s instead of 8.1 s, while the large model is consulted for 0.0% of most task types.
Release as a community preview and name the next base rather than waiting over holding the adapters until v1.3 is stable
The README announces v1.2 and nine adapters, asks for issues, and says v1.3 is about thirty-six hours away — together with the warning that adapters do not survive a base change but data sets do.
Keep adapters swappable over one base over one fine-tuned model per task
Nine adapters cost 1.96 GB in total on top of a 28.6 GB setup, which is the number that makes the cascade affordable rather than a second large model.
Read fromREADME and the repository tree of firelex/jeff, read 2026-10-02.
Build log
4 stages- 01
Four days from an empty repository to a preview with nine adapters
The repository was created on 2026-09-28 and holds 14 commits, twelve in September and two in October. It is published as a community preview rather than a release: the README announces Jeff v1.2 with nine adapters, asks readers to file what works as issues, and states that a stable long-term-support base, v1.3, is due in about thirty-six hours with the official adapters retrained shortly after. It also states the constraint that follows from that schedule rather than burying it — adapters do not carry between base versions, so they have to be retrained, while the data sets do carry over if they follow the published guidelines.
- 02
Who wrote it: nine commits out of fourteen carry a model’s name
Thirteen of the fourteen commits are the author’s and one is a second contributor’s. Nine carry a co-author trailer and all nine name Claude Opus 5.5 at 1M context — two thirds of the history. That is the highest proportion in this batch, and it is worth stating what it does and does not mean: it is evidence that the model wrote a substantial part of the code in those commits, not evidence about the training data or about how the benchmark numbers were produced.
- 03
The claim is a comparison, and it is printed in full
The README does not argue that a small model is good; it puts the two configurations side by side on the same test rows. Across eight adapters, the 27B deciding everything scores 86.6%, and Jeff with adapters deciding while the 27B handles only the unsure cases scores 95.3%. Time per decision falls from 8.1 s to 0.25 s, wrong answers from 13.4% to 4.7%, and the memory cost is stated as a number rather than a claim: +1.96 GB with all nine adapters loaded, 6.9% on top of 28.6 GB. A second table breaks the same comparison down per task — guard 84.0% to 98.0%, triage 81.3% to 91.0%, support intents 86.0% to 95.3%, tool choice 90.3% to 98.0%, spam 88.0% to 98.7% — and one row is the interesting one: grounding moves from 96.7% to 96.3%, the only task where the small model is not better, and the column that says why is the one showing 1.7% of queries passed on.
- 04
What is in the repository, and what is not
The 182 files are dominated by media rather than code:
videos/accounts for 24 files and 58.8 MB, andassets/previewsanother 2.9 MB, against adocs/data-sources.mdof 7 KB that documents where training data comes from. The model’s own size is given in the README as a byte count — the v1.2 base is 1,706,027,688 bytes, and the legal-clauses adapter 41,459,776 bytes for weights plus its readout. Between the videos, the adapter sizes and the data guidelines, the repository’s shape says what it is for: showing what the adapters do, and telling someone how to build their own.
Adjacent records
All records →No. 141
Kev
A family of small decision models built on Qwen3.5/3.8 that answer narrow questions with calibrated probabilities, trained and run on hardware you own — and 134 of its 333 commits carry Devin’s name.
No. 131
agent-memory
A long-term memory runtime for AI agents that keeps plain Markdown files as the single source of truth, ranks them locally without calling a model, answers recall with file paths the agent opens one level at a time, writes at conversation boundaries rather than on the agent’s initiative, and runs an independent sleep-time layer that may add and update on its own but can only ever file a deletion as a proposal — one store shared by Claude Code, Codex CLI and Hermes, with no API key.
No. 128
Easel
An open-source content workbench for social media creators: one agent runs the whole loop — aggregate the hot lists, plan a topic, generate the copy, the cards, the voice and the video, publish the finished file to an account that is already logged in on seven Chinese platforms, then read the numbers back into the account profile that shaped the next round.