Skip to content

Kev

A family of small decision models built on Qwen3.5/3.8 that answer narrow questions with calibrated probabilities, trained and run on hardware you own — and 134 of its 333 commits carry Devin’s name.

GitHub avatar of jaredpalmer, not the project’s own logo — taken from github.com on 2026-10-02.

Screenshot of Kev
Editor screenshot, 2 Oct 2026Kev ↗

What it is

Kev is a family of small decision models built on top of Qwen3.5/3.8 that you train and run yourself, aimed at the narrow questions an application asks over and over rather than at open-ended generation. The repository carries its own evaluations — evals/round5 and evals/devtools-v1 at 14.6 MB and 12 MB — alongside 22 KB of notes on automated research, in a tree of 3,049 files that is mostly run artefacts rather than library code.

Who built itThe repository’s 333 commits are almost entirely his — Jared Palmer on 312, the Devin integration bot on 12 — with a single commit each from eight other people. 145 commits carry a co-author trailer and 134 of them name Devin: 127 as Devin, six as Devin AI and one as the integration bot. The remaining eleven name people or a Claude model.

How it is put together

The parts · 4

The work is organised as rounds rather than as a library. Each round is a directory holding its own evaluation data, results and runs, so a change in the recipe is measured against the previous round instead of against a moving baseline. The models themselves are small enough to be trained locally on top of Qwen3.5/3.8, and the repository keeps the research notes — 22 KB of them — next to the artefacts they produced.

evals/round5
Nine files and 14.6 MB — the evaluation data and results of the fifth round, the largest single directory in the repository.
evals/devtools-v1
Four files and 12 MB for the developer-tools evaluation, kept as its own versioned set rather than folded into a shared test harness.
runs/r5-combined
The run outputs behind round five — a project sixteen days old that already keeps per-round run artefacts rather than a single result file.
docs/autoresearch.md
22 KB of notes on the automated research loop, sitting next to the artefacts that loop produced.

Choices, and what they beat

  • Keep every evaluation round as its own directory over one test suite overwritten each time

    The tree holds round-five data at 14.6 MB beside a separate devtools evaluation at 12 MB, so a result can be compared with the round before it after the fact.

  • Train small models for narrow decisions over prompting a large model for every call

    The README frames the family as models you train and run yourself, which places the cost at training time rather than at every request.

  • Build on Qwen3.5/3.8 rather than from scratch over a bespoke base model

    Stated in the one-line description: the family is built on top of those bases, so the work is the decision task and the training recipe rather than the foundation model.

Read fromREADME, docs/autoresearch.md and the repository tree of jaredpalmer/kev, read 2026-10-02.

Build log

3 stages
  1. 01

    Three hundred and thirty-three commits in sixteen days, most of them Devin’s

    The repository was created on 2026-09-17 and has been pushed to every day since: 333 commits, of which 322 landed in September and eleven in the first days of October. The authorship is the part worth reading first. 145 commits carry a co-author trailer and 134 name Devin — 127 as Devin, six as Devin AI, one as the integration bot — against eleven naming a person or a Claude model. Jared Palmer holds 312 of the 333 commits himself. The reading is not ambiguous: this is one person directing an autonomous agent through a fast research loop, and the trailer is the record of it.

  2. 02

    A repository shaped like a series of experiments

    Of 3,049 files, the largest directories are results rather than code: evals/round5 at nine files and 14.6 MB, evals/devtools-v1 at four files and 12 MB, and runs/r5-combined behind them. A tree that size, in a project sixteen days old, is what an iterative evaluation loop leaves behind — each round keeping its measurements rather than overwriting them, so a comparison between rounds is possible after the fact. The documentation concentrates in one place: docs/autoresearch.md at 22 KB.

  3. 03

    What it is for, and what it refuses to be

    The README’s one-line description is narrow on purpose — small decision models you can train and run yourself, built on Qwen3.5/3.8 — and the name it reaches for situates it against a known shape: a Jev-like family. The emphasis is on owning both ends: the training and the running happen on hardware the reader has, which is the same position several other records in this batch take, arrived at from the opposite direction — not by shrinking a large model’s runtime, but by training small models for questions that do not need one.

Adjacent records

All records →