Skip to content

Engram

A learning engine that installs into a coding agent: a curriculum architect breaks a topic into a first-principles concept map, a tutor makes you predict, attempt and explain before it explains, a blind assessor grades your verbatim free recall and writes a receipt for every verdict, and a deterministic FSRS-4.5 core in one Python file decides when each concept comes back — with explorable HTML built only for the concepts whose content rewards manipulation.

Screenshot of Engram
Editor screenshot, 1 Oct 2026Engram ↗

What it is

Engram is a learning system that installs into a coding agent. Given a topic it builds a first-principles concept map, makes the learner predict, attempt and explain each concept before it is explained, then hands the learner’s verbatim answers to a separate agent that grades them blind: the assessor receives the claim, the rubric, the probe, the production and a confidence the learner picked beforehand, and never sees the lesson. Every verdict is written to disk as a receipt, and nothing moves state without one. One stdlib-only Python file carries the FSRS-4.5 scheduler, the state machine, the statistics and 315 self-checks; it computes every date and stability value itself and contains no network code, which a permanent self-test proves by parsing the file’s own syntax tree. Threshold concepts get interactive HTML explorables built to a seven-clause contract. The same skills and engine run on nine agent platforms, from Claude Code, where it began, to Codex, Antigravity and DeepSeek Harness.

Who built itA single maintainer who works under a handle rather than a name: the GitHub profile carries 38 public repositories, 75 followers, an account opened on 2022-01-03 and a three-word bio, and nothing else. He wrote 189 of the repository’s 204 commits, under two commit names, and the README’s closing section lists four sibling plugins from the same workshop — one of which it credits as the source of Engram’s verification patterns: oracle-driven loops, receipts and re-anchoring.

How it is put together

The parts · 6

The organising idea is that the model may talk and the engine may not be talked out of anything: every number a learner is shown about their own memory is computed by one stdlib-only Python file, and no state advances without a written receipt from a grader that never saw the lesson. That splits the system into a tutor that runs in the main conversation and is forbidden to grade, an assessor spawned in a fresh context with only the rubric and the learner’s words, a coach that may adapt only from receipts, and a scheduler that owns every date. The learner’s whole state is human-readable JSON in one directory, so the tutor is re-anchored from disk at the start of a session instead of trusting its own memory of the learner. The second idea follows from the first: because the grader is the load-bearing instrument, it is itself audited against a public adversarial gold set, and the engine refuses to certify a number it cannot defend — a fitted parameter that does not beat the current one is discarded, and export refuses to run behind an unaudited grader rather than warning about it.

scripts/engram.py
783,600 bytes, and the only part of the project that decides anything: one stdlib-only command-line program carrying the FSRS-4.5 scheduler, the state schemas, the append-only receipt log, the crash-safe stash, the adherence and retention arithmetic, the self-contained HTML dashboard, the grader-audit statistics and 315 self-checks — pinned by a build-time check that parses the file’s own syntax tree and fails if a network import is ever added.
skills/ and agents/
Three skill files — learn at 26,845 bytes, review at 24,600, coach at 40,217 — plus a shared directory holding the dialogue grammar, the subagent-spawning rules, the problem grammar and the Explorable Contract; then three Claude Code subagent definitions and three TOML ports of the same three for Codex. The assessor’s specification is the longest of them at 13,463 bytes and reads like a grading manual: list what is missing before crediting what is present, round down when torn, verify steps by running them, and never let a cap become a floor.
hooks/
The entire ambient surface: one SessionStart script that loads the learner model and the due queue from disk, prints a one-line nudge when reviews are due and says nothing at all otherwise. It exists in several variants because the hosts speak differently — a TypeScript pair for OpenCode, a shell script for Hermes, a JSON-emitting wrapper for DeepSeek Harness, and an OpenClaw hook pack — and there is deliberately no session-end hook, because receipts are written at the moment of grading and nothing needs flushing.
The platform layer
Nine hosts, each with thin glue rather than a fork: .claude-plugin/ with a marketplace entry, .codex-plugin/ and .agents/plugins/ for Codex’s own marketplace, .opencode-plugin/ at 106 KB where the V1 and V2 entry points share a single npm package, .zcode-plugin/, a pi/ extension with three prompt templates, and dsh/ with a hook bridge and a patch block. One engine, many hosts, and the differences kept in manifests instead of branches.
gold/ and experiments/
The instrument and the trials: assessor-gold.jsonl at 103,538 bytes holding 89 adversarial items with their answers stripped by construction and shaped exactly like a real settle payload, so an audit grades the real assessor rather than a special case; three experiment presets (contrast-first, probe-variation, topic-reconstruction) feed the pre-registered n-of-1 trials whose verdicts only the engine is allowed to compute.
docs/ and the release machinery
Seventeen documents at 451 KB — foundations, prior art, architecture, a vision, an audit of the visual-encoding literature, two target architectures, two roadmaps — beside a 54,318-byte release protocol, a 276,967-byte changelog, six release audits and five written user sessions. The protocol is the most revealing artifact here: seventeen gates in order, each one added because it caught a bug the gate before it could not see, and a numbered list of the seven bug classes the project is not allowed to ship.

Choices, and what they beat

  • The tutor is the main conversation, not a subagent over a tutor agent with its own context

    The architecture document states the reason: the tutoring relationship has to persist in context, and only grading needs fresh-context isolation — which is why the assessor was made a subagent and the tutor was not.

  • Scheduling arithmetic in code rather than in the model over letting the model compute intervals and dates

    Stated as a rule: FSRS maths, state validation and schema migration are deterministic, so they run as code and never as model arithmetic. The engine owns every number, and the lesson inherited from the project this one borrows its verification patterns from is that everything checkable must be checked by something executable.

  • Menus for navigation, never for knowledge over the parent project’s arrow-key pattern applied to questions as well

    Written as the deliberate inverse of the parent design: logistics may be arrow keys, but retrieval is always open-ended production, because a multiple-choice question measures recognition rather than recall.

  • A whitelist rather than a blacklist for the shared export over deleting the learner’s own words before writing the file

    A blacklist is a promise that has to be kept on every release, while a whitelist is one kept by construction: every field is built by name, so no code path exists by which a production could arrive. The stripped-field list ships inside the exported file, and the export refuses — rather than warns — when the grader is unaudited.

  • Attribution rather than anonymity in the Commons over a salted anonymous hash in the payload

    The README’s argument: a salted anonymous hash riding inside a signed envelope would be theatre the moment the envelope is signed, and one-keystroke upload and anonymity cannot both be had, so the project picks one and says which. Attribution is also described as the stronger science — a retention study lives on following the same learner across months, so attributed n=100 beats anonymous n=500.

Read fromdocs/03-architecture.md (17,199 characters), docs/09-target-architecture.md (28,990 characters), skills/_shared/explorable-contract.md, agents/engram-assessor.md, RELEASE_PROTOCOL.md and README.md.

Build log

6 stages
  1. 01

    Fifty-five releases in fifty-three days, under two commit names

    The repository was created on 2026-07-05 and its last commit is dated 2026-08-27: 204 commits in fifty-three days, 162 of them in July and 42 in August. They resolve to four accounts — nagisanzenin with 182, nagisanzeninzz with 7, luanweslley77 with 11 and mertso13 with 4 — while the author’s own commits are signed with two different names, nagisanzenin for 124 and quan-vin for 65, and 180 of the 204 share one Gmail address that a contributor greets in issue #10 as “Hi Quan”. The project published 55 releases and 55 tags. On 2026-07-11, fourteen releases landed between 06:00 and 13:06 — v0.6.0 through v0.6.4 inside forty-one minutes, then v0.7.0 through v1.0.0 over the next four hours. On 2026-07-24, nine releases — v1.3.0 through v1.9.1 plus the 2.0 release candidate — were published inside forty seconds of one another, each titled with that same date. The repairs are as fast as the features: the release titled “what the post-release review caught” follows the release it corrects by thirteen to eighteen minutes, four times over. Of the 204 commits, 144 carry a co-author trailer and 143 of those name a Claude model — Fable 5 on 90, Opus 4.8 on 43, Opus 5 on 6. Around it: 1,439 stars, 105 forks, 5 watchers, 4 open issues, MIT and 2,221 KB, most of it a 783,600-byte scripts/engram.py, which is why GitHub labels the repository Python.

  2. 02

    What a receipt contains, and what it takes to earn one

    The evidence that you really recalled something is a JSON object on disk, written at the moment of grading. A receipt carries the topic and the node, the kind of encounter (encode, review, transfer, audit or pretest), the probe that was asked, the learner’s production verbatim, the confidence they stated before any sign of correctness, the grade (recalled, partial or lapsed) and the rating it maps to, one line per distinct misconception, a criterion-by-criterion note quoting the rubric, and the scheduler’s numbers: stability before and after, the interval in days, the retrievability estimate and the next due date. Two fields exist purely as guards. sid, the stash id, rides stash to assessor to receipt so that re-applying a settle file is a no-op instead of double-counting a review; grader holds the assessor’s stable identity, and the specification forbids a model to name its own weights. What makes a receipt count is structural: the assessor never sees the lesson, only the claim, the rubric, the probe and the words, and no state advances without a receipt — a node’s state, fsrs and artifact fields are stripped from anything an external payload supplies. The README offers its own worked example as the argument: a session the tutor thought went well came back 1 recalled, 4 partial and 1 first-retrieval, and the schedule believed the assessor.

  3. 03

    The scheduler is code, and a fit that loses is thrown away

    Everything that looks like a date is computed by engram.py and never by the model. The README frames it as the gap a chat cannot close — no memory of you, no test of whether you got it, and no plan for the forgetting that starts the moment you close the terminal — so the plan is FSRS-4.5, with per-node stability and difficulty, a target retention of 0.90 and one fit per learner. refit starts with a single interval multiplier once there are fifty usable reviews, then fits the learner’s own parameters: initial stability at 64 usable reviews, the full weight vector at 400, and a fit that does not beat the current one is refused rather than shipped. The first issue in the repository is a reader’s code reading: next_difficulty mean-reverted toward init_difficulty(4) where FSRS-4.5 reverts toward init_difficulty(3), so the engine had been quietly running FSRS-5’s difficulty rule under a 4.5 label; it was fixed in v0.3.0 and pinned by a fixed-point self-test. The same report caught that a receipt was written after the graph was saved, so a crash could leave an unverified advance; the receipt is now written first, making the worst case a harmless re-review. The first two intervals after encoding are capped so at least three spaced sessions land inside the first month, and a capped queue is ranked by expected retention saved per minute rather than by most overdue.

  4. 04

    The grader is graded, and the gold set failed first

    Because the verdict drives mastery, retention, calibration and the schedule itself, the project ships an adversarial gold set of 89 items — fluent-but-empty, terse-but-correct, confident-and-wrong, right-answer-wrong-reason and three analogy-alignment traps added in v1.14 — and reports one number: 0 of 258 blind judgments where the grader awarded more credit than a strict reading of the rubric, earned across three runs on the 86-item version, guarded by a stale-gold check that expires the badge when the set changes. How it arrived at that number is the part worth keeping. v0.7.0 shipped a QWK 0.93 badge, then a post-release reviewer graded the gold set with a deliberately fooled grader; the fooled grader scored higher, 1.000 against 0.990, because five lenient adjudications by the set’s own author credited adjacent facts as partial credit. Correcting them moved agreement from 0.889 to 0.965, and the engine now prints the circularity on every audit: the corrections were prompted by the grader’s own disagreements, so the agreement that follows measures the author’s willingness to concede. One genuine disagreement is deliberately left in, with the reasoning published: an instrument with no disagreement left in it measures nothing. The engine refuses to certify a grader on consistency alone, because a judge can be reproducible and systematically wrong at once.

  5. 05

    Explorables built to a contract and registered on the graph

    Interactive HTML is generated for the concepts whose content rewards manipulation — a parameter to drag, a process that unfolds — and never because a learner called themselves a visual learner. Each artifact must satisfy a seven-clause contract: nothing unlocks before a committed guess; at least one guided manipulable model runs the predict, act and explain cycle, with a worked drive first for novices; two free-recall prompts sit inside it with their answers hidden until attempted; decoration is zero and text never runs over motion; the file is self-contained, opens offline and stays under about 120 KB; it ends by asking the reader to rebuild the argument from nothing; and its header records the node, the date and the learner-model inputs it was built from. The file is registered on the graph with artifact set, which validates that it exists, and a receipt stamps whether an artifact existed at grading time — which is what stats.modality compares, explorable-encoded recall against dialogue-only recall, with at least six items per arm. examples/ ships two of them, one generated by the artifact-smith in a real v0.5.0 session and one hand-authored as the reference implementation; both are offered with a hosted link that returns 404 at the collection date, and the only other trace of that publishing route is a zero-byte .nojekyll at the repository root.

  6. 06

    Two outside contributors, and five weeks of quiet

    Four accounts have committed here and two of them are the author’s. luanweslley77 wrote eleven commits and five pull requests, several through multiple rounds of review — moving the plugin source out of the directory it extracted into, replacing the instruction file with a versioned marker block, and twice proposing a deletion that would have taken user data with it. mertso13 added Antigravity support and produced agy plugin validate output from version 1.1.4. The sharpest writing is in the bug reports: tyteachestech found that stash add clipped productions at 800 characters before the blind assessor ever saw them, so a thorough answer lost two rubric criteria (fixed in v1.11.2, with the ceiling raised to 2400); DorusKeijzer reported that probes and rubrics routinely disagreed, which became the probe_gap field and a rubric-repair command in v1.10.0; and SK-DEV-AI filed two defects against the OpenCode install with file and line numbers, both fixed in v1.13.2 after the maintainer reproduced the reporter’s counts. The maintainer files against himself as well, and one release is titled “a regression my own fix caused”. The last commit is 2026-08-27. An issue arrived on 2026-09-22 and has no reply, four issues are open, and nothing is archived: five weeks of quiet after eight weeks at that pace is why this record says maintained rather than active.

Adjacent records

All records →