agent-memory
A long-term memory runtime for AI agents that keeps plain Markdown files as the single source of truth, ranks them locally without calling a model, answers recall with file paths the agent opens one level at a time, writes at conversation boundaries rather than on the agent’s initiative, and runs an independent sleep-time layer that may add and update on its own but can only ever file a deletion as a proposal — one store shared by Claude Code, Codex CLI and Hermes, with no API key.

What it is
A long-term memory runtime for AI agents whose knowledge lives in plain Markdown files, with every index beside them treated as a cache that can be thrown away: rm -rf .index/ && mem rebuild is enforced by a test to lose zero knowledge. One memory is one file, carrying a stable name, a one-sentence abstract, a type with schema fields, timestamps, links, weight and provenance, and a validity interval instead of a status field, so a replaced file stays in the store and recall --as-of can still answer from it. Retrieval is local and ranked — BM25 over an FTS5 index by default, an optional vector index fused in by reciprocal rank — and it answers with paths rather than pasted text: recall returns a one-line list, and the agent opens a hit only as deep as the task needs. Writes fire at conversation boundaries instead of waiting for the agent, the full trace is copied before distillation, and an independent sleep-time layer consolidates on its own clock and files deletion as a proposal. Claude Code, Codex CLI and Hermes share one store, and no API key is needed.
Who built itAn organization account rather than a person: the repository carries MIT copyright to Tigerless Labs, and eight accounts wrote its 154 commits between 2026-09-01 and 2026-09-28. One member dominates the history — the account liruihan000 accounts for 103 of them, committing under three separate addresses and two display names — with faj-design5260 on 33 and importcpp on 7; the other five accounts have one or two each. Ninety of the 154 commits carry a co-author trailer, and every one of those names a Claude model: Opus 5 on 59, Fable 5.1 on 28, Opus 5.5 on 2 and Fable 5 on 1.
How it is put together
The parts · 6A runtime split into three layers that are allowed to fail independently: Markdown files as truth, a ranked local index over them, and a management layer on its own clock. The consequence that shapes everything else is that the index is disposable — the architecture is built so that rm -rf .index/ && mem rebuild loses zero knowledge, enforced by a test rather than promised in a document — and that a memory is a file with a validity interval rather than a row with a status. Writes have a single path, so agent writes and management rewrites go through the same validate, hash-diff and reindex pipeline, and reads never mutate truth: usage statistics land in an access log and weight is settled back into frontmatter in batch. Retrieval keeps three traditions in one store — relations as links inside the memories, a local FTS5/BM25 index with an optional vector plugin, and an ordinary directory an agent can walk — so a miss on one track is not a miss. The library core contains no model client; judgement is borrowed from the host agent’s own command-line interface, which keeps the core testable without a network and keeps every write visible in the user’s own transcript. Above it sit six packages with an invariant that the adapters carry no algorithm, so the same request through any entry yields the same result.
- packages/core/ — 44 files, 212 KB
- The truth and retrieval layers.
store.pyat 27.6 KB is the single write path;manage.pyat 31 KB andprompts.pyat 17 KB are the sleep-time layer and the prompts it reasons with;reconcile.py,schema.py,frontmatter.py,placement.py,record.pyandmigrate.pyhold the memory model;recall.py,indexer.py,search_index.py,vector_index.py,chunking.pyandembeddings.pyhold ranking;ledger.py,archive.py,distill.py,reasoning.py,sessions.py,trace.py,observation.py,locking.pyandwatermark.pyhold the rest of the machinery. - packages/cli/, packages/mcp/, packages/adapters/
- Three ways in that collapse into the same core calls: a 22.6 KB command-line entry point, an MCP server whose tool module is 10.6 KB and exposes nine memory tools, and host adapters of 6.1 KB of hook entry plus setup, capture, moments and transcript modules. The invariant is that adapters carry zero algorithm, so a request through any entry yields the same result.
- packages/executor/
- The only place a model is called, kept out of the library core: host dialects, reasoners, the distiller and credential handling, reasoning through the host’s own CLI by default or a configured endpoint. This is also what the fresh-install fix re-pointed, from a hosted default inside the organization to the firing host.
- packages/harness/ — 17 files, 86 KB
- The measurement apparatus: an exam driver of 25.5 KB, judge and host-system modules, a benchmark loader, dataset, metrics, report, coverage, interop, framing and sampling. It is what turns claims about recall and about write strategies into recorded arms with fixed host and judge pairs.
- tests/ — 33 unit files, 11 system files
- 205 KB of unit tests and 93 KB of system tests, plus a red-team file of 6 KB for memory-poisoning payloads and a read-exposure gate in the tools directory. The system tests cover entry equivalence, concurrency, the fixed exam, interoperability between hosts and the management entry points.
- docs/ and skills/
- Four design documents — management operation boundaries, batch-write boundaries, raw evidence reads and an index — eight plan documents of 18 KB in total, a 8.9 KB architectural document at the root, a 4.7 KB skill for agents, and a 510-byte continuous-integration workflow that runs the tests, the linter and the type checker.
Choices, and what they beat
Markdown files as the truth and every index as a cache over a store whose primary form is a database
Stated with its reasons: migration freedom, git-ability, compatibility with a host’s own automatic memory, and user sovereignty over their own knowledge. It is enforced by a test rather than asserted, and rebuilding the index from the files must lose zero knowledge.
A validity interval on the file instead of a stored status over a status field with a third state
A file is active or invalid and nothing in between, and invalidation comes only from replacement or deletion; the repository’s reasoning is that a third state that changes nothing is a lie, and that partial staleness inside one file poisons the whole file. Replaced and deleted files stay in the store for
recall --as-ofand trace, and a stored status that older files carry still loads.Borrow judgement from the host agent’s own command-line interface over shipping an LLM client inside the library
Zero keys to install and no billing surface, one extraction pipeline for every host, and a core that stays testable without a network — with the side benefit that every distillation and every rewrite is visible in the transcript the user is already reading.
Keep vector retrieval optional, off by default over making the vector index the default retrieval path
Measured before deciding: optional fusion improved Recall@5 from 79.0% to 86.6% on a fixed 120-query set while median retrieval latency rose from 5.1 ms to 139.2 ms, and answer-level replays scored 17/24 against 18/24 and 28/36 against 27/36 — the answer-level experiments did not establish an end-to-end accuracy gain, so BM25 stayed the low-latency baseline.
An unattended pass may add and update, and may only propose a deletion over a consolidation pass with delete authority
Memory poisoning is treated as a persistent attack surface, and with no human in the loop the repository names reversibility, rate limits and audit as what keep an unattended run safe: rule-only work applies, executor work files a proposal, each kind is capped per sleep, one sleep is one commit, and physical removal is a human-run command management cannot reach.
Read fromCLAUDE.md (8,884 characters), README.md (14,855 characters), skills/agent-memory/SKILL.md, docs/design/management-operation-boundaries.md, docs/design/batch-write-boundaries.md, docs/design/raw-evidence-reads.md, the eight plan documents under docs/plans/, the bodies and comment threads of the issues and pull requests, .github/workflows/ci.yml (510 bytes), and the complete 149-file tree with sizes.
Build log
6 stages- 01
Twenty-nine days, 154 commits, and no release
The repository was created on 2026-09-01, and its oldest commit — logged four seconds earlier the same evening — is titled “Scaffold repo: CLAUDE.md, design tree, roadmap v0.1”: invariants and a plan tree before any code. All 154 commits fall inside that September, the newest being the merge of pull request 46 on 2026-09-28, with the last push recorded on 2026-09-30. There are no releases and no tags at all, and the README says there is no release on PyPI yet, so installation is from a checkout; the version badge still reads 0.1.0. The month produced attention — 1,974 stars, 124 forks, 66 watchers, ten open issues in the metadata — and a burst of work on the front door. Pull requests 21 to 25 all landed on 2026-09-08 and all of them are README or licence work. One added an install section after noticing that
mem setupwrites a baremem-hookcommand into the host configuration whileuv synconly puts that executable inside the virtual environment, so a session started outside it recorded nothing. Another added a section demonstrating the read path in three commands and paid for it by cutting a section whose claims were made elsewhere. The licence change records that the repository had been public with no licence file, which leaves it all rights reserved by default, and adds MIT under Tigerless Labs to the root and all six package manifests. - 02
The file is the atom, and the index is disposable
The design rules are written down in
CLAUDE.mdas invariants a change must not break: Markdown is the single source of truth and every index is a rebuildable cache; agent writes and management rewrites share one validate, hash-diff and reindex pipeline; reads never mutate truth, so usage statistics go to an index access log and weight is settled back into frontmatter in batch; raw material is append-only; the library core contains no LLM client; management never destroys information; the file boundary is the invalidation atom; and the read path is held fixed across write experiments. A memory is one file: a name, a one-sentence abstract, a type, timestamps, links, weight and provenance in the frontmatter, free Markdown in the body, and a validity interval ofvalid_fromplus optionalinvalid_atwhere a status field used to be. The store holdsMEMORY.md, the only resident injection, a configuration file that refuses unknown knobs at load, memories at<type>/<group>/<name>.md, append-only archive directories for provenance and session traces, a rebuildable index, and a state directory holding the distillation watermark and the write lock. Pull request 40 simplified the read and management halves together: the deep raw search path, its index, its configuration keys and its adapter arguments were deleted, andlimitbecame the one way to widen a recall. - 03
Recall answers with paths, and the ranking fights behind that
Recall does not paste text into a context window. It answers with an L0 list — one-line abstract, file path, anchor and score, eight entries by default — and the agent opens a hit at the depth the task needs:
mem read <name> --level outline, or abstract, or full, withmem contextdoing both in one call andmem tracereaching the cited raw messages. Each rung costs an order of magnitude more than the last, and a long file adds two free rungs: the anchor and an outline computed at read time. Three read tracks sit behind it: deterministic injection ofMEMORY.mdat session start, BM25 over an FTS5 index with a vector plugin fused in by reciprocal rank, and the plain directory tree thatlsandgrepcan still walk. One change made scope match complete path components, because a scope ofuserwas matchingusername/settings.md. Another moved the scope filter ahead of the candidate-pool limit: recall used to take the global top matches and then discard the out-of-scope ones, so a narrow scope could return nothing while a matching memory existed. Vector recall arrived as an optional plugin — install the extra, switch it on,BAAI/bge-small-en-v1.5by default, SQLite with exact cosine search — and measured 79.0% to 86.6% Recall@5 on a fixed 120-query set, against a median retrieval latency rising from 5.1 ms to 139.2 ms; it stayed off by default. - 04
What the sleep-time layer is allowed to do
Management is an independent layer on its own clock, with authority tiers. T0 is rule-only — dates, weight, links and directories — and applies directly; T1 is decided by the library executor and only creates new files or marks old ones invalid, filing a proposal instead of acting. Each kind of operation is capped per sleep and one sleep is one git commit. A dream report records what moved, what was proposed and the evidence pointers, and a decision ledger holds the verdicts; the repository argues that the ledger is truth rather than cache, because a lost rejection verdict means the operator’s refusal is forgotten and the same proposal returns on the next sleep. Pull request 36 found that ledger append was an unlocked read-modify-write where two concurrent appenders could each write back only their own line, and put it under the store lock with the new content staged into a sibling pending file. Deletion only arrives as a proposal confirmed with
mem decide, and physical removal is a human-run command management cannot reach. Reversibility, rate limits and audit are named as what keeps an unattended run safe, because memory poisoning is treated as a persistent attack surface — there is a red-team test file for poisoning payloads. The weight of that layer shows in the code:manage.pyat 31 KB andprompts.pyat 17 KB are the two largest files in a core package of 212 KB. - 05
A fresh install that recorded nothing, and 340 sessions in three minutes
The best-written entry is a pull request that fixes several defects found while preparing a demo, each reproduced by a test that failed first. On a fresh install no conversation became a memory and no new session recalled anything: the hook launched
mem distill --store …, but the store flag is a top-level argument, so the parser refused the command and the error was discarded. The default reasoning executor was a hosted model on a private project inside the organization, so outsiders got nothing; it became the firing host’s own command-line interface. That host reasoner’sclaude -pthen fired the same lifecycle hooks and distillation recursed — the pull request records roughly 340 sessions in three minutes — so executor runs now carry an environment variable the hook stands down under. A second entry is a model echoing what it had just been told: the reasoner added agroupvalue restating a schema field, the store refused it as an unknown field, the repair round repeated the mistake, and the conversation sat pending; reconciliation now drops such an echo. The last defect is that the rule-only tier’s token-set fingerprint contributed no tokens for non-ASCII text, so a sleep could invalidate unrelated Chinese memories; the replacement is an exact identity over parsed text, type, semantic fields, validity start, author, links and provenance. - 06
The month the outside fixes arrived
The record reads thirty issues and pull requests, numbered 21 to 50, seven still open. The Windows port came from outside:
core/locking.pyimportedfcntlat module level, so every entry point —mem,mem-hookandmem-mcp— died with an import error before doing anything; the fix switches on the platform and locks one byte at position zero with the Windows call, catching the plain operating-system error it raises rather than the narrower subclass. Another outside contributor sent a pair about concurrency: a correction that read a record with no lock and then wrote it under a second lock, so a concurrent writer’s edit and provenance could be lost, and a delete path with the same shape. A third filed an issue and a pull request within a minute: frontmatter did not round-trip, because double-quoted values were never un-escaped and every rewrite added a backslash, while strings such astrue,null,42or1e3were written unquoted and read back as another type. The same defect reached two people three weeks apart, and one noticed the other’s pull request had already merged the fix they had written, closed their own and credited the test covering the boundary case. The suite grew with it: 388 passing tests at 90.91% coverage on 2026-09-10, 426 in a re-validation the next week, and 544 with the linter and type checker clean across 69 source files by 2026-09-23.
Adjacent records
All records →No. 073
Engram
A learning engine that installs into a coding agent: a curriculum architect breaks a topic into a first-principles concept map, a tutor makes you predict, attempt and explain before it explains, a blind assessor grades your verbatim free recall and writes a receipt for every verdict, and a deterministic FSRS-4.5 core in one Python file decides when each concept comes back — with explorable HTML built only for the concepts whose content rewards manipulation.
No. 055
LeanCTX
A local layer that sits beside a coding agent and decides what reaches the model: file reads are compressed and cached, command output is compressed by per-command rules, session findings persist across chats, and a local proxy rewrites each request without breaking the provider’s prompt cache — with a savings ledger, a budget and a dashboard for what it measured.
No. 113
Lemmalog
A Rust Datalog engine that treats agent memory as a deductive database rather than a bigger vector store: facts asserted at the extraction boundary, stratified rules deriving closures and temporal views, provenance back to the source episode on every derived fact, and views maintained one epoch at a time — served to Claude Code and Kimi CLI as twelve MCP tools.