Skip to content

Lemmalog

A Rust Datalog engine that treats agent memory as a deductive database rather than a bigger vector store: facts asserted at the extraction boundary, stratified rules deriving closures and temporal views, provenance back to the source episode on every derived fact, and views maintained one epoch at a time — served to Claude Code and Kimi CLI as twelve MCP tools.

Screenshot of Lemmalog
Editor screenshot, 1 Oct 2026Lemmalog ↗

What it is

An embedded Datalog engine for agent memory, written in Rust. Memory is modelled as a deductive database: an extraction model turns episodes into annotated base facts, a stratified rule layer derives temporal projections, transitive closures and contradiction candidates, and any derived fact can be asked why — the answer is a proof tree ending at the source episodes. Facts are bi-temporal, so supersession closes a validity interval rather than deleting; annotations are a semiring, confidence multiplied along a proof chain and provenance unioned as a set of episode ids. Rules are parsed at runtime, not compiled: an agent installs them as versioned batches it can uninstall, and installing one marks the program dirty so the next run backfills every rule against the store. Evaluation is seminaive per epoch, retraction is a scoped negative delta over transitive dependents only, and point queries can use magic sets instead of a full fixpoint. It ships as the crate, a REPL, a headless CLI, an agent skill and a twelve-tool MCP server, with 44 tests and benchmark scripts in the tree.

Who built itHis account was opened in 2016, holds 76 public repositories, four gists and 88 followers, and links to pwning.systems; his bio reads “Popping the stack all day, everyday.” He wrote 33 of this repository’s 46 commits — 28 of them from a local address GitHub ties to no account, 5 from the linked one. Three other people landed the remaining 13 commits between them, and five commits carry a co-author trailer crediting “Claude Opus 5 (1M context)”.

How it is put together

The parts · 6

A Rust interpreter rather than a compiled rule program, with the fixpoint kept pure and everything non-monotone pushed to the edge. Rules are parsed at runtime, which is what lets an agent install, version, revert and backfill them mid-session; the store is row vectors with per-position hash indexes backing every read path, and rule bodies backtrack through an undo trail instead of cloned environments. The evaluator is seminaive per epoch with scoped negative deltas for retraction, magic sets for demand queries, and a change log that both the context assembler and external projections read. Above it sit the extraction boundary — a trait, memoised by episode id, degrading to zero facts on provider error — a deterministic update policy that resolves what it can and escalates the rest, a positional context assembler, an MCP server, a REPL and a skill. Two consequences shape the code: installing a rule marks the program dirty, so the next run clears derived relations and backfills every rule against the existing store; and derived relations are never persisted, so loading a snapshot replays base facts and rebuilds the views.

src/eval.rs
The engine, 86 KB and much the largest file in the repository: the store of row vectors with per-position indexes, annotation merge, stratification, trail-backtracking seminaive evaluation, scoped negative deltas, the epoch change log, pattern queries, ask and ask_deep, and proof trees.
src/agent.rs
The agent layer, 50 KB: the extraction boundary with its mock and model-backed extractors, the ADD / UPDATE / NOOP / escalate policy, the escalation queue, the positional context assembler and the AgentMemory facade that walks observe, policy, maintain, ask, context and why.
src/bin/
Four binaries in 88 KB: a 50 KB benchmark runner, the 31 KB MCP server speaking stdio JSON-RPC with twelve tools, a 7 KB headless CLI that writes to the same snapshot so sub-agents and cron jobs can reach the engine, and an 815-byte REPL entry point.
The smaller src modules
Interner and values, the rule AST with a hand-written parser, magic-sets rewriting, the semantic index behind an Embedder trait, entity canonicalisation with read-side canonical views, hybrid retrieval, a deterministic long-horizon scenario generator with ground truth, the REPL command surface, the extraction prompts, and the benchmark adapter.
tests/
Eleven files and 98 KB of tests, led by the 25 KB agent tests, the 17 KB engine tests and the 15 KB differential tests that generate random programs and compare them against a brute-force oracle, with annotation, canonical, retrieval, aggregate, eval, model, semantics and session tests behind them.
README.md, the design document and skills/
A 36.7 KB README with a feature-status table and two benchmark reports, a 22 KB design document with an implementation plan, its risks and a status section, a 42 KB single-file page under docs/, an 11 KB agent skill, nine example programs, and two Python benchmark scripts, one of which traces every loss to a bucket.

Choices, and what they beat

  • Datalog, with the model kept outside the fixpoint over an LLM reasoning over retrieved context, or an LLM predicate inside rule evaluation

    The design document states the boundary as something discovered rather than assumed: no established system puts an LLM call inside a Datalog fixpoint, and for good reason, because LLM calls are non-monotone and expensive. Extraction therefore happens at the boundary and is memoised by episode id, the fixpoint stays pure, and the deterministic update policy resolves what rules can resolve and escalates the rest to the agent as a work item.

  • An interpreter with runtime-loadable rules over the compile-time rule binding used by Crepe and Ascent

    Compile-time binding rules out the thing the project needs most — rules an agent installs, versions, reverts and backfills in the middle of a session. The cost is paid inside the evaluator, which now carries per-position indexes, row-id lookups and an undo trail so that rule-body joins still select the smallest bucket.

  • One concrete annotation product first over general annotation polymorphism

    The risk section defers the general case in writing: fix one concrete annotation product first and generalize after the eval harness exists. The polymorphism arrived as a pull request on 2026-09-14, once that harness existed, and the author’s reply records the ordering — the design document’s phase 3 arriving exactly when the risk note said it should.

  • Indexes and a nested-loop evaluator now, triejoins later over writing worst-case-optimal joins first

    The cyclic-join benchmark decided it: with 2,000 nodes and 8,000 arcs, triangle detection was cheap on the sparse buckets, while materialising a full transitive closure dominated everything at 3.9M facts and about 67 seconds. The README reads that as an argument against blindly materialising dense closures and for demand queries, not as an argument for triejoins at agent-memory scale.

  • Scoped retraction instead of rebuilding the derived views over recomputing every derived relation when a fact is superseded

    Only predicates that transitively read the retracted one are cleared, level by level in stratum order, and a level is only propagated into when its input’s key set actually changed. That is what keeps an idle turn in the microseconds and an incremental turn near 50 ms on the 124,750-fact chain, and it is the mechanism the differential harness later forced onto predicates read negatively.

  • Rank the selection rather than keep extracting over pouring a richer memory into the same context budget

    After per-item enumeration extraction grew the store about 2x and dropped F1 to 0.435, the conclusion written into the README is that selection, not extraction, is the binding constraint at the context boundary — and the fix was to scale the selection budget from 1,800 to 3,200 tokens rather than to extract less.

Read fromREADME.md (36,718 characters), datalog-context-engine-design.md (22,066 characters), skills/lemmalog/SKILL.md (10,975 characters), the complete 49-file tree with sizes, and the issue and pull request threads that name the files and commits they changed.

Build log

5 stages
  1. 01

    Three weeks, forty-six commits, and no releases at all

    The repository was created on 2026-08-27 and its 46 commits split almost evenly across two months: 24 in August, 22 in September, the first of them stamped 06:53:34Z — thirteen minutes before the repository’s own creation time. There are no releases and no tags, so the version line never started. Around that sit 326 stars, 29 forks and 3 watchers, with zero open issues at the end. Four contributors appear: 33 of the 46 commits are authored under the name Jordy Zomer, 28 of those from an address GitHub links to no account, and the other three — Bergmann89, Frank Noirot and Josh Biggley — landed 13 commits between them, each arriving attached to a detailed report. Five commits carry a co-author trailer, all of them crediting “Claude Opus 5 (1M context)”. The ending has a shape of its own: on 2026-09-15, between 08:34:57 and 08:35:03, four issues were closed inside six minutes, after a reply to a 2026-09-03 pull request that reads “Sorry for the late response, I was on holiday.” The newest commit is titled “Close the triage: _Y variables, shadow-install warnings, installer”. The README describes itself as carrying an honest status log, and its feature table still lists leapfrog triejoins and DBSP-style streaming deltas as future phases.

  2. 02

    Why Datalog, and where the author’s own numbers argue against it

    The design document opens by rejecting the framing it argues against: context is treated as a buffer, not a database, and its predecessors — Zep/Graphiti, Mem0, GraphRAG, Letta — are faulted not for storing the wrong thing but for deriving nothing: closure and contradiction detection are redone per query, or absent. It cites the context-rot literature for the claim that reliable context is 4k–32k tokens whatever the advertised window. Datalog is chosen for four reasons: recursion is the shape of context queries; guaranteed termination makes it safe to hand an LLM; incremental evaluation is solved; and semiring annotations carry provenance, confidence and recency together. Two limits are written down. No established system puts an LLM call inside a Datalog fixpoint, because LLM calls are non-monotone and expensive, so the model sits outside it. A positioning table on temporal facts, derived rules, incrementality, provenance and a safe query language gives Lemmalog all five. The measurements cut the other way, and the README says so: selection, not extraction, is the bottleneck at the context boundary. What shipped instead is a three-signal ranker — in-crate BM25 over facts and verbatim episodes, entity-match boosting at +1.5 for the named entity and +0.4 for one-hop neighbours, and assembly giving 60% of the tokens to ranked facts.

  3. 03

    What a fact carries, and how far the provenance reaches

    Every base tuple holds the triple plus valid_from, valid_to, asserted_at, a confidence and a provenance pointer. Bi-temporality means an update is retraction by annotation — valid_to is set and nothing is deleted — so history and as-of queries survive. The annotations are a semiring rather than columns: confidence in the unit interval fused by a product t-norm, provenance as a set of episode ids under set union, salience as an interval lattice. Re-derivation merges rather than replaces, taking the maximum confidence, the union of provenance and a deduplicated support list. why() returns a proof tree with cycle protection, and on an aggregated fact it shows the rule plus a contributing row witness. The granularity the design settles on is the episode, and provenance witnesses per fact are capped on purpose, because why() needs one derivation rather than the exponentially many paths. Two conventions make that granularity usable by a model: the agent skill anchors evidence as located(Entity, "file:line") so a reference survives derivation, and it spells out the arithmetic of a product — four hops of verified facts land at 0.66 — so read-and-verified facts are tagged 1.0, inferences 0.4 to 0.7, and the default is 0.9. Snapshots keep episodes verbatim and never persist derived relations: those are rebuildable projections, recomputed on load.

  4. 04

    Epochs, scoped retraction, and the join that was not written

    Each run() is an epoch. Facts asserted since the previous one form the delta seeds; seminaive evaluation fires each rule once per positive body atom bound to the delta and iterates to fixpoint, so a run with no new assertions derives nothing. Retraction is the asymmetric half, and the author chose the scoped version: only derived predicates that transitively read the retracted one are cleared and rebuilt, level by level in stratum order, and a level is only propagated into if its input’s key set actually changed. Every addition, retraction and clear is stamped with an epoch in a change log, which backs the context assembler’s “new in memory since last turn” section and a change feed for external projections; what_if assumes temporary facts, answers a goal, and restores the store byte-identically. On an M-series laptop a 500-node chain closure of 124,750 facts fixpoints in about 17 seconds while an incremental turn costs about 50 ms and an idle turn microseconds. The same benchmarking is used to argue against an element the document had already chosen: with 2,000 nodes and 8,000 arcs, triangle detection was cheap on the sparse buckets while materialising a full closure dominated everything at 3.9M facts and about 67 seconds — an argument for demand queries, not for triejoins at agent-memory scale.

  5. 05

    Silent successes, the union trap, and what the harness caught

    Several of the eight reports share a failure shape: the engine reports success and produces nothing. Issue 7 is the sharpest: the author’s answer contradicts the report. Narrowing a rule appeared to leave its old tuples queryable, but uninstall plus install withdraws them correctly: the trap is Datalog union semantics, because installing the corrected rule as a new batch leaves the earlier batch active and both definitions union. Nothing warned the caller; the tool and CLI now do. Issue 6, filed the same day, is the same shape: _Y, the Prolog don’t-care convention, was parsed as a constant, so the rule derived nothing while install reported success; it now parses as a variable. The largest report is Josh Biggley’s: one clock was both the valid-from stamp for facts being written and the present now(T) reads, so replaying history out of order pushed the read clock backwards and 79 of 823 edges vanished from current. A differential harness — 450 random programs against a brute-force oracle and a 2,000-case parser fuzz — found a soundness bug: growth of a negatively read predicate never invalidated earlier derivations. Not everything worked: counting rules did not fix the counting questions, whose bottleneck is extraction recall rather than aggregation, and a prompt that restored knowledge-update to 0.587 was reverted after crashing temporal to 0.20.

Adjacent records

All records →