Skip to content

Memmy

A local memory hub: one service on the machine holds the traces, policies, world model and skills that Claude Code, Codex, Cursor, DeepSeek Harness, OpenClaw, Hermes and OpenCode all read and write, so context carries between agents instead of staying inside whichever one was open.

Screenshot of Memmy
Editor screenshot, 1 Oct 2026Memmy ↗

What it is

A local memory hub that several coding agents read from and write to. One service — http://127.0.0.1:18960, its data in ~/.memmy/memory-service/memory.sqlite — holds the memory, and hooks, plugins, the memmy-memory CLI and the desktop app all go through it, so Claude Code, Codex, Cursor, DeepSeek Harness, OpenClaw, Hermes, OpenCode, WorkBuddy and Pi can share one store. Each host is connected one of three ways: a hook and workspace bridge written into its own directory, a plugin in its own format, or a read adapter that scans the session files the host already keeps. Memory is layered — L1 traces, L2 policies, an L3 world model and skills — and promoted between layers by measured gain, with reward propagated back through an episode after a 30-second feedback window. Built by the MemTensor organisation in 76 days: 1,279 commits, 20 releases, 2,439 files, 2,000 stars.

Who built itAn organisation account rather than an individual one: MemTensor ships Memmy as a product, with staff addresses on memtensor.cn and memtensor.com, a site at memmy.bot, a Discord server and an X account. Its 1,279 commits come from 25 contributors, and one of the contributing accounts is named Memtensor-AI. The largest single committer is Wang-Daoji with 223 commits, followed by hijzy at 163, syzsunshine219 at 130 and ZongYue99 at 118. Of the 1,279 commits, 921 are linked to an account and 375 are not. The organisation is not the main writer: 153 commits carry a co-author trailer, 54 of those saying Cursor and 64 naming a Claude model, with another 35 naming accounts that also appear in the contributor list.

How it is put together

The parts · 6

A local-first memory service with a per-host integration layer around it and a desktop application on top. The centre is one HTTP service on loopback — Memory/src/server/index.ts, default port 18960, SQLite and sqlite-vec for storage and vectors — that owns a four-layer memory model and the background jobs promoting memory between those layers, so nothing about recall depends on which agent is in front of it. Around the service sit three kinds of adapter: install targets that write a hook and a workspace bridge into each host’s own directory, plugins in each host’s native format, and read-only scanners that import history a host has already written to disk. The desktop app is an Electron shell with a React frontend and a Fastify local backend keeping application state in SQLite, and the agent runtime is a separate workspace with its own model providers, CLI and TUI. The consequence of putting the service in the middle is that the hard problems end up in the adapter layer rather than in the memory itself: every host names its files, formats its sessions and gates its hooks differently, which is why that one directory holds ten targets with large per-host files, and why the reading code shared by the desktop backend and the memory service lives in its own workspace instead of being copied into both.

Memory/
The memory service: 228 source files and 3037 KB, with 122 test files at 1765 KB beside them. Inside src/service/ the biggest files are session-turn-service.ts at 151 KB, feedback-experience.ts at 116 KB, memory-service.ts at 115 KB, retrieval-service.ts at 103 KB, skill-pipeline.ts at 68 KB and span-pipeline.ts at 64 KB. The largest single file in the repository sits in the same tree: src/algorithm/plugin-algorithms.ts at 263 KB, next to an agent-source runtime of 58 KB. Alongside it sit Memory/adapters/ (a Cordis patch for DeepSeek Harness, an OpenClaw plugin manifest, a Python provider for Hermes), Memory/agent-contract/ with a 26 KB DTO module and the event and episode-status types, installers for shell and PowerShell, and a 71-file viewer.
App/
Four workspaces, 1,695 files. frontend is the Electron UI, 480 files at 15.7 MB, whose Chinese and English message table alone is 217 KB and whose agent API client is 105 KB; memmy-agent is the runtime, CLI and TUI, 733 files at 8.2 MB, including the model providers, an OpenAI-compatible server and the terminal interface; shell is the desktop main process, 97 files at 5.8 MB, where main.ts is 204 KB and runtime-services.ts 93 KB, with a 127 KB packaged-runtime boundary test; backend is the local API, 385 files at 2.8 MB, holding the Fastify routes, SQLite application state, source scanning and skill writing.
AgentSourceCore/
Twenty source files shared by the desktop backend and the memory service rather than duplicated: one <host>-source-turn.ts per agent for Claude Code, Codex, Cursor, DeepSeek, Hermes, OpenClaw and OpenCode, plus JSONL line handling, a secret redactor and a memory token budget. Its tests include a 22 KB review test over the source-turn contract, which is how seven different session formats are held to one shape.
docs/
The documentation, in parallel trees: 39 Chinese pages at 117 KB and 38 English pages at 76 KB, twelve architecture documents, and 38 assets at 60.6 MB — the largest directory in the repository by size. The memory overview runs to 17,032 characters in English and 15,391 in Chinese, and the page on memory sources to 20,262 characters in English.
Knowledge/, Migrations/ and scripts/
Three supporting workspaces: a knowledge store whose single UI page is 112 KB, eighteen migration sources with fourteen test files beside them at 152 KB and 136 KB, and thirteen top-level scripts plus 31 under scripts/internal covering development startup, packaging and release plumbing.
.github/
Six workflows and a CODEOWNERS file: a 50.5 KB draft-release workflow, a Linux CLI installer, Windows memory validation, a memory release job, and the Yunxiao-to-GitHub sync, next to seventeen release-note files ranging from 455 bytes to 4.7 KB.

Choices, and what they beat

  • One local service as the memory store over a separate memory store inside each agent

    The documentation is explicit that hooks, plugins, the memmy-memory CLI and the desktop app all read and write the same service at 127.0.0.1:18960, and that this is what lets different agents share one memory. It also makes the service separable from every integration: install --service-only registers the service and verifies its health without installing a skill or adapter into any agent.

  • Install the memory service without configuring the agents it finds over wiring every detected agent during installation

    The README states that the installer initializes memory without changing Codex, Claude Code, Cursor or other agents, and that attaching a skill and the supported hook or plugin is a separate, explicit command, either for all detected agents or for one named agent. The memory core is therefore usable on its own, and the invasive half is opt-in.

  • Layered memory promoted by measured gain, with a negative archive threshold over one undifferentiated store of past turns

    L1 traces become L2 policies only above a gain of 0.02 and are archived below -0.05, clusters of L2 become an L3 world model that must reach a confidence of 0.2 before it can be recalled, and skills crystallize from policies at an eta of 0.1. Promotion is a threshold rather than an editorial act, and demotion has a number too, which is what keeps four layers from filling with everything.

  • Recalled memory is rendered as historical context, with the current request authoritative over treating recalled text as part of the instruction

    The recalled items are rendered in a fixed order and the documentation says the current user request always remains authoritative, with the injected Markdown stating that the material is historical memory to be checked against the present request and repository state. The same instinct excludes the current session’s own L1 traces from recall, so a turn cannot echo back through the long-term path.

  • A model filter with a deterministic fallback after mechanical ranking over letting the filter decide alone

    The final semantic filter is on by default and is allowed to drop every candidate, but an invalid response or a failed call keeps the top six mechanically ranked results, and the diagnostics expose whether the filter ran, was skipped or fell back. Filtering improves the result without becoming a single point of failure for recall.

Read fromdocs/en/memory/overview.mdx (17,032 characters) and docs/en/concepts/architecture.mdx, the Chinese copies of the same two pages, README.md (9,833 characters), the Memory/src/agent-source/integration/ and Memory/adapters/ trees, Memory/readme.md, and the complete 2,439-file listing with sizes and the two-level directory summary.

Build log

6 stages
  1. 01

    One local store, ten hosts, twenty releases in eleven weeks

    The repository was created on 2026-07-16 and its first commit is dated 2026-07-17: memmy-agent: Let every AI remember the same you. Version v1.0.1 shipped the same day, and by 2026-09-30 the line had reached v1.2.0 — twenty releases in eleven weeks, the longest quiet stretch being the nine days between v1.1.1 on 2026-08-26 and v1.1.2 on 2026-09-04. The commit curve is 240 in what was left of July, 450 in August and 589 in September, 1,279 in all across 76 days. Around the code sit 2,000 stars, 194 forks, 5 watchers, 37 open issues and 25 contributors, and the README carries a Product Hunt daily top-post badge. The claim the project makes for itself is in the repository description — in its words, “all AI remember the same you” — and the mechanism is a single service rather than one store per tool: the endpoint is http://127.0.0.1:18960, the data is ~/.memmy/memory-service/memory.sqlite, and the documentation states that hooks, plugins, the memmy-memory CLI and the desktop app all read and write that one service, which is what lets different agents share the same memory.

  2. 02

    Three ways into a host, because no two hosts agree on anything

    A memory hub is only useful if it can get in and out of other tools, and the repository shows three distinct routes. The first is an installed target per agent: Memory/src/agent-source/integration/ holds one directory per host — claude-code at 13.6 KB, codex at 12.4 KB plus a 9.6 KB hook-trust.ts, cursor at 9.0 KB, deepseek-harness at 7.9 KB, opencode at 7.5 KB, and much larger ones for hermes at 78.8 KB and openclaw at 64.3 KB, with pi and qwenwork under 600 bytes each — alongside templates that emit what gets written: a 43.8 KB resume hook, a 39.0 KB OpenCode plugin and a 25.2 KB DeepSeek Harness plugin. The second route is a plugin in the host’s own format: Memory/adapters/openclaw/openclaw.plugin.json, a Cordis patch file for DeepSeek Harness, and a Python memmy_provider package for Hermes. The third route only reads: one adapter per host that scans the session files a tool already keeps. How little uniformity there is between them shows in those adapters — Cursor’s workspace state is a SQLite file, Codex keeps rollouts, OpenClaw and OpenCode have databases of their own, and DeepSeek Harness writes compressed session.jsonl.zstd files.

  3. 03

    What controlled and local-first mean in the code

    The service can be installed without touching any agent: memmy-memory install --service-only downloads the pinned runtime, registers a user-level service, starts it and checks /api/v1/health, and the README says it does all of that without installing a skill or adapter anywhere. The Linux installer initializes memory and leaves Codex, Claude Code, Cursor and the rest alone until the user runs memmy-memory init, or memmy-memory init --agent <agent>. Two settings turn memory reading and writing off independently — enableMemorySearch and enableMemoryAdd — and the documentation notes that neither forces the other. Configuration is a single file at ~/.memmy/config.yaml, and recall, evolution and model settings can be reloaded with memmy-memory reload-config; changing storage instead returns requiresRestart: true rather than pretending to have applied. The surface for checking what the machine actually knows is small and inspectable: memmy-memory search "..." --verbose prints the raw candidates gathered by each channel, the hits left after fusion and thresholds, the memories really rendered into the agent’s context, and whether the model filter ran, was skipped or fell back.

  4. 04

    Four layers, six channels, and a two-hour definition of a conversation

    The memory is layered, and every promotion between layers has a number attached. turn.complete writes a raw turn and one L1 trace per captured step, keeping up to 4,000 characters of text and 2,000 per tool output; consecutive follow-ups merge into one episode with a maximum gap of two hours; when an episode closes the service scores a reflection, waits a 30-second feedback window, computes a task reward and propagates it back through the episode’s traces with a 0.9 gamma and a 30-day decay half-life. L1 traces above a value of 0.005 and a similarity of 0.65 become L2 policies when their gain clears 0.02, and are archived below -0.05; clusters of L2 at 0.3 similarity become an L3 world model that needs a confidence of 0.2 to be recalled; skills crystallize from policies at an eta of 0.1. Recall uses four families — vectors, SQLite FTS5, a short-pattern channel built for two-character Chinese fragments and short ASCII terms, and structural fragments for error signatures and paths — over three tiers sized 3, 5 and 2, which is also the default limit of ten results, then MMR at 0.7 relevance against 0.3 redundancy, then an optional model filter that keeps at most eight and falls back to six when it fails. Vector search builds its window from the most recent 2,000 rows, and the documentation says that number is a fixed constant rather than an option.

  5. 05

    Release branches with the boundaries written into the pull request

    The release process is unusually explicit for a project this young. Work lands on a version branch, and the pull request that cuts a release states what it will not do: a v1.1.9 candidate in September carries a section headed “Release boundaries” — draft pull request only, no version tag pushed, no GitHub release published, no installer uploaded. Two candidates were abandoned and superseded within the day, one of them for targeting the version branch instead of main. The merge pull request for v1.2.0 carries three review handoff markers, each an HTML comment holding a hash and a timestamp, with the note that only reviews submitted strictly after that timestamp qualify. The task tracker is not GitHub: work arrives as identifiers such as MEMMY-491, MEMMY-487, MEMMY-607 and MEMMY-623, which the pull requests describe as tasks in Yunxiao and which a workflow named for Yunxiao mirrors onto GitHub, so requirements are filed in one system and code lands in another. Release notes live in the repository and shorten as the pace rises — 4,731 characters for v1.1.2 in early September gave way to 455 for v1.1.8 and 695 for v1.2.0 — and one candidate PR includes a fix to two stale regression assertions that still pinned the backend to 1.1.7 while the root version was already 1.1.8.

  6. 06

    A numbered bug list, synthetic samples, and what the fixes cost

    Most contributors come from outside the organisation and work from a numbered list of known problems; the pull requests cite entries such as item 2, item 5, item 39 and item 60, and most of those fixes are still open rather than merged. An analytics call was sending the search query, the memory body and the memory ID to a cloud endpoint as the “action” name — the skill this project installs into agents teaches those exact forms — and the fix reports a known subcommand or nothing. Closing a session woke the background worker only when the response happened to carry a closed episode, so newly frozen L3 jobs sat; the pull request measures the wait at 2 minutes 41 seconds and replaces it with one unconditional call. read_file counted a trailing newline as an extra line while apply_patch did not, so a model patching a one-line file wrote a diff that could not apply. Turn capture excluded every turn the relation model labelled as ending a topic, discarding a user’s feedback when it shared a sentence with a goodbye; only a bare end command is excluded now. The samples in these pull requests are synthetic and the authors say so, and where testing stopped short they say that too: the change adding subscription plans as providers has not been called with a real key. An offer of a hosted namespace behind the local hub, filed as issue 547, has no reply.

Adjacent records

All records →