OpenWolf
A local .wolf folder that holds one project’s task state, file index, notes and known fixes for whichever coding agent is running, injects only the changed parts of it at session start, on a prompt and around compaction, and reports the provider token counters it reads back out of Claude transcripts, Codex rollouts and OpenCode plugin records.

What it is
OpenWolf keeps a coding project’s working memory in a folder called .wolf beside the code and connects it to the agents a developer already runs. It installs lifecycle hooks for Claude Code and Codex CLI and a native plugin for OpenCode, so a session opens with a short project index and whatever saved task evidence is still relevant, an indexed file comes with a description and line ranges before it is read, and an edit refreshes the entry it touched. The folder holds the task checkpoint, the file index, project notes, candidate conventions, a bug log, handover packets between agents, archives of older notes with restore pointers, and a token ledger. The documentation states that memory, indexing and usage processing run locally without extra model calls, and that the only network traffic is an optional worker that checks npm for a compatible runtime. Usage figures are the counters the agents themselves wrote, grouped by agent and model, with missing counters left unavailable.
Who built itThe account that owns the repository holds 78 of its 123 commits, and 72 of those come from one address under two spellings of a name — Farhan Palathinkal Afsal on 55 commits and Farhan on 17. Eleven more commits arrive from cytostack@192.168.1.11, an address with no linked GitHub account at all, and a second account, doctorfarhan, holds 21. Fourteen people appear in the contributor list; apart from those two accounts and meketreve on 2 commits, the other eleven names account for one commit each.
How it is put together
The parts · 6A folder of local files plus a set of lifecycle hooks, with an optional background daemon for the dashboard and maintenance work. The organising constraint is stated in the first line of the documentation: no second model is run to maintain memory or generate summaries, so everything the tool knows was either written down by the coding agent itself or read out of that agent’s own records. That single decision explains the shape of the repository — a hook layer large enough to live inside other people’s event systems, and a file-format layer with locks, migrations, archives and restore pointers so that concurrent sessions cannot lose each other’s writes. The measurement side is a reader rather than a meter: usage is taken from provider counters in Claude transcripts, Codex rollouts and OpenCode plugin records, and where a counter is absent the documentation says it stays unavailable instead of being reconstructed.
- src/hooks/
- The integration surface, and the largest directory in the repository: 37 files and 271 KB.
shared.tsis 48,050 bytes on its own; thenpost-write.ts25,670,session-start.ts17,439,anatomy-store.ts16,796,ledger-math.ts14,695,pre-read.ts13,059,post-bash.ts11,596,handoff-state.ts8,864,anatomy-lock.ts7,715,pre-write.ts7,407,stop.ts7,275,rule-reinjection.ts5,021,session-end.ts2,473 anduser-prompt-submit.ts2,222 — one file per event the supported hosts expose. - src/templates/
- Everything an install writes out, 45 files and 232 KB:
OPENWOLF.md3,014 bytes,config.json2,501,cron-manifest.json1,604,STATUS.md1,471,cerebrum.md763,memory.md219,identity.md341 and three rules-file snippets forAGENTS.md,GEMINI.mdand OpenCode; five skills; a 31,141-byte frameworks document; and the second copy of the runtime for OpenCode, 25 files whoseshared.tsis also 48,050 bytes. - src/tracker/
- Nine files and 48 KB doing the reading and the arithmetic:
usage.ts14,966 bytes,pricing.ts11,420 (the same file is copied into the dashboard at the same size),cache-attribution.ts6,200,transcript-usage.ts6,077,waste-detector.ts3,958,usage-report.ts2,743,token-ledger.ts2,736,token-estimator.ts776 andusage-worker.ts267. - src/cli/
- 22 files and 170 KB of commands:
update.ts35,607 andinit.ts30,407 are the two biggest files in the repository after the shared hook runtime, followed bybench.ts10,828,index.ts10,335,memory-migrate.ts9,219,daemon-cmd.ts8,651,cron-cmd.ts8,367,report.ts7,071,map.ts6,221,status.ts5,565,find.ts4,599,registry.ts4,299,hook-manifest.ts4,072,handoff.ts3,048,scan.ts2,327,bug-cmd.ts1,740. - src/daemon/ and src/dashboard/
- The optional half. The daemon is six files and 43 KB —
wolf-daemon.ts19,212 bytes,cron-engine.ts14,291,context-audit.ts5,082,file-watcher.ts2,977,source-watcher.ts1,829,health.ts1,099 — and it owns the cron manifest, hook health, the watchers over source and Git, and the token auth insrc/utils/dashboard-auth.ts(2,952 bytes). The dashboard it serves is 31 files and 149 KB of React, led byTokenUsage.tsxat 25,640 bytes,ProjectOverview.tsx17,297,useWolfData.ts13,029 andCronStatus.tsx8,640. - docs/, docs/audit/ and tests/
- Eighteen documents and 378 KB in
docs/, among themclaude-codex-handoff-plan.md14,903 bytes,session-visibility-plan.md13,548 andrepair-operations.md11,608, behind a VitePress theme whose landing component is 43,135 bytes;docs/public/CNAMEcarries a custom domain and.github/workflows/docs.ymlpublishes it.docs/audit/adds six JSON files and 346 KB, of whichdiscussion-items.jsonalone is 268,270 bytes.tests/is thirty files and 199 KB, includinghandoff.test.ts12,900,anatomy-store.test.ts12,155,concurrency.test.ts11,286,usage-integrity.test.ts10,544 andp2-honest-results.test.ts9,298.
Choices, and what they beat
Maintain memory with no model calls at all over running a second model to summarise sessions and write the notes
The documentation opens by ruling it out — the tool “does not run a second model to maintain memory or generate summaries” — and the code enforces it: the cron engine throws on a retired task type with the message “ai_task is no longer supported: OpenWolf makes no model calls. Remove this task from .wolf/cron-manifest.json.” The limit that buys is stated rather than hidden: the agent “still needs to save useful semantic summaries; OpenWolf cannot infer every decision from a file edit.”
Report the agents’ own counters instead of a local meter over reconstructing usage from what the tool can observe itself
Local output-size arithmetic is kept explicitly apart from provider counters — “Local output-size calculations remain estimates, separate from provider token counters” — and the reconciled figure is labelled for what it is: “an API list-price estimate with stated assumptions, not a subscription invoice or a measurement of remaining quota.” Coverage is reported per agent so that missing counters can be seen as missing.
Treat what a past session saved as untrusted until a verifier approves it over loading earlier notes as instructions on the next run
Stated in the documentation as two sentences: “Saved evidence is marked as untrusted context. Durable instructions are loaded automatically only when the independent protected-memory verifier approves them.”
src/hooks/trusted-memory.ts(3,516 bytes) is that check, and the same file is duplicated into the OpenCode plugin copy of the runtime.Send only the changed context at each boundary over reloading the saved notes on every prompt
The handover design document states the intent — “before the next task turn, load only changed facts” — and issue #126 is a report that the implementation does not hold it:
recentis part of the change detection, so in a project where no checkpoints were ever written the hook injects an evidence block of “~400–700 tokens each” on every prompt, made up largely of echoes of the agent’s own recent tool calls. No fix appears in the material that was read.Ship a second copy of the hook runtime so OpenCode can run it as a plugin over supporting only the hosts that call external hooks
The two copies are visible in the tree — 25 files under
src/templates/opencode-plugin/,shared.tsat 48,050 bytes in both, kept in step byscripts/sync-plugin-runtime.mjs(622 bytes) — and so is the bill. Issue #129 records that when OpenCode V2 changed its plugin API the installed shape stopped loading, and issue #130 records the same five-second SessionStart timeout configured for Claude Code and Codex alike on Windows.Project notes in version control, runtime data ignored over keeping the whole folder out of the repository
The split is stated as a rule rather than left to the generated ignore file: “Project notes can be shared through version control after review. Local usage, runtime and session data follow the generated ignore rules.” The README repeats the review step for the reader — “Review notes, bug records and indexed content before sharing them” — which is the same instinct as marking saved evidence untrusted.
Read fromdocs/how-it-works.md (5,257 characters, printed in full), README.md (16,080 characters, of which the report printed the first 6,000), docs/claude-codex-handoff-plan.md as quoted in issue #126, the complete 274-file tree with sizes, the two-level directory summary, and the pull request and issue bodies quoted above.
Build log
6 stages- 01
Twelve releases, none of them in the first four months
The repository was created on 2026-03-15, and its 123 commits are piled into the back half: 21 in March, 4 in April, 5 in May, 2 in June, 40 in July, 39 in August and 12 in September, the last at 2026-09-15T16:00:52Z, a documentation commit that marks version 2.5.2 as published. The twelve releases tell the same story from the other end: the oldest is v2.0.0 on 2026-07-14T22:54:31Z, twenty-five minutes before v2.0.1, and no 1.x release exists on the repository at all, even though three separate reports are filed against version 1.0.4 by users who then upgraded to 2.5.2. Eight of the twelve landed in August 2026 — two on 2026-08-19 forty-four minutes apart, four on 2026-08-20 inside three hours, two more on 2026-08-29 — and the order of that last pair is worth keeping: v2.5.1 was published at 20:45:43Z, v2.5.0 seven minutes later at 20:52:06Z. Each release has a matching tag, and none is marked as a prerelease or a draft. Around the version line sit 2,369 stars, 215 forks, 11 watchers and 60 open issues; fourteen people appear in the contributor list, and 112 of the 123 commits have a linked GitHub account. Forty-five commits carry a co-author trailer, and 37 of those name Claude Fable 5, with Opus 5, Opus 4.7 and Opus 4.8 accounting for five more.
- 02
The memory is a folder of plain files, with a lock around every writer
What gets installed is a directory, not a service.
docs/how-it-works.mdlists what lives inside.wolf/— the project index and its readableanatomy.md,memory.md,STATUS.md,cerebrum.md,buglog.json,handoff/for checkpoints and packets,archive/for older notes with restore pointers,token-ledger.jsonandhooks/— and splits it in two: “Project notes can be shared through version control after review. Local usage, runtime and session data follow the generated ignore rules.” The line that shapes the feature set is drawn in the same document: “A checkpoint records task state separately from the full conversation”, so the file is not a transcript, and the agent is told it “still needs to save useful semantic summaries”, because “OpenWolf cannot infer every decision from a file edit”. Loading is meant to be incremental — “OpenWolf can return changed context without repeatedly loading every old note” — and what comes back is untrusted by default: “Saved evidence is marked as untrusted context. Durable instructions are loaded automatically only when the independent protected-memory verifier approves them”, which is whatsrc/hooks/trusted-memory.ts(3,516 bytes) and the plugin copy of it are for. Concurrent writers get locks and event records, and the document is careful about what that buys: “These controls reduce lost updates; they do not replace a backup.” - 03
The token ledger reads the agents’ own counters, and says so
Measurement here is a reader rather than a meter.
docs/how-it-works.mdsays the tool “reads available provider counters from Claude transcripts, Codex rollouts and OpenCode plugin records”, that it “reconciles repeated records and reports coverage by agent”, and that “Missing counters remain unavailable” — nothing is extrapolated to fill a gap. The rules are in the same place: “Input totals include cached input. Output includes reasoning where reported. Pricing separates the relevant input categories and uses each recorded provider and model.” The status of the resulting figure is spelled out rather than left to the reader: “The result is an API list-price estimate with stated assumptions, not a subscription invoice or a measurement of remaining quota.” Two other kinds of number are kept apart from it on purpose — the README describestoken-ledger.jsonas “Operational records and separate content-size estimates”, and the Bash output governor’s arithmetic is called local, because “Local output-size calculations remain estimates, separate from provider token counters”. The code matches that split:src/tracker/is nine files and 48 KB, withusage.ts14,966 bytes,pricing.ts11,420 (the same file is copied into the dashboard at the same size),cache-attribution.ts6,200,transcript-usage.ts6,077 and atoken-estimator.tsof 776 bytes. - 04
What the ledgers actually recorded, and on what basis
Every usage number here comes from a user’s own ledger, with its basis. Issue #118 measured a single reinjection of three rule reminders at 4,392 and 3,030 tokens across two sessions in
token-ledger.json, against a configured session-digest budget of 1,500 — several times the budget it sat beside, quoted from its own ledger. Issue #126 puts the evidence block injected on every prompt at “~400–700 tokens each” and multiplies that into “~30k tokens” for a 60-prompt session; the first is his measurement, the second his arithmetic, and neither comes from the project. Issue #131 counted 1,221 session entries for 110 distinct session ids, up to 65 copies of one session, in one project used for about five months, and traced it to a Stop hook that fires every turn but records each one as the session’s end. Issue #114 reports anatomy hits and misses of 5/32 (14%) on one project, where 43 of 113 misses were reads that could never have been indexed. No saving percentage appears anywhere in the material: the README excerpt stops short of any results, and the A/B harness that would produce them sits in the tree unnumbered,scripts/benchmark/run-ab.mjs(5,926 bytes) driving six task files. Pull request #115 reports 60 concurrent post-read hooks keeping 6 of 60 reads before its locking work and 60 after, measured “on real multi-process runs”. - 05
One hook runtime, plus a second copy of it for the host without hooks
The integration list in the README is itself the design: “Hooks: Claude Code, Codex CLI · Plugin: OpenCode · Compatible hooks: Grok Build · Context only: Cursor, Gemini CLI, Antigravity” — two hosts that call external hooks, one that loads a plugin, one reached through Claude-compatible discovery, three with only a rules file.
src/agents/is ten small adapters for that spread,codex.tsat 6,838 bytes down togrok.tsat 490. OpenCode cannot call the hooks, so it gets the runtime as a plugin, and the runtime is checked in twice:src/templates/opencode-plugin/holds 25 files, itsshared.tsis 48,050 bytes — the same size assrc/hooks/shared.ts— kept aligned byscripts/sync-plugin-runtime.mjs(622 bytes). Pull request #113 split OpenCode’s state file per valid session id and collected session files older than seven days; the maintainer adopted the collection directly: “Splitting state per session without it would have grown the directory without bound.” The issues record what it costs: #129 reports that OpenCode V2 replaced its plugin API while the installed plugin still exports the V1 shape, and #130 that on Windows the SessionStart hook can exceed the five-second timeout configured for Claude Code and Codex alike. #122 and #127 extend the same runtime rather than adding another, the second by reusing.wolf/hooks/and reporting 327 tests passing. - 06
Sixteen reports, one pull request, and the disagreement left on the record
The busiest stretch of the project is one contributor’s bug reports. Pull request #115, “2.5.1: multi-writer safety, project boundaries, and honest results”, closes the sixteen defects filed as #78 through #93 — “Fifteen were found and reported by @davdittrich, one by @krsfer, each with a reproduction” — and the same reporter had opened a fix for every one of them (#94, #98 through #103, #105 through #113), all still open. The maintainer says which came first: “Those PRs came first and should be weighed against this branch before merging either.” Two were adopted wholesale and named — the descriptor leak on read failure from #103 and the session-file collection from #113 — and #103 got a sentence changelogs rarely carry: “The version that shipped had the same leak until I read yours.” The verification claim is 271 tests passing, each new test first run against the unfixed code and observed to fail there. Where the two disagreed it is written down: #106 restricted daemon control to the PM2 process name, and the reply explains that 2.5.1 uses a
.wolf/daemon.pidownership record so a daemon started byopenwolf dashboardcan be stopped too. Later rounds: #118’s rule cap “was adapted in the 2.5.2 branch and is covered by tests”, #104’s work on sessions sharing one.wolfwent in except forextra_roots, and #121’s corrected dashboard warning “has not been applied yet.”
Adjacent records
All records →No. 070
delegate-skills
A skills package in which every coding-agent CLI gets its own delegation skill: the orchestrating agent writes a self-contained brief, a separate CLI edits a real working tree, and the human keeps the review and the commit.
No. 108
agy-staff
A plugin that hires Google’s Antigravity CLI as a member of staff: the host agent — Claude Code, Codex or Pi — keeps the decisions and hands the surveys, reviews and scoped edits to a Gemini 3.8 Flash worker that runs in the background and answers through a job id.
No. 109
Whiteboard
A desktop app that puts an agent and a person on the same canvas: the agent draws the review — sequence diagrams, entity relationships, code peeks pinned to specific commits — and every shape on that canvas links back to the code it was drawn from.