Skip to content

mindwalk

A local Go binary that reads Claude Code, Codex and pi session logs and replays them as light moving across a deterministic 3D map of the repository, so the files an agent searched, read and edited glow and everything else stays dark.

Screenshot of mindwalk
Editor screenshot, 30 Sep 2026mindwalk ↗

What it is

A local-first visualizer for coding-agent sessions, written in Go with a React and Three.js frontend. It reads Claude Code, Codex and pi session logs, normalizes them into one ordered stream of file-touch events, builds a deterministic 3D map of the repository that was worked in, and plays the session back over it as light: files the agent searched, read or edited glow according to how deeply and how often they were touched, and everything else stays dark. Anything the session touched that is no longer in the repository lingers as a wireframe ghost, and a histogram under the playback deck separates observation from mutation. When a session launched subagents, each child trace can be replayed over the same map as its own lens. The one feature that leaves the machine is an optional evaluation: it sends a summary of the session to the model behind the user’s own claude or codex CLI in at most two sealed calls. Reports are cached one per session, and no verdict is the model’s to decide.

Who built itRicko Yu, who publishes as cosmtrek, wrote 53 of the repository’s 62 commits under cosmtrek@gmail.com, the first on 2026-07-09 and the last on 2026-08-10. Four other accounts made up the remaining nine: yearth with six, and one each from jinwik, tcsenpai and elf-pavlik. Thirty-four of the commits carry a co-author trailer and thirty-three of those name a Claude model — twenty-nine say “Claude Fable 5” and four “Claude Opus 4.8”.

How it is put together

The parts · 6

Three artifacts kept apart on purpose, with a one-way flow between them. A trace is a session log normalized into an ordered stream of file-touch events; a citymap is a deterministic layout of the repository; a report is an evaluation of one session. A local Go server joins them and serves a React and Three.js client, and the separation is written down as a rule: adapters do not know about rendering, citymap generation does not depend on playback, the judge reads only the normalized trace and never a raw session log, and the server mostly connects data sources to the web client. Two properties follow. Determinism — the same tree always produces the same map — is what makes two sessions comparable at all, because the picture only measures attention if the ground under it does not move. And because the only thing that leaves the machine is an evidence document built from a normalized trace, the judge can be sealed: a subprocess running the user’s own CLI, with no tools, no MCP servers, no project or user settings, no session persistence, and findings as its only output.

internal/adapter/
Fourteen files and 223 KB, one package per session format: a shared adapter.go of 30.5 KB beside a Claude Code adapter of 10.8 KB with an 11.8 KB agent-correlation file, a Codex adapter of 24.5 KB with 12.6 KB of the same, and a pi adapter of 13.4 KB — plus roughly as much test code again. One open September pull request proposes splitting the shared file into topic files and replacing positional tool-call pairing with explicit by-id pairing.
internal/citymap/
Two files, 52 KB, of which a 30.2 KB builder and a 22.9 KB test. It classifies the root by shape and holds each class to a hard budget, because an unbounded walk of a home directory does not finish; it is the module the determinism claim rests on and it depends on nothing in the playback path.
internal/judge/
Ten files and 93 KB: the sealed CLI runner, the evidence and prompt builders, the rubric phase with its own evaluation document, the report model, and a cache that writes one report per session to ~/.mindwalk/reports and lets a report go stale rather than re-running it automatically.
internal/server/
Twenty-three files and 1.34 MB, by far the largest directory — a 30.9 KB server against a 51 KB test file — because it also carries internal/server/static, the embedded frontend build: a 533 KB three.js chunk, a 192 KB React chunk, 36 KB of CSS and five web fonts. The README tells contributors never to edit that directory by hand and to regenerate it instead.
web/
The React, Vite and Three.js source behind that embedded build: 26 files and 252 KB, led by a 33.7 KB application shell and a 47.4 KB stylesheet, two scene files of about 25 KB each, the playback reducer and recorder, ten files under the interface directory, and one 25.5 KB Playwright spec.
schema/, cmd/ and the repository root
schema/ holds four JSON contracts — trace, citymap, report and agent graph — that mirror the exported shapes and are meant to change with their tests in the same commit. cmd/mindwalk is the command line (serve, open, map, build, trace, analyze) and cmd/rubriceval a 9.4 KB harness for the rubric layer. The root carries AGENTS.md at 4.1 KB, a Makefile, a GoReleaser configuration, an install script that verifies a checksum, and a Claude Code skill for verification.

Choices, and what they beat

  • Take the data-layer slice of trace health and leave the panel out over shipping the standalone Trace Health view

    The contributor proposed health scoring with an indicator and a panel in #14. The author wrote that the data layer was the part mindwalk needed “regardless of how it is surfaced” — recorded, inferred and unknown outcomes — but that he was not settled on the layer above it, whose payoff comes from its consumers: the judge’s evidence document, cross-session comparison, or annotations where a metric is read. The contributor then closed #14, reopened exactly the adapter-level change as #16, and left the consumer out of scope until its product role was decided.

  • Choose a scan mode from the shape of the root over one walker for every root

    An unbounded walk of a home-directory session never returned on the author’s own machine, because consent dialogs, FIFOs and files that exist only as cloud stubs all block on open. So a git root is streamed with git ls-files, capped at 100,000 paths; a project root without git gets a bounded level-order walk; a home-like root has only selected subtrees mapped.

  • Roll verdicts up in Go from finding severities over letting the model decide them

    Stated in the README and in AGENTS.md: the judge contributes findings only, every verdict is derived mechanically, and an unverifiable criterion loses coverage rather than gaining a warning — a blind spot, not a failure. The fixed four dimensions are enforced in Go no matter what the drafted rubric contains, and the trace handed to the judge is treated as untrusted input.

  • Mirror pi’s own loader over approximating the format from its documentation

    The issue that asked for pi support pointed at pi’s documented session format and warned that its branching tree could add complexity. The adapter reproduces the loader’s recognition rules instead — the header must be the first line that parses as JSON, with a string-typed id, and blank or malformed lines may not consume that slot — then linearizes the trunk of the id-and-parent tree.

  • Reject cross-site non-GET requests and pin the host over trusting anything that arrives over loopback

    Because the analyze endpoint spends real tokens and minutes, so a page the user happens to have open must not be able to trigger it: a host allowlist on every request, origin and Sec-Fetch-Site checks on writes, and strict scheme, host and port equality — while requests carrying no origin header at all are passed through untouched.

Read fromAGENTS.md (4,134 characters), README.md (9,810 characters), docs/dynamic-rubric-evaluation.md, the bodies and comment threads of pull requests 10, 14, 16, 19, 20, 22 and 23, the release workflows and Makefile, and the complete 120-file tree with file sizes.

Build log

6 stages
  1. 01

    Twelve days to a first release, then seven weeks of quiet

    The repository was created on 2026-07-09 and its first commit, “feat: bootstrap mindwalk”, lands the same afternoon. v0.1.0 follows two days later, on 2026-07-11; then v0.2.0 on 2026-07-15 and v0.3.0 on 2026-07-18, a two-week gap, and v0.4.0 on 2026-08-03 with v0.5.0 on 2026-08-07. That is five releases in the first month and sixty-two commits in total: fifty-five in July, seven in August. The last one is dated 2026-08-10T02:06:46Z and is the merge of pull request #16, “fix(trace): preserve tool outcome certainty”. Seven weeks later, when this record is written, nothing has been added since, and the repository is not archived: 1,363 stars, 120 forks, seventeen open issues, MIT. The pause is not explained anywhere in the material — not in the README, not in AGENTS.md, not in any of the thirty issues and pull requests the report captured. What kept moving is the inbox. On 2026-08-07 the author told a maintainer of the Crush terminal agent that extracting the adapters into a reusable library was fine; two days later that library, agent-trace, appeared under MIT with mindwalk’s copyright preserved. On 2026-09-09 a contributor sent a nine-part series of pull requests adding a fourth session source, none of them merged in the capture, and the newest item in the report — an adapter for Antigravity sessions, dated 2026-09-22 — is open too.

  2. 02

    One adapter per harness, and a refusal to guess at the format

    Everything a session log becomes passes through internal/adapter, one package per agent format, each turning JSONL into the same trace model. Claude Code and Codex came first; pi was added in #20 as the third source, roughly 490 lines of implementation against 600 lines of tests. The recognition rules are the interesting part: rather than approximating pi’s schema, the adapter mirrors pi’s loader. The first line that parses as JSON must be the session header, with a string-typed id and a working directory older sessions lack; blank or malformed lines are skipped without consuming that slot, and a first entry that parses but is not a header rejects the file. Pi sessions are id-and-parent trees, so the adapter linearizes the trunk. The other recorded change here is #16, about what a normalized event cannot say: it carried only isError, so when that was false a consumer could not tell a definite success from a log that never recorded an outcome. The contributor had first proposed a whole Trace Health view in #14; the author replied that the data layer was the part mindwalk needed “regardless of how it is surfaced” and that he was not settled on what should surface it, so #14 was closed and that slice came back alone as #16. He then tested it against eleven real local sessions holding 822 events: removing the new field left every normalized trace identical.

  3. 03

    A map that always comes out the same, and the walk that never returned

    The map is the part everything else rests on, and the README states its one hard property: internal/citymap is deterministic, so the same tree always produces the same map and two sessions can be compared by looking at them. Without that, glow would measure the layout as much as the agent. Height encodes lines of code when a repository is opened with no session attached, and with one, glow follows how deeply and how often a file was touched. What walking a repository costs is the subject of #22, and it is unusually concrete. A session whose working directory was $HOME — a pi session started from the home directory — made the builder walk the whole tree, read every file end to end and never finish: consent-gated folders block open() behind a dialog that a background process never sees, FIFOs and sockets block on open, and files that exist only as cloud stubs block until they are materialized. The author wrote that he verified this on a real machine and the build never returned. The fix stops trying to have one walker: it classifies the root by shape and gives each class a budget. A git repository is read with git ls-files, capped at 100,000 paths so that a dotfiles directory cannot buffer without bound; a directory with a project marker but no git is walked level by level within a bound; anything else is treated as a workspace and only selected subtrees are mapped.

  4. 04

    Two ways of drawing a repository, and a recorder that stays in the browser

    The frontend is React, Vite and Three.js, and it ships two scenes rather than one — a radial tree and a treemap plain — each about 25 KB of TypeScript, sharing a layout module, a trail module, a texture module and directory labels. The embedded build carries a 533 KB three.js chunk and a 192 KB React chunk, which are the largest artifacts in the repository after the two screenshots. Playback is split into a reducer and a 5.7 KB recorder, and that is what lets the export menu write the playback to a .webm entirely in the browser: there is no server-side encoder and nothing is uploaded. The rest of the interface is built around not guessing. A scrubber runs over a bucketed histogram of the session in which observation stays cool and mutation glows warm, so editing phases stand out from reading. Timeline marks for context compactions, subagent launches and user turns are all click-to-jump targets. Clicking a file pins its visit history in an inspector, and clicking one of those rows moves the playhead to that moment. When a session launched subagents, an agents panel lets any child trace be replayed over the same map as its own lens — a feature with a 25.5 KB Playwright spec of its own.

  5. 05

    A judge that is not allowed to decide

    The third artifact is an evaluation, and it runs only when asked: mindwalk analyze renders a session’s normalized trace into an evidence document and hands it to the user’s own claude or codex CLI. A report has two layers — four fixed process dimensions that are the same for every session, and a task scorecard whose criteria the judge drafts from the session’s own user messages, so the yardstick is written for that task. The line the project holds is that the model supplies findings and nothing else: dimension and criterion verdicts are rolled up mechanically from finding severities, every finding must cite a timeline event the reader can click through to, and a criterion the log cannot verify loses coverage and reads “no signal” rather than counting as a failure. It is at most two calls, because the drafted rubric is cached while the task wording stays unchanged. Two pull requests from one week are the same bug twice: command summaries and judge failure detail were truncated by slicing a Go string at 500 bytes, which can split a multi-byte character into invalid UTF-8; both were changed to truncate by rune, each with a regression test whose boundary lands inside a Chinese character. A third, #10, is maintenance from the other side: Codex CLI 0.142.0 rejects a feature flag the judge still passed, so every codex-judged run failed before reaching the model.

  6. 06

    Same-origin by hand, because a judge run costs tokens and minutes

    #23 is the author reviewing his own server in four independently revertable steps, and the first step is a security argument that is only interesting because of what the endpoint does. Every request is checked against a host allowlist, to pin the name that DNS rebinding would otherwise forge, and non-GET requests are rejected when they carry a cross-site Origin or Sec-Fetch-Site header. The stated reason is that POST /analyze spends real tokens and minutes, so a page the user happens to have open in another tab must not be able to start one. Same-origin is enforced strictly: scheme plus normalized host and port must match the request exactly, so a second service on loopback does not pass, while requests with no origin header at all, such as curl, are left untouched. The same series turns boundaries that AGENTS.md states as rules into code — adapters do not know about rendering, citymap generation does not depend on playback, the judge reads the normalized trace and never a raw session log, and the server mainly joins sources to the client. The state those pieces keep, the report cache and the judge working directory, lives in ~/.mindwalk, which is what an open September issue asks to move to the XDG directories.

Adjacent records

All records →