Skip to content

CodeGraph

A code knowledge graph for agents: tree-sitter parses the source, FalkorDB stores the structure and its history, and five MCP tools answer questions about it from a Claude Code session or a browser dashboard.

Screenshot of CodeGraph
Editor screenshot, 29 Sep 2026CodeGraph ↗

What it is

CodeGraph turns a codebase into something an agent can search: tree-sitter extracts functions, classes, interfaces, imports and calls, FalkorDB stores them with validity intervals so a question can be asked as of a point in time, and five MCP tools — analyze, codebase, knowledge, query, search — expose it to an agent. It runs with an embedded database on Linux and Apple silicon or against an external FalkorDB elsewhere, parses six languages first-class with generic tree-sitter support beyond that, and ships as both an npm package and a Claude Code plugin.

Who built itCommits as Phoenixrr2113, and 516 of the 522 commits are attached to a linked account. Three addresses appear in the history — a personal Gmail, a second personal address, and one commit from a corporate domain — and the project’s npm package is published under the @agntk scope, which the repository does not explain. The volume of automated review on this repository is much larger than its audience, which makes the commit history the best evidence of how it was built.

Build log

8 stages
  1. 01

    Two thirds of the commits name their co-author

    Of 522 commits, 342 carry a co-author trailer naming a model — 66 per cent, in the same range as AI Comic Builder and below gate4agent’s 71. The mixture is dominated by one model: Claude Opus 4.7 with a million-token context on 139 commits, then Sonnet 4.6 on 100, Opus 4.6 with the same context on 59, Fable 5 on 21, plain Opus 4.6 on 18. Where AI Comic Builder recorded a project moving from one model to another and gate4agent recorded a wide spread, this one records a project that settled: four fifths of the trailers name the same three models, and the same families recur across eight months. The pull requests are signed the same way, several of them ending with the line that they were generated with Claude Code, and the repository goes further than most by shipping as a Claude Code plugin — .claude-plugin/hooks/ contains a post-tool-use hook and a pre-compact hook.

  2. 02

    What it does

    The problem is that an agent answering a question about a large codebase spends its context reading files. CodeGraph parses the code instead — functions, classes, interfaces, variables, types, imports and calls — writes the relationships into FalkorDB, and answers questions from the graph. Two design choices are worth naming. Facts carry valid_at and invalid_at timestamps, so the graph is bitemporal and a query can ask what the code looked like at a point in time rather than only what it looks like now. And the five tools are deliberately few and named for the questions they answer — analyze, codebase, knowledge, query, search — rather than exposing the graph directly. It parses TypeScript, JavaScript, Python, Go, Rust and Markdown first-class with generic tree-sitter support beyond that, and runs two ways: embedded on Linux x64 and Apple silicon, external FalkorDB anywhere else. The npm package installs two commands, an MCP server and a dashboard, and the README is explicit that they are separate entry points which can run at once against the same data path.

  3. 03

    A script that forbids the language of a claim it used to require

    Among the scripts in this repository is one that audits its own claims, and the pull request that announced the npm publication describes it turning on itself: with version 0.1.0 live, every surface had to stop being phrased as pending, and the audit script was changed to forbid the gate language it had previously required. The same pull request records that the weekly-downloads badge "honestly renders too new until npm serves counts" rather than being hidden, and that publication was verified by checking the registry tarball was byte-identical to the release artifact and that both public invocation forms boot from the registry. A later pull request corrects the documentation against the merged source and finds real drift: a README said to set an environment variable to 1 where the code checks for the literal string true; a provider that the code rejects was still recommended in the docs and was removed everywhere; embedded storage was documented as needing a running database server when the package bundles its binaries; and one package’s documented inventory of type edges did not match the eighteen-member union in the code. Documentation that is checked by a script will still drift, but it drifts visibly.

  4. 04

    The scout ran the experiments instead of reasoning about them

    One pull request is introduced by a paragraph that explains why it blocked a release: a scout — the project’s word — ran the experiments rather than reasoning about them and found three blockers. Two processes pointed at one embedded data directory spawned two database servers, and a shutdown-order test erased an indexed graph. Packing the package from its own directory produced an archive with no server in it. And an empty database was not a working first run: both binaries returned server errors and raw empty-key messages, "which is exactly what the first live dashboard load hit". The fix landed in five sequential waves, and the first of them is the interesting one — a shared embedded store where the first process takes an atomic owner lease and records its socket, later processes attach as clients over that socket, stale leases from killed owners are reclaimed safely, start races resolve to a single owner, and forged leases are rejected. The stated payoff is that Claude Code and the dashboard can hold the same database open at once.

  5. 05

    A CI failure root-caused to an ambient environment variable

    The bootstrap release run failed twice with the same message, that concurrent binaries did not share exactly one embedded server socket. The root cause is the kind of thing that would normally be fixed by retrying: the validation job exports database host and port variables for its external-storage lanes, the smoke test passed the ambient environment into both installed binaries, and driver selection prefers a remote host whenever those variables are set — so both processes were talking to the CI’s remote database and the embedded-socket count was deterministically zero. The interesting part is what the author did before believing that: to rule out an ownership bug he wrote a stress harness that cold-boots two binaries simultaneously sixty times, and ran it before and after the fix, one hundred and twenty runs in total. Every run produced exactly one lease, one socket and one server, and the two binaries won the race at different times, fifty-eight to two. That is what it looks like to distinguish a race from a bug and then prove which one you had.

  6. 06

    Incremental by contract

    A bug report describes something only a user would notice: double-clicking a node to expand its neighbours made the whole graph jump. The expansion callback had run a full-graph layout with fitting enabled and then animated a second fit to the neighbourhood, so every existing node moved and the camera relocated. The fix is written as a contract rather than a patch — the planner returns a flag meaning preserve the viewport, with no layout and no fit; the viewport is captured before the neighbour request; existing positions are untouched; and new neighbours are seeded in a deterministic ring around the expanded node. Fit and reset keep their whole-graph behaviour, and the reasoning is given: those genuinely change the picture, so they are allowed to move it. Verification is stated in the same breath: red-first planner tests — written to fail before the fix — the dashboard suite at 88 of 88, and both the Vite and Turborepo builds at 21 of 21.

  7. 07

    Numbers re-measured rather than remembered

    The landing page was rewritten product-first, and the pull request records that every metric on it was re-measured during the rewrite rather than carried over: a thousand-node project window came out at 750,767 bytes under the previous format and 305,248 under the newer compact one, 59 per cent smaller. The same change replaced synthetic illustrations with five real screenshots, sanitised. It is a small thing and it is the difference between a page that describes the software and a page that describes what the software did at some point in the past — the failure mode this archive spends its time documenting in other people’s projects.

  8. 08

    Seven stars and a lot of robots

    The audience is tiny: seven stars, one fork, no watchers. The pull-request queue, on the other hand, is busy — almost all of it Dependabot, proposing updates to Next.js, Hono, Zod, Vitest, Shiki, tree-sitter grammars and the GitHub Actions themselves, and almost none of it merged. One was closed with the note that the dependencies were updatable another way and the change was no longer needed; the rest sit open, some for a month. Every one of them collects comments from three more bots — Vercel reporting a deployment preview, Augment posting a summary, and a tool posting five separate audit reports covering security evidence, risk taxonomy, reference-set readiness, promotion readiness, config and harness checks. Those bot messages repeatedly admit that they cannot publish their results, because check permissions were never granted. Twelve issues are open. The result is a repository with more automated review than human readership, which is worth recording plainly: the process here is unusually good, and the process is not the thing that finds out whether anyone needed the software.

Adjacent records

All records →