Skip to content

Graft

It reads a repository with tree-sitter, writes what it finds as a folder of linked markdown nodes plus a per-symbol code graph, and hands the matching slices to whichever coding agent you already use — up front in the prompt, or through six MCP tools the agent calls itself.

Screenshot of Graft
Editor screenshot, 30 Sep 2026Graft ↗

What it is

Graft is a command-line context layer for coding agents, written in TypeScript under the MIT licence. It reads a repository twice: a deterministic tree-sitter pass that produces a per-symbol graph with cross-file call and import edges, and an optional model pass that summarizes every file and groups those summaries into a few dozen readable markdown nodes joined by typed links. The result is a gitignored folder, and every query re-checks the working tree against the last build’s fingerprint before answering — about three milliseconds when nothing moved — so answers describe uncommitted edits as readily as committed ones. Agents reach it three ways: a marker-fenced block or an owned skill file written by an init command that knows nine agent hosts, six MCP tools the agent calls itself, and Claude Code hooks that pull matching nodes into each prompt and print an edited file’s blast radius. Deep summaries use whichever provider key the user supplies; the structural graph never calls a model at all.

Who built itThe repository belongs to the trailhq organisation, so the name here is the heaviest contributor rather than an owner. Of its 512 commits Shrish Dwivedi wrote 235 and Anirudh Kumar 160, against 25 for Dependabot, 25 for Alex Matthews and 20 for Frankie-Xu; thirty-six accounts appear in the contributor list, and the 111 co-author trailers add about ten more names, most of them once. The project grew out of a NanoNets repository — the npm package is still published under the NanoNets scope, and its own development instructions still clone NanoNets/context-graph-engine.

How it is put together

The parts · 6

The organising idea is that a map of a codebase should be a folder of files an agent reads the way it reads the repository, not an index it queries. One build produces two graphs: a deterministic per-symbol code graph from tree-sitter, which needs no key and no network, and an LLM-written markdown node graph that a deeper, explicitly opt-in pass adds. Retrieval then has three front doors onto the same artifacts — an instruction or skill file that names the graph and tells the agent how to re-ask, six MCP tools the agent calls itself, and Claude Code hooks that inject matching nodes before a prompt and print an edit’s blast radius afterwards. Freshness is handled at query time instead of by a watcher, because the parse is cheap and content-hash cached, so every command can afford to compare the working tree with the last build’s fingerprint first. Nothing here is per-host: the wiring layer is a registry of readers and writers across nine agent hosts, which is why the same instruction block has to exist both as a fenced section inside files the user owns and as wholly-owned rule or skill files for hosts that expect one.

src/graph/
The engine, and the largest directory in the repository: 45 files and 455 KB, led by a 118 KB extractor, 34 KB of binding code, 33.6 KB of symbol resolution and 29.7 KB of workspace handling, with sixteen tree-sitter query files under queries/ and separate modules for the extraction cache, fingerprints, invariants, the optional language-server enrichment and the shells that answer callers, skeleton and blast-radius questions.
src/cli.ts and src/ai/
The whole command surface sits in one 88 KB file, the largest file here; beside it are provider adapters for Anthropic, OpenAI-compatible endpoints, LiteLLM and OrcaRouter, a module that recovers structured tool calls, and the context assembly, pricing and savings code under src/context/.
test/
138 files and 1,169 KB, larger than any source directory except src/graph, and unusual in composition: test/ask.test.ts is 52 KB and the Java graph test 53 KB, while three files are probes rather than assertions — a review-process probe, a reference-scale probe and a Node 24 probe.
src/brain/ and src/app/
Twenty-eight files and about 274 KB for the hosted half attached to Trail: sign-up, push, pull, upload, watch, chunked transfer, the GitHub app, the server and the review workers, with a 10 KB github-app document and an App Runner deploy script.
src/telemetry/ and TELEMETRY.md
Eleven files around a 13.8 KB contract module, an on-disk queue with a daily flush, a gate, identity and notice modules, and a 10.8 KB contract test — the anonymous-stats promise as a tested, documented artifact rather than a footnote, with a build-time script that stamps the key.
viewer/, .github/ and assets/
An eight-file prebuilt viewer served by graft viz, six workflows including a 15 KB publisher that puts per-pull-request blast-radius pages on GitHub Pages and a 9.3 KB composite action, and 40 MB of demo material, the largest single piece a 17.6 MB recording of the post-edit hook.

Choices, and what they beat

  • A folder of markdown files as the graph over embeddings and a similarity index

    The README’s argument is that the agent should open, grep and follow the map the way it reads any other file in the repository — no embeddings, no similarity search, no index to keep warm. Each node carries a plain-English summary, a crux excerpt lifted from the source, hashed sources and typed wikilinks, on the claim that an address-only map still sends the agent to read the file itself.

  • Deterministic tree-sitter by default, the model behind an explicit flag over making the graph an LLM artifact

    A plain build, a check and an ask need no key and no network, and the automatic sync driven by the Claude Code hooks is documented as never calling the model on its own. The deep pass is opt-in, and the provider is chosen by environment variable or flag, with the OpenAI wire format pointable at OpenRouter, Fireworks, Groq, a LiteLLM proxy or a local server.

  • Refresh before every query over a warm index kept by a watcher

    A stat against the last build’s fingerprint costs about three milliseconds when nothing moved, so answers can describe the working tree with uncommitted edits included and there is no stale index to babysit. The exception is deliberate: graft check never refreshes, because it is the drift report and exits 1 when the graph has fallen behind.

  • The crux stored as source text over a line range

    Stated in the README in one sentence: line numbers drift whenever unrelated code above them shifts, while the lines that matter do not, so keeping the text instead of the numbers keeps the excerpt correct as the file around it moves.

  • graft/ as a gitignored local cache over committing the graph

    It is described as regenerable, like node_modules, so what a team shares is the wiring the init command drops in and each teammate builds their own graph. The cost of that choice is visible in the issue list, where users object to a build that silently gitignores the very wiring files the same policy expects them to commit.

  • Generated blocks fenced, and the user’s CLAUDE.md never touched over writing the instructions as a file the project owns outright

    A selected agent gets a marker-fenced section inside the shared instruction file or a wholly-owned rule or skill file; Claude Code gets its own skill file and the user’s CLAUDE.md is left alone, and with no terminal to prompt on, the init command writes nothing at all and prints the command to run instead.

Read fromThe README, fetched in full from raw.githubusercontent.com (44,245 bytes): its sections on how the graph gets built, what is in a node, what runs where, agent integration, the CLI reference, search and orient, monorepos and multi-repo folders, and the benchmark appendix. Plus the recon report’s 357-file tree with sizes and its two-level directory summary, and the pull request bodies that state the reasoning behind the concurrency flag and the model-adapter fixes.

Build log

6 stages
  1. 01

    Thirteen weeks, two people, fourteen tags

    The repository was created on 2026-07-03 and its first commit is Initial commit: Context Graph Engine; by 2026-09-30 it held 512 commits — 243 in July, 213 in August, 56 in September, a first rush that slows once the project must be maintained rather than built. Two people wrote most of it: Shrish Dwivedi at 235 commits and Anirudh Kumar at 160, against 25 for Dependabot, 25 for Alex Matthews and 20 for Frankie-Xu. Thirty-six accounts appear in the contributor list, and the 111 co-author trailers add about ten more names, most of them once: 47 name a Claude model, 20 say Cursor, 25 are the dependency bot. Versions are tags rather than releases — GitHub records no releases at all, while the tags run from v0.7.1 through v0.8.1, v0.8.2, v0.9.0 and v0.12.1 to v0.21.1, with v0.10, v0.11 and v0.17 absent from the record, and a 39,765-byte CHANGELOG.md carrying what release notes would. Upgrades are quiet by design: the CLI checks npm once a day, announces a new version, and the next session refreshes the repository’s own wiring once it is installed. The project also runs on itself, visibly — a bot comments a blast-radius diagram on each pull request and links to a hosted page for it, which is what the blast workflow, its cache workflow, a 15 KB Pages publisher and a 9 KB composite action exist for. Around all of it: 9,431 stars, 865 forks, 188 open issues.

  2. 02

    Two tiers, two caches, one three-millisecond stat

    Understanding is built in two tiers with unrelated costs. The first is deterministic tree-sitter: every function, class and call edge, no model and no key, written as graft/.graph/wiring.json and per-file cards. The optional deep pass adds a per-symbol summary and crux, then groups the file summaries into concept nodes joined by typed links. Every pass, the parse included, is cached by content hash, so the README’s own figures for this repository — 124 files — are 0.74 s cold, 0.18 s after one file is edited and 0.18 s with nothing changed; --no-reuse forces the cold path. That cheapness pays for the retrieval design. Rather than a watcher keeping an index warm, every query stats the tree against the previous build’s fingerprint — about 3 ms, structural, no model call — and rebuilds only if bytes moved, which is why ask, grep, callers, skeleton and map describe uncommitted and unstaged edits alike, and why GRAFT_REFRESH=hash exists for anyone who distrusts size and mtime. graft check is the deliberate exception: it never refreshes and exits 1 on drift, because it is the drift report. Injection rides the same artifacts: a marker-fenced block or an owned skill file written by an init command that knows nine hosts, six MCP tools the agent calls itself, and Claude Code hooks that probe each prompt and print an edited file’s blast radius afterwards.

  3. 03

    The speed and cost claims, and how each was measured

    The headline figure is the project’s own: up to 4× cheaper and 3× faster, with methods printed beside it. The first ran three variants of a Claude Sonnet 5 agent with the same file tools — cold, a graft ask --source bundle pushed up front, and a pull variant with only the MCP tools — scored by an Opus 4.8 judge against a required-keyword floor, with cache-aware costing at roughly 0.1× for reads and 1.25× for writes, over 162 runs, two repositories and three trials each. It reports cost per task falling from $0.0429 to $0.0292 (+32%), tokens from 8,070 to 4,650 (+42%), tool calls from 4.2 to 2.3 (+46%) and latency from 39.8 s to 15.8 s (+60%), correctness equal at 93%, and 98% for the pull variant that injects nothing. The second is SWE-bench Verified: 50 instances, the same model on both arms, the official 4.1.0 grader — 27 of 50 resolved against 33, 142.0M tokens against 109.4M, $52.34 against $42.43, 1,370 tool calls against 1,031, 13,094 s of wall clock against 8,922 s. The third re-implements five merged PocketBase pull requests from their base commits across 15 tasks: $13.91 down to $11.02, 2,044 s down to 1,762 s, all five reproduced. The table at the top of the README mixes the first two — its correctness row, 54% to 66%, is the SWE-bench result, while the controlled sweep found correctness unchanged. Nothing in the material is a third-party replication of any of it.

  4. 04

    What users measured back

    Several of the thirty most recent issues are counter-measurements rather than bug reports, and they are the sharpest material in the history. One user compared graft ask --source on a 291-file Swift and Kotlin application against their own baseline of one ripgrep plus a targeted read: the build took 3.7 s and produced 3,217 nodes and 9,743 edges, the right code was in the pack for all three questions, but the top hit was correct for only one, and the packs were 2.5× to 4× larger than the grep path — a ranking complaint with numbers attached. Another paired 331 prompt-hook outputs with the prompts that produced them and found that most of the injected lines saying there was no strong match came from machine-generated background-task notifications rather than from user prompts. A third documented a plain folder of repositories, never indexed, where every turn marked it dirty and every Stop spawned a build that hit its 120-second timeout: 0.8 of a core and 2.5 GB of RAM held continuously, the stats file reading syncedAt: null with zero nodes. That one also carries the clearest admission in the material — a folder polluted by an earlier version stays adopted, so the new gate still lets it through until the stale cache is deleted by hand.

  5. 05

    Generated files that overwrite what people edited

    Two of the loudest bugs share a cause: graft writes into files that are not wholly its own. In one, the session-start wiring refresh read .claude/settings.json through a bare JSON.parse inside a try/catch that fell through to an empty object whenever parsing failed, then wrote the merge back through a truncate. A file left holding merge-conflict markers — or read while another writer had it — therefore came back with the user’s permissions.deny, permissions.allow, outputStyle and hooks gone, and modified enough to commit. The pull request fixing it removes the fall-through, on the principle that a settings file which fails to parse must never be overwritten. The second case is one user filing seven issues in a single sitting about the generated instruction block: regeneration silently reverts hand-written corrections inside the graft marker fences, deleting the generated ignore file was itself reverted within the same session, graft build appends the wiring files to .gitignore while its output mentions only the cache directory, the template points at graft/INDEX.md while the build writes graft/index.md — a dead path on any case-sensitive filesystem — and the ignore file re-admits the card tree through an unanchored pattern. The node files were designed for exactly this pressure: anything a user writes below the generated block survives regeneration.

  6. 06

    Adapters, one closed pull request, and a metric checked first

    Much of the pull request traffic is one genre: making a model gateway behave. One reports that OpenAI-compatible gateways answer HTTP 200 with an error object and no choices, which threw a TypeError inside the response parser, and that some models cannot serve a forced tool_choice at all, leaving files and concept batches empty, reproduced against OpenRouter. The follow-up retries a structured call a reasoning model cut off at its length limit, with four times the token allowance, capped at 32,768. A third fixes a crux pass where models echo the whole target row back as the identifier, dropping every summary. The newest parallelizes concept-synthesis batches behind a new --synth-concurrency flag defaulting to four, with a second knob because phase one is hundreds of small calls and phase two a few 48,000-character requests; the graph stays byte-identical to the serial run. Around those sit the maintainer’s habits: a 52-file pull request bundling a SQLite stats store, five host adapters and a raised Node floor was closed unmerged, asking for an issue first and focused changes keeping the Node 20 floor; the telemetry event added for the hosted Trail workflow carries closed sets and buckets only and is pinned by a contract test; and before adding a count of suggestions waiting for review, the author checked the database, found zero such events in a week, and shipped it anyway.

Adjacent records

All records →