Skip to content

ripwire

A C++23 command-line tool and MCP server that parses a repository into a ranked call graph, then answers what a symbol reaches, what a change breaks, which tests reach it and what the change made worse.

Screenshot of ripwire
Editor screenshot, 30 Sep 2026ripwire ↗

What it is

A C++23 binary that reads a repository and prints a ranked, deterministic map of it, for the moment a coding agent is about to spend tokens reading files. One serial crawl and one tree-sitter query engine per language feed a Personalized PageRank pass; the verbs read that graph — --callers, --deps, --arch, --impact and --situ before a change, --edit-check and --quality-delta after one, and --for=TASK or --pack-task to serve a token-budgeted bundle for a named task. The same computation is exposed to agents over MCP: 33 verbs that are, in the architecture document’s words, "a thin front door onto the same computation and the same renderer as a CLI sibling." Build it with cmake -S . -B build && cmake --build build -j; register the server as {"type":"stdio","command":"ripwire","args":["--mcp"]}; skills/install.sh --hook installs the routing hooks and skill files. Output is minified XML on stdout and byte-identical between runs, and the crawler discloses what it dropped — size ceilings, git-ignored files, unresolved calls — rather than letting it disappear.

Who built itThe repository is owned by the GitHub organization redhat-et, whose name expands to Red Hat Emerging Technologies. That is the whole of the attribution: the repository metadata carries no owner field, and the README — 209,734 characters — does not contain the string "Red Hat" anywhere, so the affiliation rests on the organization name and on the paths inside the repository rather than on anything the project says about itself. The report’s only direct Red Hat marker is in the commit-email histogram, where dbrewste@redhat.com appears on 1,651 of the 4,135 commits. Which team inside Red Hat this came out of is not covered by the material. The work itself is one-sided: joyful-ii-V-I accounts for 3,554 of the commits with a linked account, and the 25-name contributor list is led by barefootski (346), quaterniondrift (93) and lennix1337 (33).

How it is put together

The parts · 6

A pipeline of five stages with no back edges — ingest → graph → rank → serialize → cli / mcp — where each stage’s output is a plain data structure, which is what makes one stage testable on its own and lets every verb be a different reading of the same graph. Three consequences run through the code. The crawl is serial and sorts every candidate path before it assigns node IDs, because node IDs are indices into that list and any dependence on directory order would make IDs, top-K cutoffs and every diff-aware verb churn between identical runs; the parse pool is where the threads are. Extraction is one query engine over a queries/<language>/tags.scm per language rather than a traversal written per language — the architecture document refuses the other route by name. And what the crawler drops is a decision with a number behind it rather than a default: a generic 4 MB ceiling, JSON skipped over 256 KB or past 512 nesting levels, YAML given its own 512 KB line because JSON’s would drop hand-maintained configuration, no TOML ceiling at all because a measurement put the largest of 321 files at 57,759 B, and every size drop counted into the header’s skipped_oversize= rather than vanishing quietly.

src/ingest*.h
The ingest stage: one translation unit with src/ingest.cpp (29,564 bytes) as its spine and its families in guarded sections — ingest_cache.h (255,885 bytes), ingest_relations.h (140,685), ingest_names.h (130,361), ingest_sidecap.h (130,077), ingest_binds.h (123,635), ingest_crawl.h (116,419), ingest_astquery.h (108,592), plus the prewarm, parse-pool, document post-pass and model tails. src/ overall is 151 files and 11,275 KB.
src/graph.h and src/pagerank.cpp
The graph and rank stages. graph.h is 453,160 bytes; pagerank.cpp is 10,939 with a 2,174-byte header, kept small because the determinism rule forbids floating-point reassociation in that translation unit. Symbols are ranked by Personalized PageRank, and that ordering is what most other verbs rest on.
src/main.cpp, src/cli.h and src/verbs_*.h
The verb families and the shared output machinery. serialize.h is 564,798 bytes, cli.h 523,742 and quality.h 541,224, with main.cpp at 347,197. The families themselves are verbs_for.h (287,589), verbs_report.h (241,378), verbs_navigate.h (175,046), verbs_quality.h (160,168), verbs_lint.h (126,375), verbs_change.h (91,385), verbs_grep.h (71,657) and verbs_doctor.h (66,917). src/infra/ holds 39 files, including sortutil.h (9,889) — the helper that five comparator sites were routed through after #342.
src/mcp*.h
The MCP surface: 33 verbs that are, by the architecture document’s own description, a thin front door onto the same computation and the same renderer as the CLI. mcpverbs.h is 352,422 bytes, mcp.h 201,355, mcprefusal.h 96,085, mcpedit.h 86,563, mcpindex.h 77,824, mcpjson.h 43,368 and mcpserver.h 30,626. A 101-byte .mcp.json registers the server for this repository itself.
queries/ and third_party/
Extraction data and the vendored tree. queries/ holds 23 <language>/tags.scm files from 1,367 bytes (bash) to 28,032 (C++). third_party/deps holds 211 files and 242,952 KB — tree-sitter core, doctest, and grammars for bash, C, C++, C#, CUDA, Dart, Elixir, GDScript, Go, Java, JavaScript, JSON, Kotlin, Lua, Markdown, Objective-C, PHP, Python, Ruby, Rust, Swift, TOML, TypeScript and YAML. third_party/patches/ carries 14 patches against those vendored scanners, and third_party/ itself the header-only pieces, with CMakeLists.txt at 100,647 bytes.
test/, skills/ and hooks/
The gate suite and the agent-side plumbing. test/ is 718 files and 13,573 KB, of which 661 are *check.sh gates driven by test/pargates.py, registered in test/regression.sh and audited by test/manifestcheck.sh; test/fuzz/ alone is 90 files. skills/ holds install.sh (33,469 bytes) and 17 skill documents, hooks/ five shell hooks with ripwire-nudge.sh at 107,068 bytes, and .claude/skills/, .codex-plugin/ and .coderabbit.yaml carry the per-agent configuration.

Choices, and what they beat

  • The crawl is serial and folds its candidate list before any node ID exists over assigning IDs in directory order and parallelising the walk

    Node IDs are indices into the sorted candidate list, so the document states the consequence rather than the preference: directory order would make node IDs, top-K cutoffs and every diff-aware verb churn between identical runs. It is equally explicit about where the parallelism went — the crawl "is the cheap half and is deliberately not parallelized; the parse pool is where the threads are."

  • One query engine reading a tags.scm per language over a bespoke AST traversal per language

    The architecture document names the alternative and forbids it: "a bespoke AST traversal per language — is forbidden here", because that route is "five fragile walkers that break on every grammar bump instead of one query loop that survives them." The per-language work is then data — a query file plus a capture-name table — which is also why the repository can carry 23 of them.

  • Rails schema columns are definitions with no edges over letting create_table columns take call edges and PageRank weight

    Stated in the pull request as "Definitions only (maintainer decision after review)." A column minted from a create_table block is recognized by content rather than by path and is answered by --uses, --whereis, --grep and the map, but buildGraph’s byName skips Section && Lang::Ruby, so --callers=id reports defs="14" count="0" — the definition is found and the edge is refused.

  • TOML gets no lane-specific ceiling, YAML gets 512 KB rather than JSON’s 256 KB over giving every configuration format the same ceiling

    Written as "TOML has no lane-specific ceiling, and that is a measured decision rather than a missing sibling": over 90 public repositories and 321 .toml files the largest is 57,759 B, so a ceiling could not sit both above the observed maximum and below the generic 4 MB skip. YAML is the opposite case — JSON’s 256 KB line would drop real hand-maintained configuration, since NeMo’s cicd-main.yml is 293 KB.

  • Directory symlinks are not followed at all over following them and tracking inodes to break cycles

    "There is no inode tracking, because with symlink-following off there is nothing for it to do." The walk is opened with skip_permission_denied only, so a symlinked directory is never descended into and a cycle cannot arise in the first place.

Read fromdocs/ARCHITECTURE.md (50,755 bytes, 50,456 characters — the report prints the first 12,000), AGENTS.md (2,930 characters), the architecture-document listing, the complete 2,951-file tree with per-file sizes, and the pull requests that state a decision in their own words (#339, #359, #363).

Build log

6 stages
  1. 01

    A repository that appeared at the end of July and shipped twenty releases by September

    The repository was created on 2026-07-29 and its oldest commit is dated 2026-07-31, titled import from internal development tree — the only account in the material of where the code came from before that. From there it moved at a rate this archive rarely records: 4,135 commits, of which 17 fall in July, 1,294 in August and 2,824 in September, against 2,374 stars, 151 forks, 8 watchers, 58 open issues, Apache-2.0 and C++. It has published 20 releases, every one marked neither prerelease nor draft, from v0.1.0 on 2026-08-02 — titled ripwire v0.1.0 — the ripgrep of AI context — to v0.6.5 on 2026-09-27, and eight of them land within 29 hours of each other: v0.3.0 at 23:31 on 2026-08-11 through v0.3.8 at 04:11 on 2026-08-13. The tag list carries v0.3.7, which has no release entry, and does not carry v0.1.0; it is 20 entries long, the same as the release list. 3,205 co-author trailers were counted, led by Claude Opus 5 at 1,417, Claude Fable 5 at 698 and Claude Fable 5.1 at 543, with Qwen3.8 Max at 2. The three commit histograms disagree with one another — joyful-ii-V-I is 3,993 commits by author name and 3,554 by linked account, and the email tally is led by dbrewste@redhat.com at 1,651.

  2. 02

    Five stages with no back edges, and a gate for every claim

    The architecture document describes five stages — ingest → graph → rank → serialize → cli / mcp — and states its shape: "Five stages, in that order, with no back edges. Each stage’s output is a plain data structure, so any stage can be tested in isolation and every verb is a different way of reading the same graph." Two choices in the first stage carry the weight. The crawl collects every candidate path, sorts them by byte and only then assigns node IDs, because the IDs are indices into that list; directory order instead would make IDs, top-K cutoffs and every diff-aware verb churn between identical runs. And the walk stays serial — the "cheap half", not parallelized, with the threads in the parse pool. Extraction is one query engine over a queries/<language>/tags.scm per language, 23 files in the tree; the other route is refused by name: "a bespoke AST traversal per language — is forbidden here". The same instinct governs the tests. AGENTS.md states it as "Gate first, code second", and a new test/*check.sh must be added to test/regression.sh in the same commit, or test/manifestcheck.sh fails. test/ holds 718 files, 661 of them *check.sh gates, and the count moves weekly: 649 gates in #337, 650 in #341, 651 in a contributor’s report, and 665 gates, 657 pass, 4 environment skips and 4 failures in a full pargates.py -j 6 run on a loaded machine.

  3. 03

    The home directory that cost 67 GB

    Issue #350, opened by KilimcininKorOglu on 2026-09-27, is the failure the memory guard exists for. With the MCP server registered in Claude Code as {"type":"stdio","command":"ripwire","args":["--mcp"]} and a session opened in the home directory, one grep call drove the --mcp process to a 67 GB phys_footprint in seven hours; swap reached 28 GB with 458 MB free, on an Apple M1 Max with 64 GB of RAM running macOS 27.0. The cause is in the report and was confirmed by the maintainer: in a root that is not a git repository, .gitignore cannot be applied, so the crawl maps everything under the directory and nothing bounds the memory that takes. #363 answers it in layers. A root nobody chose is refused — $HOME itself even when it is a git repository, a filesystem or drive root, the parent of the home directories, an OS tree such as /System, /usr, /etc, /proc or %WINDIR% — with one line, "no project root: <dir> is a home/system directory; pass a project path", and the same rule covers an MCP request’s path=. Proving it needed a test seam: RIPWIRE_TEST_MEMGUARD=hard:N makes the Nth hard-limit reading count as over the limit, and arm B14 runs two tiny roots under hard:1 and asserts one ingest-cache blob and a stderr line naming root 1 of 2. Arm B12 runs with AddressSanitizer’s quarantine off, because the quarantine inflates the footprint being measured.

  4. 04

    A sanitizer lane that had never once completed

    On 2026-09-27 llvm-x86 opened #342 with the finding that the documented Linux sanitizer ritual cannot complete on main, and filed six sub-issues in the same minute — #343 through #348. The cause is a library detail: libstdc++ computes string_view ordering as n1 - n2 in size_type, which wraps when one string is a prefix of a longer one, and the G1 lane runs -fsanitize=integer with -fno-sanitize-recover=all, so the documented self-run aborts on “unsigned integer overflow: 4 - 16 cannot be represented in type size_type”. Five sites compared std::string_view operands with a raw < — pathInIgnoreSet in src/ingest_crawl.h, two in src/situ.h and two in finalizeNamedIdents in src/mention.h. The fix routes all five through rw::sortutil::svLess; the maintainer’s note says what was and was not broken: "The ordering is still correct, so release binaries answer correctly." Two were not about C++. Two gate scripts had been committed with mode 100644 in a suite where every other *check.sh is 100755, so invoking them directly exited 126 — and CI, which runs them through a shell, never noticed. And the documented GCC -DRIPWIRE_ASAN=ON path did not build, because a constexpr predicate fed a pointer from findByField into a static_assert. #351 closed #342 and all six sub-issues in one commit, with the new arm #6c red on main and green with the fix.

  5. 05

    What a report has to look like to be trusted

    #353, filed by toasterbook88, is a report about a report. Running --scan-skills over a 2,077-file skill home returned 156 findings, 107 of them EXFILTRATE:net-exfil; the reporter classified all 107 by destination and isolated the trigger to three lines. The maintainer agreed: the rule fires on "a network verb plus any $VAR on the same line" and cannot tell where a variable’s value came from, so documented API calls and loopback checks were raised as critical. The second comment is what makes it worth keeping. The reporter re-ran the scan with "a single frozen script" and published a corrected histogram, writing that the first version "was wrong in two places — one from over-claiming, one from a too-narrow var-trace — and I’d rather correct it explicitly than leave a number standing that a reader can’t reproduce." After the correction about 89 of the 107 rows were documented service-API calls. The maintainer’s reply is a standard worth quoting: "It’s exactly the kind of care that makes a report trustworthy. You withdrew the claims you couldn’t back, froze the classifier, and published the row-level TSV." The follow-through goes the same way: #359 credits NVIDIA’s SkillSpector in docs/LINEAGE.md, taking only its code-based checks — "only the code checks come across. The agent’s LLM reads ripwire’s output and does the judging".

  6. 06

    Numbers the author publishes with their own limits attached

    docs/EVALS.md is 1,132,024 bytes and docs/COMMANDS.md 720,784, which is what a project looks like when measured claims are a deliverable rather than an afterthought. The README also carries a field report, added by #361 and collapsed near the top, credited to "Claude Fable 5.0, the frontier model orchestrating ~20 coding agents over two days on a ~1,500-file C++/Metal codebase". Its numbers are large and its labels are explicit: roughly half the audit and research token spend, called "the operator’s whole-phase estimate"; a single --callers query that returned zero production callers; --edit-check flagging 5 of 6 call sites a text search had missed; "a dozen-plus real code-quality defects fixed, not waived". The pull-request body states the boundary in the same breath — one engagement, on an earlier version, "the orchestrating model’s own report and not a controlled measurement" — and the full report is a click-to-view image, docs/assets/field-report-multi-agent.jpg at 1144×1600 and 567 KB, with a text version underneath added for accessibility. The README’s own figures sit in the same frame: an index built in 0.25–0.45 s, a strict file@10 of 58.3% against 40.0%, signatures at 74.7% fewer bytes than bodies, and 14–25% from profile-guided optimisation — all of them the author’s, and the README says the timing columns have not been re-measured.

Adjacent records

All records →