ripwire
A C++23 command-line tool and MCP server that parses a repository into a ranked call graph, then answers what a symbol reaches, what a change breaks, which tests reach it and what the change made worse.

What it is
A C++23 binary that reads a repository and prints a ranked, deterministic map of it, for the moment a coding agent is about to spend tokens reading files. One serial crawl and one tree-sitter query engine per language feed a Personalized PageRank pass; the verbs read that graph — --callers, --deps, --arch, --impact and --situ before a change, --edit-check and --quality-delta after one, and --for=TASK or --pack-task to serve a token-budgeted bundle for a named task. The same computation is exposed to agents over MCP: 33 verbs that are, in the architecture document’s words, "a thin front door onto the same computation and the same renderer as a CLI sibling." Build it with cmake -S . -B build && cmake --build build -j; register the server as {"type":"stdio","command":"ripwire","args":["--mcp"]}; skills/install.sh --hook installs the routing hooks and skill files. Output is minified XML on stdout and byte-identical between runs, and the crawler discloses what it dropped — size ceilings, git-ignored files, unresolved calls — rather than letting it disappear.
Who built itThe repository is owned by the GitHub organization redhat-et, whose name expands to Red Hat Emerging Technologies. That is the whole of the attribution: the repository metadata carries no owner field, and the README — 209,734 characters — does not contain the string "Red Hat" anywhere, so the affiliation rests on the organization name and on the paths inside the repository rather than on anything the project says about itself. The report’s only direct Red Hat marker is in the commit-email histogram, where dbrewste@redhat.com appears on 1,651 of the 4,135 commits. Which team inside Red Hat this came out of is not covered by the material. The work itself is one-sided: joyful-ii-V-I accounts for 3,554 of the commits with a linked account, and the 25-name contributor list is led by barefootski (346), quaterniondrift (93) and lennix1337 (33).
How it is put together
The parts · 6A pipeline of five stages with no back edges — ingest → graph → rank → serialize → cli / mcp — where each stage’s output is a plain data structure, which is what makes one stage testable on its own and lets every verb be a different reading of the same graph. Three consequences run through the code. The crawl is serial and sorts every candidate path before it assigns node IDs, because node IDs are indices into that list and any dependence on directory order would make IDs, top-K cutoffs and every diff-aware verb churn between identical runs; the parse pool is where the threads are. Extraction is one query engine over a queries/<language>/tags.scm per language rather than a traversal written per language — the architecture document refuses the other route by name. And what the crawler drops is a decision with a number behind it rather than a default: a generic 4 MB ceiling, JSON skipped over 256 KB or past 512 nesting levels, YAML given its own 512 KB line because JSON’s would drop hand-maintained configuration, no TOML ceiling at all because a measurement put the largest of 321 files at 57,759 B, and every size drop counted into the header’s skipped_oversize= rather than vanishing quietly.
- src/ingest*.h
- The ingest stage: one translation unit with
src/ingest.cpp(29,564 bytes) as its spine and its families in guarded sections —ingest_cache.h(255,885 bytes),ingest_relations.h(140,685),ingest_names.h(130,361),ingest_sidecap.h(130,077),ingest_binds.h(123,635),ingest_crawl.h(116,419),ingest_astquery.h(108,592), plus the prewarm, parse-pool, document post-pass and model tails.src/overall is 151 files and 11,275 KB. - src/graph.h and src/pagerank.cpp
- The graph and rank stages.
graph.his 453,160 bytes;pagerank.cppis 10,939 with a 2,174-byte header, kept small because the determinism rule forbids floating-point reassociation in that translation unit. Symbols are ranked by Personalized PageRank, and that ordering is what most other verbs rest on. - src/main.cpp, src/cli.h and src/verbs_*.h
- The verb families and the shared output machinery.
serialize.his 564,798 bytes,cli.h523,742 andquality.h541,224, withmain.cppat 347,197. The families themselves areverbs_for.h(287,589),verbs_report.h(241,378),verbs_navigate.h(175,046),verbs_quality.h(160,168),verbs_lint.h(126,375),verbs_change.h(91,385),verbs_grep.h(71,657) andverbs_doctor.h(66,917).src/infra/holds 39 files, includingsortutil.h(9,889) — the helper that five comparator sites were routed through after #342. - src/mcp*.h
- The MCP surface: 33 verbs that are, by the architecture document’s own description, a thin front door onto the same computation and the same renderer as the CLI.
mcpverbs.his 352,422 bytes,mcp.h201,355,mcprefusal.h96,085,mcpedit.h86,563,mcpindex.h77,824,mcpjson.h43,368 andmcpserver.h30,626. A 101-byte.mcp.jsonregisters the server for this repository itself. - queries/ and third_party/
- Extraction data and the vendored tree.
queries/holds 23<language>/tags.scmfiles from 1,367 bytes (bash) to 28,032 (C++).third_party/depsholds 211 files and 242,952 KB — tree-sitter core, doctest, and grammars for bash, C, C++, C#, CUDA, Dart, Elixir, GDScript, Go, Java, JavaScript, JSON, Kotlin, Lua, Markdown, Objective-C, PHP, Python, Ruby, Rust, Swift, TOML, TypeScript and YAML.third_party/patches/carries 14 patches against those vendored scanners, andthird_party/itself the header-only pieces, withCMakeLists.txtat 100,647 bytes. - test/, skills/ and hooks/
- The gate suite and the agent-side plumbing.
test/is 718 files and 13,573 KB, of which 661 are*check.shgates driven bytest/pargates.py, registered intest/regression.shand audited bytest/manifestcheck.sh;test/fuzz/alone is 90 files.skills/holdsinstall.sh(33,469 bytes) and 17 skill documents,hooks/five shell hooks withripwire-nudge.shat 107,068 bytes, and.claude/skills/,.codex-plugin/and.coderabbit.yamlcarry the per-agent configuration.
Choices, and what they beat
The crawl is serial and folds its candidate list before any node ID exists over assigning IDs in directory order and parallelising the walk
Node IDs are indices into the sorted candidate list, so the document states the consequence rather than the preference: directory order would make node IDs, top-K cutoffs and every diff-aware verb churn between identical runs. It is equally explicit about where the parallelism went — the crawl "is the cheap half and is deliberately not parallelized; the parse pool is where the threads are."
One query engine reading a
tags.scmper language over a bespoke AST traversal per languageThe architecture document names the alternative and forbids it: "a bespoke AST traversal per language — is forbidden here", because that route is "five fragile walkers that break on every grammar bump instead of one query loop that survives them." The per-language work is then data — a query file plus a capture-name table — which is also why the repository can carry 23 of them.
Rails schema columns are definitions with no edges over letting
create_tablecolumns take call edges and PageRank weightStated in the pull request as "Definitions only (maintainer decision after review)." A column minted from a
create_tableblock is recognized by content rather than by path and is answered by--uses,--whereis,--grepand the map, butbuildGraph’sbyNameskipsSection && Lang::Ruby, so--callers=idreportsdefs="14" count="0"— the definition is found and the edge is refused.TOML gets no lane-specific ceiling, YAML gets 512 KB rather than JSON’s 256 KB over giving every configuration format the same ceiling
Written as "TOML has no lane-specific ceiling, and that is a measured decision rather than a missing sibling": over 90 public repositories and 321
.tomlfiles the largest is 57,759 B, so a ceiling could not sit both above the observed maximum and below the generic 4 MB skip. YAML is the opposite case — JSON’s 256 KB line would drop real hand-maintained configuration, since NeMo’scicd-main.ymlis 293 KB.Directory symlinks are not followed at all over following them and tracking inodes to break cycles
"There is no inode tracking, because with symlink-following off there is nothing for it to do." The walk is opened with
skip_permission_deniedonly, so a symlinked directory is never descended into and a cycle cannot arise in the first place.
Read fromdocs/ARCHITECTURE.md (50,755 bytes, 50,456 characters — the report prints the first 12,000), AGENTS.md (2,930 characters), the architecture-document listing, the complete 2,951-file tree with per-file sizes, and the pull requests that state a decision in their own words (#339, #359, #363).
Build log
6 stages- 01
A repository that appeared at the end of July and shipped twenty releases by September
The repository was created on 2026-07-29 and its oldest commit is dated 2026-07-31, titled
import from internal development tree— the only account in the material of where the code came from before that. From there it moved at a rate this archive rarely records: 4,135 commits, of which 17 fall in July, 1,294 in August and 2,824 in September, against 2,374 stars, 151 forks, 8 watchers, 58 open issues, Apache-2.0 and C++. It has published 20 releases, every one marked neither prerelease nor draft, fromv0.1.0on 2026-08-02 — titledripwire v0.1.0 — the ripgrep of AI context— tov0.6.5on 2026-09-27, and eight of them land within 29 hours of each other:v0.3.0at 23:31 on 2026-08-11 throughv0.3.8at 04:11 on 2026-08-13. The tag list carriesv0.3.7, which has no release entry, and does not carryv0.1.0; it is 20 entries long, the same as the release list. 3,205 co-author trailers were counted, led byClaude Opus 5at 1,417,Claude Fable 5at 698 andClaude Fable 5.1at 543, withQwen3.8 Maxat 2. The three commit histograms disagree with one another —joyful-ii-V-Iis 3,993 commits by author name and 3,554 by linked account, and the email tally is led bydbrewste@redhat.comat 1,651. - 02
Five stages with no back edges, and a gate for every claim
The architecture document describes five stages —
ingest → graph → rank → serialize → cli / mcp— and states its shape: "Five stages, in that order, with no back edges. Each stage’s output is a plain data structure, so any stage can be tested in isolation and every verb is a different way of reading the same graph." Two choices in the first stage carry the weight. The crawl collects every candidate path, sorts them by byte and only then assigns node IDs, because the IDs are indices into that list; directory order instead would make IDs, top-K cutoffs and every diff-aware verb churn between identical runs. And the walk stays serial — the "cheap half", not parallelized, with the threads in the parse pool. Extraction is one query engine over aqueries/<language>/tags.scmper language, 23 files in the tree; the other route is refused by name: "a bespoke AST traversal per language — is forbidden here". The same instinct governs the tests.AGENTS.mdstates it as "Gate first, code second", and a newtest/*check.shmust be added totest/regression.shin the same commit, ortest/manifestcheck.shfails.test/holds 718 files, 661 of them*check.shgates, and the count moves weekly: 649 gates in #337, 650 in #341, 651 in a contributor’s report, and 665 gates, 657 pass, 4 environment skips and 4 failures in a fullpargates.py -j 6run on a loaded machine. - 03
The home directory that cost 67 GB
Issue #350, opened by
KilimcininKorOgluon 2026-09-27, is the failure the memory guard exists for. With the MCP server registered in Claude Code as{"type":"stdio","command":"ripwire","args":["--mcp"]}and a session opened in the home directory, onegrepcall drove the--mcpprocess to a 67 GBphys_footprintin seven hours; swap reached 28 GB with 458 MB free, on an Apple M1 Max with 64 GB of RAM running macOS 27.0. The cause is in the report and was confirmed by the maintainer: in a root that is not a git repository,.gitignorecannot be applied, so the crawl maps everything under the directory and nothing bounds the memory that takes. #363 answers it in layers. A root nobody chose is refused —$HOMEitself even when it is a git repository, a filesystem or drive root, the parent of the home directories, an OS tree such as/System,/usr,/etc,/procor%WINDIR%— with one line, "no project root: <dir> is a home/system directory; pass a project path", and the same rule covers an MCP request’spath=. Proving it needed a test seam:RIPWIRE_TEST_MEMGUARD=hard:Nmakes the Nth hard-limit reading count as over the limit, and arm B14 runs two tiny roots underhard:1and asserts one ingest-cache blob and a stderr line naming root 1 of 2. Arm B12 runs with AddressSanitizer’s quarantine off, because the quarantine inflates the footprint being measured. - 04
A sanitizer lane that had never once completed
On 2026-09-27
llvm-x86opened #342 with the finding that the documented Linux sanitizer ritual cannot complete onmain, and filed six sub-issues in the same minute — #343 through #348. The cause is a library detail: libstdc++ computesstring_viewordering asn1 - n2insize_type, which wraps when one string is a prefix of a longer one, and the G1 lane runs-fsanitize=integerwith-fno-sanitize-recover=all, so the documented self-run aborts on “unsigned integer overflow: 4 - 16 cannot be represented in typesize_type”. Five sites comparedstd::string_viewoperands with a raw<—pathInIgnoreSetinsrc/ingest_crawl.h, two insrc/situ.hand two infinalizeNamedIdentsinsrc/mention.h. The fix routes all five throughrw::sortutil::svLess; the maintainer’s note says what was and was not broken: "The ordering is still correct, so release binaries answer correctly." Two were not about C++. Two gate scripts had been committed with mode 100644 in a suite where every other*check.shis 100755, so invoking them directly exited 126 — and CI, which runs them through a shell, never noticed. And the documented GCC-DRIPWIRE_ASAN=ONpath did not build, because aconstexprpredicate fed a pointer fromfindByFieldinto astatic_assert. #351 closed #342 and all six sub-issues in one commit, with the new arm #6c red onmainand green with the fix. - 05
What a report has to look like to be trusted
#353, filed by
toasterbook88, is a report about a report. Running--scan-skillsover a 2,077-file skill home returned 156 findings, 107 of themEXFILTRATE:net-exfil; the reporter classified all 107 by destination and isolated the trigger to three lines. The maintainer agreed: the rule fires on "a network verb plus any$VARon the same line" and cannot tell where a variable’s value came from, so documented API calls and loopback checks were raised as critical. The second comment is what makes it worth keeping. The reporter re-ran the scan with "a single frozen script" and published a corrected histogram, writing that the first version "was wrong in two places — one from over-claiming, one from a too-narrow var-trace — and I’d rather correct it explicitly than leave a number standing that a reader can’t reproduce." After the correction about 89 of the 107 rows were documented service-API calls. The maintainer’s reply is a standard worth quoting: "It’s exactly the kind of care that makes a report trustworthy. You withdrew the claims you couldn’t back, froze the classifier, and published the row-level TSV." The follow-through goes the same way: #359 credits NVIDIA’s SkillSpector indocs/LINEAGE.md, taking only its code-based checks — "only the code checks come across. The agent’s LLM reads ripwire’s output and does the judging". - 06
Numbers the author publishes with their own limits attached
docs/EVALS.mdis 1,132,024 bytes anddocs/COMMANDS.md720,784, which is what a project looks like when measured claims are a deliverable rather than an afterthought. The README also carries a field report, added by #361 and collapsed near the top, credited to "Claude Fable 5.0, the frontier model orchestrating ~20 coding agents over two days on a ~1,500-file C++/Metal codebase". Its numbers are large and its labels are explicit: roughly half the audit and research token spend, called "the operator’s whole-phase estimate"; a single--callersquery that returned zero production callers;--edit-checkflagging 5 of 6 call sites a text search had missed; "a dozen-plus real code-quality defects fixed, not waived". The pull-request body states the boundary in the same breath — one engagement, on an earlier version, "the orchestrating model’s own report and not a controlled measurement" — and the full report is a click-to-view image,docs/assets/field-report-multi-agent.jpgat 1144×1600 and 567 KB, with a text version underneath added for accessibility. The README’s own figures sit in the same frame: an index built in 0.25–0.45 s, a strict file@10 of 58.3% against 40.0%, signatures at 74.7% fewer bytes than bodies, and 14–25% from profile-guided optimisation — all of them the author’s, and the README says the timing columns have not been re-measured.
Adjacent records
All records →No. 091
OKF Agent Memory
Keeps what a coding agent learns as plain Markdown inside the repository — an OKF v0.2 knowledge bundle searched in-process by BM25 — so the memory can be diffed and reviewed instead of living in a database.
No. 090
Reverify
A verifier for an AI’s claims about binaries: the model proposes, a deterministic toolkit checks the claim against the actual bytes, and what comes back is VERIFIED, REFUTED or INCONCLUSIVE with the evidence it read — the verdicts, not the prose, are what survives a reset.
No. 081
pgbot
A single static Go binary that connects to PostgreSQL read-only, reads the server’s own statistics views and prints a graded, findings-first health report — and, because every run saves a local baseline, tells you what changed since last time; the same deterministic findings are served to AI agents over MCP, and the optional AI layer may only explain them.