Skip to content

Headroom

A compression layer that sits between an agent and the model it calls: tool output, logs, file reads, RAG chunks and conversation history are shrunk on the machine that runs it, and the originals are kept in a local store so the model can ask for any block back.

Screenshot of Headroom
Editor screenshot, 1 Oct 2026Headroom ↗

What it is

Headroom is a compression layer for the traffic an agent sends to a model. It comes in three shapes — a library whose whole API is one compress() call in Python or TypeScript, a FastAPI proxy that any OpenAI-compatible client can point at, and an MCP server exposing headroom_compress, headroom_retrieve and headroom_stats — plus wrapper commands that start the proxy and configure fifteen named coding agents. Inside, a ContentRouter decides what each content block is and hands it to exactly one compressor: SmartCrusher for arrays of JSON records, a search compressor, a log compressor, a diff compressor, an HTML extractor, tabular and config compressors, a text crusher, and an opt-in ModernBERT model called Kompress for prose. The heaviest of them run in a Rust extension loaded through PyO3, and Kompress runs through ONNX Runtime. Only the newest blocks are compressed, earlier turns are forwarded byte for byte so a provider prefix cache survives, and every compression is reversible because the original is written to a local store the model can retrieve from.

Who built itThe repository belongs to headroomlabs-ai, an organisation account rather than a personal one — its commit emails carry the headroomlabs.ai domain, and the README sells a supported or fully managed deployment alongside the open-source core. The largest single committer is chopratejas (Tejas Chopra), with 1,234 of the 3,000 commits in the sample; JerrettDavis follows with 373, abhay-codes07 with 169, gglucass with 139 and rodboev with 102, and dependabot accounts for 76. The contributor list names fifty accounts, and 2,984 of the 3,000 commits carry a linked GitHub account.

How it is put together

The parts · 6

A compression layer placed in front of whatever an agent is talking to. One request lifecycle is shared by the library, the SDKs and the proxy, and inside it a short ordered pipeline runs: an opt-in tool-result interceptor, a detector for prefix drift, and then the ContentRouter, which is the stage that actually rewrites content. Everything else follows from three consequences. The first is that compression must not disturb the provider cache, so only the newest content blocks are compressed and earlier turns are forwarded byte-identically — which is why the default mode is the one called cache. The second is that compression has to be reversible, so originals go into a local hash-indexed store and the model is handed a retrieval tool. The third is that a wrong compressor is worse than no compressor, so every transform fails open, the guard rails are configurable per savings profile, and the heaviest compressors were moved into a Rust extension while reference implementations stay in Python and are compared against recorded fixtures.

headroom/
The Python package, 472 KB across 33 top-level files and about thirty subpackages. transforms/ holds the compression itself — 39 files led by a 370 KB content_router.py, a 112 KB code compressor, a 99 KB Kompress wrapper and a 64 KB SmartCrusher. proxy/ holds 123 files behind a FastAPI app, including a 564 KB handlers/openai.py, a 311 KB handlers/anthropic.py, a 191 KB helpers.py and a 97 KB streaming handler. cache/ tracks prefix stability and cache TTL, ccr/ the Compress-Cache-Retrieve store and its MCP server, providers/ one slice per wrapped agent, and cli/ the commands — among them a 359 KB wrap.py — alongside evals/, memory/, image/, pricing/, telemetry/, tokenizers/, integrations/, learn/ and install/.
crates/
Five Rust crates. headroom-core carries the Rust implementations — SmartCrusher, the log, search, diff and code compressors, the live-zone dispatcher, content detection, a tokenizer registry and CCR back ends. headroom-proxy is the Rust proxy, with native Bedrock SigV4 and Vertex routes, an SSE state machine and a cache-stabilisation module. headroom-py is the PyO3 extension exposed to Python as headroom._core, and headroom-parity and headroom-simulators exist to compare and to fake both sides.
tests/ and tests/parity/
666 files. The 231 under tests/parity/ are recorded fixtures — the cache aligner, CCR, code-aware compression, content detection, diffs, Kompress, log compression, SmartCrusher, text crushing, the tokenizer and the Codex and OpenAI contract — regenerated by scripts whose names begin with record_. A separate tests/fixtures/fidelity_golden/ holds a generator, a baseline and 16 KB of cases, and tests/gateway/ drives a fake provider and a fake gateway through contract and invariant tests.
benchmarks/ and headroom/evals/
The measurement surface: 31 scripts and 626 KB, including the proof table generator with its checked-in result file, a 46 KB latency benchmark, worst-case and adversarial benchmarks, an i18n compression evaluation, a text-quality evaluation, cache-bust and prefix-cache traces, and Claude session comparisons; plus an eval suite of 26 files with datasets, adversarial grids, memory evaluations and a report card.
docs/, wiki/ and REALIGNMENT/
Three documentation layers rather than one. docs/content/docs/ holds 54 MDX pages, led by a 70 KB proxy page, a 35 KB configuration page, an API reference and a limitations page; wiki/ holds 35 pages plus three dated plan documents; and REALIGNMENT/ holds the fourteen-document internal audit, from a 26 KB bug list to a phase-by-phase plan and an index.
.github/workflows/, scripts/ and sbom/
The release and review machinery: 23 workflows, of which release.yml is 50 KB and ci.yml 28 KB, plus docker.yml, rust.yml, eval.yml, security.yml and a PR-health workflow; scripts for changelog generation, version syncing, pull-request governance, release smoke tests, ruff and tool-hash verification, and for recording the parity fixtures; and a committed SBOM directory of 8.5 MB holding CycloneDX and SPDX manifests and vulnerability scans.

Choices, and what they beat

  • Compress only the live zone, and never drop history over scoring the whole conversation and deleting old turns to fit the window

    The realignment document calls the dropping model the wrong mental model in as many words and traces five cache-killer bugs to it; what it prescribes is deleting the context manager, the scoring, the relevance and the rolling-window stages and compressing only the newest blocks. The architecture documentation records the result: the stage was removed, the pipeline no longer runs any dropping or scoring logic, and the classes are gone from the Python package — while the TypeScript SDK still exports the equivalent field names.

  • CacheAligner as a detector that never rewrites a prompt over extracting volatile parts and moving them to the end of the message

    The older architecture document describes the moving version in detail, down to the before-and-after of a dated system prompt. The current documentation says the stage never mutates, moves or rewrites content, is disabled by default and is hard-disabled inside the proxy, and exists to publish prefix-stability metrics. The behaviour was kept and the mutation was taken away, which is the more interesting half of the change.

  • Cache mode as the default over token mode, which maximises raw removal

    The two modes are documented side by side: cache compresses only the newest delta in a turn and forwards prior turns byte-faithfully so the provider prefix cache is never invalidated mid-conversation, while token mode may recompress earlier turns and trades cache stability for savings. The default follows from the project’s own observation that prefix caching is where most of the cost savings live on multi-turn workloads.

  • Reversible compression on by default, with opt-outs over a lossy pipeline with no retrieval path

    CCR writes the original of anything compressed into a local hash-indexed store and injects a retrieval tool so the model can ask for it back, and it is on by default; the markers and the tool can be switched off with --no-ccr, and a marker-free, format-native lossless mode exists behind --lossless. The default is the lossy one, and the escape hatches are documented next to it rather than buried.

  • Report output-token savings as an estimate over presenting a single measured figure

    The README states the reason: what the model would have written is never observed, so the figure is counterfactual and is therefore printed with a confidence band and the label estimated. A ten per cent holdout of conversations is offered as the way to get a measured number instead, and the dashboard is documented as switching the card from estimated to measured when it is set.

Read fromdocs/content/docs/architecture.mdx (8,593 characters), wiki/ARCHITECTURE.md (37,720 characters), REALIGNMENT/00-overview.md, plugins/headroom-oauth2/SPEC.md, README.md (about 33 KB), the complete 2,489-file tree with sizes, the 666-file test suite with its 231 parity fixtures and fidelity golden files, and the 31 benchmark scripts.

Build log

6 stages
  1. 01

    A nine-month-old repository with 74,192 stars, and nothing in the material that explains it

    The repository was created on 2026-01-07 and the newest commit in the sample is dated 2026-10-01; in between sit 3,000 commits, and the pull-request numbering has already passed #3896. The GitHub API on 2026-10-01 reports 74,192 stars, 5,734 forks, 219 watchers, 529 open issues, an 85 MB repository and an Apache-2.0 licence. Nothing in the material accounts for the size of that number. What the material does show is the machinery around it: a documentation site, a Discord, a PyPI package, an npm package that ships only the TypeScript SDK, a model published on HuggingFace, a Docker image, a README carrying a Trendshift badge reading “#1 Repository Of The Day”, a changelog of 533 KB, and a description that leads with two savings figures rather than with what the tool does. The one count that sits oddly beside the stars is the watchers: 219 of them against 5,734 forks.

  2. 02

    One router, one compressor per content type, and savings figures with no method attached

    Three entry points feed one pipeline: compress() for Python and TypeScript, a FastAPI proxy whose per-provider handlers each run the same pipeline before forwarding, and framework adapters for LangChain, the Vercel AI SDK, Agno, Strands, LiteLLM and MCP adapters. A short ordered pipeline runs on every request — an opt-in tool-result interceptor, then a CacheAligner that only reports prefix drift, then the ContentRouter, which does essentially all of the work. The router detects the type of each block and dispatches it to exactly one compressor, and the project’s own architecture document prints a “typical savings” figure beside each: SmartCrusher for JSON arrays at 70–90%, search results at 80–95%, build and test logs at 85–95%, diffs at 40–80%, HTML at 70–90%, tables at 60–90%, structured config at 40–70%, plain text at 30–60%. Those numbers are the project’s own; they are labelled typical rather than measured, and the document supplies no corpus, run or method behind them. Every transform fails open — on an error the content is returned unchanged and the request still goes out.

  3. 03

    The proof table states a method, and the accuracy table carries a caveat the author wrote himself

    The README has a Proof section built from four scenarios, and unlike the table above it says how it was produced: real MCP server output formats, the provider tokenizer, the shipped compress(), seeded and offline, with the command uv run python benchmarks/index_proof_table.py --seed 20260902. Code search 17,199 tokens to 13,597 (21%), an SRE incident 55,957 to 24,340 (57%), codebase exploration 58,801 to 33,895 (42%), GitHub issue triage 46,067 to 32,429 (30%) — every one of those figures is the project’s own. The same section says savings scale with repetition, that repeated JSON arrays and log lines clear 90% in benchmarks/bench_latency.py, that prose and already-dense output compress very little, and that compression costs 0.21 ms at the median on a 10K-token JSON search result. Accuracy comes from python -m headroom.evals suite --tier 1 at N=100 per benchmark, and the section then argues against its own table: “At N=100 a delta of ±0.03 falls inside the confidence interval, so TruthfulQA shows no detectable difference rather than an improvement.”

  4. 04

    What the issue tracker records about fidelity

    The compression has documented failures, and the clearest is #3880, headed “Severity: silent data loss”: Kompress deleted whole records from single-line JSON tool output and left valid JSON behind, “so no parser, log line or model can tell”. The reporter describes the consequence in his own deployment — an agent repeatedly and confidently told a user that a data collection did not exist, and it did — and names the only mitigation then available, HEADROOM_DISABLE_KOMPRESS=1, which also switches off prose compression. The fix keeps record-bearing JSON away from the prose model on every route that reaches it, and along the way replaces a scanner that walked the same bytes repeatedly with a single linear walk: on a 58.8 KB brace-heavy block, 3.43 s became 0.011 s, checked against the previous walk on 4,000 random strings. Two smaller reports are the same theme: #3881 found one code path comparing a word count against a token count, so an unchanged code block looked like a 55% saving, and #3893 found a passthrough from an external compressor being adopted as a result, which let the block go out uncompressed.

  5. 05

    The author’s own audit: the wrong mental model, and about 25,000 lines to delete

    A directory called REALIGNMENT holds fourteen documents and 193 KB of the project’s own diagnosis, and its executive summary opens with an indictment: Headroom “is built on the wrong mental model”, namely that compression means choosing what to drop from conversation history. The IntelligentContextManager tokenised the whole message array, scored each message and removed old ones until the budget was met, and it had been wired into the Rust proxy with frozen_message_count: 0 hardcoded, “so every compression event drops messages from index 0, busting the Anthropic prompt cache for every customer that triggers it”. The audit lists five cache-killer bugs that follow from that model, about 10,000 lines of architectural over-build, a Bedrock and Vertex parity it calls fake, CCR markers computed but never injected into the outgoing body, and X-Headroom-* headers leaking upstream. The plan is nine phases, forty pull requests, roughly thirteen weeks. The same material records the outcome: the context-manager stage was removed, and the scoring classes are gone from the Python package, though the TypeScript SDK still exports the field names “wired to nothing”.

  6. 06

    Release cadence, the review machine, and two languages kept in step by recorded fixtures

    The twenty releases in the sample run from v0.26.0 on 2026-06-16 to v0.39.1 on 2026-09-26, and one week in August shows the rhythm: 0.36.0 on the 20th, then 0.36.1, 0.36.2 and 0.36.3 on the 21st, 0.36.4 and 0.36.5 on the 22nd. Commits by month run 135, 68, 165, 676, 334, 433, 581, 324, 276 and 8. Around that sits a review machine: twenty-three workflow files including a 28 KB ci.yml, a 50 KB release.yml, plus eval.yml, rust.yml, security.yml and pr-health.yml, with scripts for changelog generation, version syncing and pull-request governance. A bot comments on pull requests that leave required sections empty; a maintainer reviews at an exact head commit, reports how many focused tests passed locally next to ruff, mypy and the diff check, and states that he has authorised the maintainer-gated hosted workflows. The two languages are kept in step by recorded fixtures — 231 files under tests/parity/, covering the cache aligner, CCR, code-aware compression, content detection, diffs, Kompress, log compression, SmartCrusher, text crushing, the tokenizer and the Codex and OpenAI contract — and the co-author trailers tell their own story: of 1,168, about 466 name a Claude model and 129 say Copilot.

Adjacent records

All records →