Nexus Agents
A control plane that sits above coding agents rather than being one: it admits work through a single entry point, puts every real fork to a multi-agent vote, records every action in a hash-chained audit log, and requires the repository owner to ratify any change to the rules that govern it.

What it is
A governance substrate for coding agents. It exposes one MCP server with 47 tools; behind it, a scheduler picks a strategy for a task, admission gates decide what is allowed to ship, a default seven-role panel votes on real forks through consensus_vote, and every tool call, voter decision and routing choice is written to a hash-chained append-only audit log that verify_audit_chain can walk. Outcomes feed back into routing through an outcome store, with a second loop that can apply a capped, decaying, demotion-only adjustment to how a CLI is scored. Claude Code, Codex, Gemini and OpenCode are the data plane — the agents that do the work; Nexus Agents is the layer that reviews, audits and routes it.
Who built itA single GitHub account, opened in 2015, with 84 public repositories and 35 followers. He wrote 4,147 of the repository’s 5,184 commits; the other 1,037 are from two bots, one of which also files pull requests on a weekly schedule. Every path in CODEOWNERS names him, including the governor section that requires owner ratification — so the gate separates agent-authored work from owner approval, not one person’s work from another’s. He opens the pull requests, and posts the adversarial review of his own change at a named commit hash underneath them.
How it is put together
The parts · 6One MCP server in front, four coding CLIs behind it, and a control plane in between whose parts all report to the same audit chain. The design decision that explains the rest is a separation of planes: the agents that write code are the data plane and are interchangeable, while admission — what is allowed to ship — is the control plane and is the product. A task enters through a single entry point, a router picks a strategy, gates decide whether the result may proceed, a panel votes when the result is a real fork, and the outcome is written down; the written record then feeds back into how the next task is routed. Around that loop sits a second, narrower one for the substrate itself: the paths that implement governance are enumerated in CODEOWNERS between two directives, and both the review gate and the ratification gate read that same parse, so a change to the machinery of review cannot be landed by the machinery of review. The project is explicit about the ceiling — “autonomic” means self-managing within bounds, every loop’s authority is capped by an authority ladder, and promotion to more authority is earned against evidence and human ratification rather than enabled by default.
- packages/nexus-agents/src/
- 3,074 files.
mcp/is the external interface (47 tools, resources, middleware);consensus/holds the voting engine, the voter roles and the decision computation;pipeline/holds the task contract, the runner, the event bus and the policy engine;audit/is the hash chain;governance/is the fitness audit, drift detection and registry coverage gates;learning/holds the outcome store and strategy distillation;cli-adapters/wraps the external CLIs as subprocesses;adapters/talks to model APIs directly;security/holds the hostile-input firewall and trust tiers;replay/re-executes traces deterministically;observability/carries swarm metrics;dogfooding/is the self-referential review tooling. - governance/
- The durable record rather than the code:
vote-records.jsonl(41 ratification records, 795 KB), the claims registry, the allowed signers a vote signature is verified against, andrequired-jobs.json— the manifest that pins which CI jobs must be wired inside the governor gate, checked by a script that is itself inside the gate. - .rules/ and scripts/
- Twenty rule files loaded by the agents and enforced by CI, plus 216 scripts — the two governor gates, the ledger evidence, append-only and signature-policy checkers, the required-jobs checker, and
inject-governance.ts, which generates a block ofCLAUDE.mdout ofAGENTS.mdand is a governor path because it decides what the agents read. - AGENTS.md, CLAUDE.md and the federation
AGENTS.mdis 84,682 characters and is the single canonical surface for agent guidance;CLAUDE.mdis 89,570 characters and carries a generated block injected from it, gated in CI. Every other harness —.cursor/rules/,.windsurf/rules/,.continue/rules/,.clinerules/,.aider.conf.yml— is a one-line redirect, and pull requests that duplicate content into a harness-specific file are refactored to redirects before merge. The file itself records the limit: content outside the generated block is ungated by construction, “which is how the TypeScript pin drifted”.- skills/ and agents/
- Roughly thirty-five skills as plain instruction files —
self-critique,reviewing-code,research-and-vote,test-driven-development,pre-push-parity,dogfooding-issues,system-review,security-advisory-response, and delegators that hand work to Codex and Gemini — alongside thirteen agent definitions and four under.claude/agents/. The skills are how a discipline written in.rules/becomes a step an agent actually runs. - .github/workflows/
- Twenty-seven workflow files, including the governor review and the scheduled system review that files an issue reporting registry coverage, documentation staleness in days, open and stale issue counts, vulnerabilities, and lint and typecheck status — the same numbers a person would otherwise have to go and collect.
Choices, and what they beat
The governor path set is parsed out of CODEOWNERS, once over a list maintained inside each gate
The instruction files state the mechanism as well as the intent: the set is read from the section bounded by the two directives and both governor gates derive from that single parse, so adding a protected path is an entry in one reviewed file rather than a change to two checkers. The comment above the entries gives the reason the gates are inside their own section — a gate an agent can quietly weaken without tripping it is not a gate.
Ratification fails closed and is bound to a commit over a label, an approval, or a merge that looked right
A governor-path change passes only when an owner approval or an owner-applied label is present and the committed ledger holds a record bound to that pull request and that head SHA at supermajority or unanimous, on a whole panel, under an absolute-quorum error policy. Binding to the SHA is what makes the record about the code that merged rather than about the branch it was proposed on.
The audit chain is opt-in, and the threat model opens by saying so over shipping it on so the guarantee holds out of the box
The document’s section 0 exists because the earlier version described the chain’s guarantees without saying whether a chain was being written; the flag is unset in every shipped configuration and the server warns at startup instead. It also separates a capability from a default state in the instruction files, which had described the log as load-bearing.
Aggregated verdicts must declare what empty means over letting the language default answer it
[].every(p)is true,![].some(p)is true andMath.min(...[])is infinity, so a check over an empty collection reports health. Helpers require awhenEmptyargument, and the rule states that the empty-input test, not the type system and not the lint rule, is what catches the class.Byzantine detection is kept off the live voting path over wiring the weighted detector into consensus_vote
The architecture document says it in both directions — the
WeightedVotingpattern detector and itsbyzantine.*events are exported and implemented, and the note beside the event table states that they are not on the liveconsensus_votepath. The same document counts five distinct strategies behind six names, becausehigher_orderis an alias, rather than reporting six.
Read fromARCHITECTURE.md (18,204 characters) and its module table, AGENTS.md (84,682), CLAUDE.md (89,570), README.md (39,076), the audit hash-chain threat model (43,380), docs/architecture/CONSENSUS_PROTOCOLS.md, CODEOWNERS, .rules/development-disciplines.md, governance/vote-records.jsonl and the directory tree.
Build log
6 stages- 01
One person, two bots, and 6,829 numbered items
The repository was created on 2026-01-03 and had 5,184 commits by 2026-09-28 — 4,147 by one human account, 932 by a workflow bot and 105 by Dependabot, with every commit linked to an account. Alongside them: 3,313 pull requests, 3,516 issues of which 157 are open, 903 releases and 913 tags, 3,897 files, and 500 MB. The commit graph has no idle month: 630 commits in January, 777 in February, 988 in September. Version numbering is not one line but three that do not agree — the oldest tagged release is
v2.2.0(2026-01-16) and the newest isnexus-agents@8.113.0, whileARCHITECTURE.mddeclares itself “Version: 2.137.0, Last Updated 2026-06-21”. The twenty most recent releases all landed on 2026-09-24, between 05:47 and 15:11, and the repository’s own scheduled review bot reports documentation staleness in days, listingARCHITECTURE.mdat 23. Against all of that, the repository has 19 stars, 2 forks and 1 watcher. - 02
It reviews its own work, adversarially, at a named commit
There is no second person, so the second reader is constructed. Open any pull request and the first comment after the body is the author’s own adversarial review of the exact commit he is proposing to merge — “Independent adversarial review of exact head
ce559a93found no substantive blocker” — followed by what he checked, what he could not check, and an explicit statement of the gap the change does not close. A fix to the npm install smoke test (#6826) reads as a list of false passes: a truncatedtools/liststream was counted fromgrepand passed a floor of 25 names; the response had been cut at 65,536 bytes whilejqreported an unfinished string. Holding stdin open returned valid JSON with all 47 tools. Independent review then produced three more ways to exit zero — an emptyinitializeresult,tools/listcarrying both a result and an error, and a healthy page with anextCursor— and each became a red fixture before it became a fix. The same comment names the reviewer’s own limits: “Focused tests 8/8 … await refreshed exact-head CI and review; do not merge before both.” - 03
The governor must not be able to weaken its own governor
The parts of the repository that enforce the rules are themselves gated, and the gate is not a paragraph.
CODEOWNERScarries a block bounded by# @governor-section-startand# @governor-section-endlisting the paths that may never be auto-merged: the audit hash chain, the governance source, the script that injects one instruction file into another, the claims registry, the signing keys a vote record is verified against, and — the entry added for this reason — “the governor’s own gates”, because as the comment puts it, “a gate an agent can quietly weaken without tripping it is not a gate.” Both gates derive their path set from that single parse. Ratification fails closed and is mechanical: an owner approval or an owner-applied label, and a record in the committed ledgergovernance/vote-records.jsonlbound to that pull request and that head SHA, approved at supermajority or unanimous, on a whole panel, under an absolute-quorum error policy. The repository also records where this went wrong:.rules/sat outside those directives until #5997 “while this paragraph claimed otherwise”. The instruction files state the design intent in one line — an agent that can quietly lower its own review bar has no review bar. - 04
The audit chain is off by default, and the threat model says so
The hash chain is the substrate under every “immutable audit” claim in the project, so there is a document whose stated job is to analyse it adversarially. Its headline finding is that the chain is tamper-evident against in-place edits but not against a whole-log rewrite — there is no external anchor, no signing and no monotonic counter — and it tells the reader to read the residual-risk lines rather than the design. Section 0 is the part most projects would omit: in a default installation the chain does not exist. The subsystem is gated on a
security.audit.enabledflag that no shipped configuration sets, so unless an operator opts in, no events are written and every threat in the document is not mitigated but inapplicable.verify_audit_chainagainst such an installation returnsok: truewithnotVerified: 'empty'— honest, as the document notes, but a reader who checks only.oklearns nothing. The server does warn at startup. The document also records that it previously described the chain’s guarantees in full, including aenableHashChaindefault oftrueand the numbered threats, without ever saying whether a chain was being written at all, and that the instruction files compound it by calling the record load-bearing — which describes a capability, not a default state. - 05
What the committed ledger actually contains
governance/vote-records.jsonlis checked into the repository: 795,042 bytes, 41 records, averaging about 19 KB each because each one carries every voter’s reasoning and confidence alongside the tally. Every record was written between 2026-09-14 and 2026-09-24, and every one of the 41 is the same shape — strategysupermajority, decisionapproved, seven approvals, zero rejections, zero abstentions. The first record is the ratification of a fix to the ratification machinery itself (#6241): the ledger gate had judged a bound record on decision, error policy, append-only status and panel coverage but had never readrecord.strategy, so a record produced at simple majority would have counted. The panel that ratified that fix is the standard seven roles — architect, security, devex, ai_ml, pm, catfish, scope_steward — wherecatfishexists to be contrarian; in the committed record it approved at 0.9 confidence along with the rest. A consensus system whose default panel includes a professional dissenter has, in its own durable record of the votes that mattered, no dissent to show. - 06
Rules that are checked, and empty cases that are named
.rules/holds twenty files totalling about 105 KB, and the README says plainly that they are “enforced by CI gates and PR-review voters, not just suggestions”. Drift detection fails the build when documented rules part company with registered behaviour — the model registry, the MCP tool list, expert types, the skills set. The disciplines are specific enough to disagree with: Red/Green TDD requires the first commit on a branch to be a test failing for the right reason; TypeScript is held to a zero-anypolicy by ESLint rather than by review; and one rule targets a class of bug rather than a style — name the empty case.[].every(p),![].some(p),Math.min(...[])anderrors.length === 0all render absence as health, so aggregation goes through helpers whose empty-case argument is required, and the rule adds that the test, not the linter, is what actually catches it. Two sentences in the instruction files are the whole thesis: “A check that cannot fail by construction is not a check”, and “a review must consume the artifact, not a description of it” — with the qualification that a bounded read is legitimate as long as the record says which portion was read. The project reports its own pull-request review experiment as “directional small-n figures, not measured rates”, and notes that human triage of two inspected false positives reclassified one as a real finding the dataset had mislabelled.
Adjacent records
All records →No. 048
gstack
Twenty-three specialist roles and eight power tools for Claude Code, all written as Markdown slash commands, plus the evaluation harness their author uses to decide whether any of it is working.
No. 061
DeepSeek Harness
DeepSeek’s agent harness, built so that the model adapter, the tool registry, the session log and the agent loop itself are plugins — swapped from a configuration file rather than a fork.
No. 060
VibeGame
Describe a game in one sentence and a team of agents divides the work — an architect plans it, a programmer builds it, an auditor checks the code against the plan, and a player has to actually play it before the task is accepted.