Procoder
A single Go binary that gives a coding agent the discipline of a senior engineer: a commit gate that treats a check it could not run as a failure, controllers that refuse to call unfinished work finished, and a loop that closes the class of every bug that escapes. It works with more than twenty agents, needs no runtime dependencies, and never edits a file behind the agent’s back.

What it is
A discipline harness for AI coding agents, shipped as one Go binary with no runtime dependencies and no server process. It gives the agent three things: a commit gate it cannot argue past, controllers that refuse to call unfinished work finished, and a loop that turns every escaped bug into a closed class. The gate counts a check that could not run as a failure — a tool that is missing, times out or returns output nobody can parse is reported as not checked rather than clean, and the count appears in the gate’s own summary line. The quality chain puts a refusing controller at every link: a specification interview, a plan, a backlog worked in sprints, and a standalone task list whose close refuses without checked criteria and recorded evidence. Everything the tool owns lives in a single directory inside the repository, and the repository’s own files beat the built-in defaults wholesale. Twenty releases in the thirty-four days between 2026-08-20 and 2026-09-23, on 530 commits and 211 stars, under Apache-2.0.
Who built itA solo maintainer: 514 of the repository’s 530 commits are his, committed as Pascal Watteel from the account piwi3910 and the address pascal@watteel.com. Of the other 16, twelve come from two bots — github-actions[bot] with eleven and dependabot[bot] with one — which is the origin of the figure of about twelve machine-named authors in the recent window; the remaining four are people with a single commit each, RealPLU, Acroaticum, ToberoCat and qinghuanandejiangshi, three of whom commit under a display name that does not match their handle. Eleven co-author trailers were extracted from the history, six of them naming github-actions[bot], and beside the names Claude Opus 5, Codex and Nadim Stub the extractor also returns prose fragments, so that list is not a reliable credit roll.
How it is put together
The parts · 5A command-line instrument that a coding agent calls, with a hook layer that calls it whether the agent wanted it to or not. The binary computes and reports; the agent acts, because an agent that experiences its tools as collaborators uses them while one that gets silently overridden routes around its harness. Everything the tool owns lives under one directory in the repository, and the repository’s own version of every rule beats the built-in default wholesale, so a project can change its thresholds without forking the tool. Two consequences follow. A report that says clean has to mean a check ran, so every domain carries a third state beyond pass and fail and it has to survive into the summary line a human reads. And the gate is one collector called from the command line, the git wrapper and continuous integration, so anything that makes the local verdict differ from the CI verdict is a defect in the shared path. The same binary carries the process layer — specifications, plans, a backlog with sprints, a task list, decision records — which is unusual for a linter and is why this is one program. Around it sit thin adapters: a launcher that fetches and verifies the binary, rule files generated from one master for each host, and host manifests that must all report the same version.
- cmd/procoder/
- The whole engine: fourteen files including a 72,448-byte command surface, the flag parsing, a local socket service, the collect path the gate reports through, and the parity tests that hold the two front ends to the same answers.
- internal/
- Fifty-nine packages and 284 Go files. The domains sit beside each other — gate, format, lint, security, deps, docs, infra, testrun, bench, ciops, maintain, codeindex, copilot — with the process layer (spec, plan, backlog, ask, adr, debt, learn, lessons, principles) and the plumbing (gitx, store, config, hook, portability, releases) in the same tree rather than in separate programs.
- .procoder/
- The repository’s own state and its opinion about how work is recorded: 34 specifications of 341 kilobytes, 7 plans led by a 37,591-byte command API plan, 311 backlog files of 421 kilobytes holding 29 epics, 6 milestones and 25 numbered sprints, 5 decision records, 25 files of pending tasks, a 23,975-byte lessons ledger, a 29,500-byte decisions file awaiting answers, and the configuration the binary reads.
- docs/
- Eighteen files directly in the directory, seventeen of them documents: an 81,671-byte command reference, domain, configuration, portability and lifecycle references, and three pages whose job is to limit claims — honest limits, positioning, and a research page that separates the premises with evidence from the premises without. The site is MkDocs with a custom domain recorded in a 19-byte CNAME file.
- The host surface
- Twelve rule copies under a drift check: ten byte-identical to the root agent file at 14,504 bytes, with Cursor’s at 14,571 and Kiro’s at 14,531. Around them sit six host manifests that carry a version, a command directory of 34 files of which 33 are mirrored into the Kilo and OpenCode trees at identical sizes, one 15,093-byte skill, hook definitions for Claude Code and Copilot, and a launcher written twice — an 8,072-byte shell script and a batch file.
Choices, and what they beat
Treat a check that could not run as a failure over reporting only what the tools were able to find
It is contract two on the architecture page, and it is stated as a rule about honesty rather than about coverage: a tool that is missing, times out or returns unparseable output is not checked, counted by the gate as failing, never collapsed into “no findings”, so that if a report says clean, the check ran. The formatter verdict carries three values rather than two for the same reason.
The binary computes and the agent writes over editing code, files and state directly
Contract one, and the reason is given in terms of what the agent will do next: an agent that experiences its tools as collaborators uses them, while an agent that gets silently overridden routes around its harness — and every change stays reviewable in one place, the agent’s own actions. Templates, rule files for other hosts, specification skeletons and tasks are printed for the agent to write; the exceptions are the tool’s own state and a cache prune that asks twice.
Build the binaries in CI and fetch them on first use over committing them with the plugin
The architecture page records both sides of the trade. Committing them made the install offline and cost 39 megabytes of git history per release; the change accepts a first run that needs the network once, and a hook that cannot fetch warns and lets the session continue instead of failing it. The release job builds all five targets from the tagged commit.
Answer only for repositories that adopted the tool over applying the full rule set to any repository the agent walks into
A decision record of its own: the full set of checks belongs to a repository with a configuration directory or an agent file naming the tool, and in somebody else’s repository only the checks that are true anywhere run — secrets, oversized files, conflict markers, junk, attribution lines — while the checks that read content see only the lines the commit wrote. Every run states which mode it was in.
Make the drift check agree with the check that already had the rule right over a configuration key that switches the host copies off
A root agent file with no host copies made the drift check demand all twelve, blocking every commit in a single-agent repository, and the only escape was deleting that file, which also switched off drift protection for the copies the repository does keep. The two halves were made to agree, and drifted or unreadable copies still block whatever was adopted: a stale rule file is another agent being told something this repository stopped believing, while a file that does not exist tells no agent anything.
Read fromdocs/architecture.md (6,449 bytes, and the source of the three contracts), README.md (10,335 bytes), AGENTS.md (14,504 bytes), docs/honest-limits.md, docs/positioning.md, docs/research.md, docs/commands.md (81,671 bytes), commands/spec.md, the five decision records under .procoder/adr/, .procoder/github/LESSONS.md, the pull request and issue bodies for numbers 278 to 303, and the complete 896-file tree with sizes.
Build log
6 stages- 01
A gate that counts what it could not check
The rule the project states before any feature is that a check which did not run is not a pass. Contract two on the architecture page: a tool that is missing, times out, or returns output nobody can parse yields not checked, counted by the gate as failing and never collapsed into “no findings”. The formatter verdict therefore has three values — clean, unformatted, unchecked — and the gate’s closing line prints them together:
procoder gate: 0 clean, 2 unformatted, 0 unchecked, 1 out of scope, 8 hygiene finding(s) (3 blocking). The rule reaches the smaller surfaces too: a barepackage.jsonwith no lockfile is an explicit unscannable gap, and an unreadable rule copy reports itself as UNREADABLE rather than missing. The gate is one code path —procoder check,procoder gitand CI call the same collector, and a local pass that fails in CI is defined as a bug — and the full set of checks belongs only to a repository that has adopted the tool, with every run stating its mode. It was strict enough to block its own repository’s fix for a dependency advisory: in 3.6.0 it named versions that existed nowhere in the working tree, so the commit that removed them was refused, and the fix stopped the scan at nested checkouts. On its own tree the counter reads 856 clean, 0 unformatted, 0 unchecked, 20 out of scope and 172 hygiene findings. - 02
Where the discipline stops being optional
Running the gate is the agent’s job; two lifecycle hooks are not. A write hook fires on every write and edit, and a session-start hook runs before the agent reads anything — which is why the project budgets their output rather than their honesty. A host inlines only the first two kilobytes of hook output, writing the rest to a file, so the hook keeps its message under that, gives the formatted body a bounded share, and drops what will not fit rather than truncating it: half a finding has a meaning nobody can trust, and half a file may be written back over the whole thing. What was dropped is counted, and
procoder checkshows it. Contract one points the same way: the binary computes and the agent writes, so formatters hand over content to review; the exceptions are its own state files. That promise is also where it hurt itself. The format command printed a header plus content when a file needed changes and only a header when it did not, exit code zero both ways, and the gate’s hint invitedprocoder format f | tail -n +2 > f; the maintainer lost two documentation files that way, both committed at zero bytes. The fix moved the verdict to stderr and put the file’s own bytes on stdout, and closed the quieter version underneath — with no header to strip, the same pipe deleted the file’s first line — by rewording the hint and refusing to print nothing for a non-empty file. - 03
The loop that closes a class of bug
The self-learning loop is the part the project calls its reason to exist. A self-review before a pull request exists — through lenses including an adversarial one and an edge-case one — catches reviewer-class findings early. Anything that still escapes becomes an entry in a lessons ledger, one file in the repository that has grown to 23,975 bytes, and writing the entry down does not close it: the adaptation it names, whether a linter rule, a rubric line or a test, has to land before the work counts as done. The bot reviewer is the fallback net, not the net, and what it catches is not lost: one command collects GitHub Copilot’s automatic review findings, strips every trace of the repository’s code out of them, files them as issues only after a terminal confirmation, and marks them unlearned until somebody writes the adaptation that closes the class. A sprint file in the project’s own backlog names the campaign that tests these instruments as one deliberate defect per class the tool claims to catch; its report runs to 17,554 bytes, six scripts drive it, and the architecture page states the rule — a checker that cannot catch its planted bad fixture is not trusted. The loop records its own holes: one pending task file says the records it writes were not actually bounded, and the no-silent-green rule has its own specification, its own backlog epic and two test files named after it.
- 04
One binary, no runtime dependencies, and the bill for it
The implementation is one Go program: 284 files across 59 internal packages, a 72,448-byte main command, and a launcher that finds it. Cross-compilation happens in CI at each tag, the release publishes them with a checksum manifest, and the launcher fetches the one binary the machine needs on first use, verifies it against the published digest and caches it beside the plugin, so later runs exec it directly with no network at hook time. That arrangement is itself a recorded trade: the binaries used to be committed, which made the install offline and cost 39 megabytes of git history per release, and the decision that changed it accepts one network fetch on first run plus a hook that warns and lets the session continue. The README has not caught up with its architecture page — it still says the binaries are cross-compiled and committed with the plugin — the kind of drift its own mirror check exists to catch. What the single binary buys is legibility: the same repository carries rule files for around fifteen hosts, six plugin manifests that carry a version, and no server process. What it costs shows in the small print: two pull requests both record that standalone JavaScript was not run for lack of a test script, and an open bug reports the self-upgrade command on 3.6.0 asking to go to 3.7.0, failing to fetch the checksum manifest on Ubuntu 22.04, leaving the upgrade undone.
- 05
Twenty releases in thirty-four days
The first release is v1.0.0 on 2026-08-20, four days after the repository was created, and its title in the release list is a sentence rather than a number: 1.0.0 — the interface is a promise now. Nineteen more follow in thirty-four days, and the first week is the interesting part of the rate: four on the day of the first one, four the next day, three inside four hours on 2026-08-24, then v3.0.0 on 2026-08-25, after a v2.0.0 the previous afternoon. Three major versions arrived inside six days, and after that the rhythm settles into roughly weekly — v3.5.0 on 2026-09-01, v3.6.0 on 2026-09-09, v3.7.0 on 2026-09-23 — while the commit count tells the other half of the story, 499 of the 530 commits arriving in August and 31 in September. Behind the tags sits a controller that refuses instead of tagging: it checks the version across nine files the repository lists, the changelog entry, a clean tree, a clean gate and a passing suite, prints every failure in one list, and on success prints the
git tagcommand for a human to run. It never tags anything itself, the whole procedure is written down in a release document where that tag command is step seven, and the release pull request for 3.7.0 shows the checking habit reaching into the changelog: it lists the nine files, the new entry, and a correction to the previous version’s heading, which had the wrong date. - 06
One outside fix, a queue of bot bumps, and two misleading failures
The outside traffic is thin: 22 open issues at the snapshot, and one outside contributor, whose OpenTofu fix for the infrastructure gate the maintainer rebased, extended with two commits and landed with the contributor’s name kept on it — noting that the extension was needed because with both binaries installed the same directory was still being sent to the wrong one, the original bug the other way round. Most of the rest is machine-opened: bump pull requests from a workflow token, each carrying the caveat that it does not start CI and has to be closed and reopened, and each answered with a version-specific verification: a checksum matched, the tool installed in isolation, the real command run. Two of the project’s own bugs are worth more than their fixes. A guard test failed once in CI and passed on a re-run of the same commit; the maintainer filed it rather than let a green re-roll close the question, and it closed only when the runner was changed to read structured test output, so a printed failure line can no longer be counted as a failing test. The other is a Windows job that failed on the v3.6.0 tag because two callers appeared to win a start race — his verdict was that the lock was right and the test was wrong, since each winner released inside its own goroutine and Windows simply spread ten of them past the five-millisecond window the other platforms had fitted inside.
Adjacent records
All records →No. 067
Reticle
An MCP server and a dev-only SDK that let a coding agent read and drive a running web or desktop app from the inside, then answer with a verdict and the file and line to fix instead of a screenshot.
No. 070
delegate-skills
A skills package in which every coding-agent CLI gets its own delegation skill: the orchestrating agent writes a self-contained brief, a separate CLI edits a real working tree, and the human keeps the review and the commit.
No. 065
GSD Core
Git. Ship. Done. — a meta-prompting, context-engineering and spec-driven development framework that runs the same five-step loop on every milestone: discuss, plan, execute, verify and ship. The heavy work is pushed into fresh-context subagents so the main session stays lean, and every decision is written into Markdown and JSON under a planning directory instead of living in the conversation.