GSD Core
Git. Ship. Done. — a meta-prompting, context-engineering and spec-driven development framework that runs the same five-step loop on every milestone: discuss, plan, execute, verify and ship. The heavy work is pushed into fresh-context subagents so the main session stays lean, and every decision is written into Markdown and JSON under a planning directory instead of living in the conversation.

What it is
A framework that sits between a person and whichever coding agent they run, and drives it through a disciplined loop instead of a chat. Each milestone repeats five steps, one phase at a time: discuss the implementation decisions before anything is planned, plan and have the plan checked against the context it will run in, execute the plans in parallel waves so that every executor starts from a clean window, verify by walking through what was built, then ship. The problem it is aimed at is context rot — the quality loss that accumulates as an agent fills its window — so research, planning and execution happen in fresh subagents and the main session is kept thin. All state is file-based under a planning directory, as human-readable Markdown and JSON: project brief, requirements, roadmap, state and configuration, with no database and no server, so it survives a context reset and can be committed. Installation is a single command, after which the installer asks which runtime to target and whether to install globally or into one project.
Who built itAn organization account rather than a person, which is why the attribution splits so unevenly. The repository has 6,046 commits and 50 listed contributors: trek-e, whose commits carry the name Tom Boucher and the address trekkie@nomorestars.com, holds 3,742 of them; glittercowboy holds 945; davesienkowski 133, Tibsfox 127, jeremymcs 122, github-actions[bot] 95, 0xdhx 82. 5,829 of the 6,046 commits are linked to an account and 217 are not. Of the 3,143 co-author trailers, the largest single entry is not a model at all — it is sim, on 540 — while 2,309 of them, spread across 19 differently named Claude models, credit one, the biggest being Claude Sonnet 4.6 on 465.
How it is put together
The parts · 6Four layers between the user and the agent runtime: prompt files that name the commands, workflow files that orchestrate, agent definitions that each receive a fresh context, and a command-line tool that owns state — all of it over a planning directory of Markdown and JSON. Because the product is prompt text and the code that emits it, most of the architecture is about making prose bounded and testable: one owner for each vocabulary, a byte ceiling for each file, lint guards that fail the suite when a forbidden write or a retired runtime name comes back, and ratchets whose limits may only tighten. State is deliberately not a service. It is files, which is why a context reset cannot lose it, why git can carry it, and why the same framework can be installed into a dozen different runtimes instead of being rewritten for one of them.
- commands/gsd/ and skills/
- Seventy-two command files, 185 KB, each frontmatter plus a prompt body, converted by the installer into whatever shape the target runtime accepts — a slash command, a skill, an agent file. Six
ns-*.mdnamespace routers sit above the concrete commands, and a parallelskills/directory carries aSKILL.mdfor each of the 72. - gsd-core/workflows/
- One hundred and eighty-two files, 2,357 KB, 89 of them top-level. The three heaviest are plan-phase at 96,328 bytes, execute-phase at 87,727 and review at 51,996, and the byte budget is visible in the layout: the workflows that outgrew a tier grew
modes/,steps/,detail/andtemplates/subdirectories that are read only when the step needs them. - agents/ and gsd-core/references/
- Sixty-four agent definitions, 1,025 KB, 29 of which have a
.compact.mdtwin, over a references directory of 132 files, 738 KB. The references exist to be read lazily by name; an executor exists to start empty, and the definitions say what it may do rather than what it should know. - gsd-core/bin/ and hooks/
gsd-tools.cjsat 290,161 bytes over 24 files, with vendored parsers —re2jsat 252,184 bytes andjs-yamlat 131,632 — and shared JSON manifests for the model catalog, the configuration schema and the exit codes. Forty-two hook files, twelve of them underhooks/lib/, covering a status line, context warnings, secret and injection guards, and the Cursor, Windsurf and shell runtimes.- src/, tests/ and scripts/
- Two hundred and fifteen compiled
.ctsmodules, 7,228 KB in total, of whichstate.ctsalone is 376,952 bytes andphase.cts267,763; 1,247 test files; and 143 scripts, 2,195 KB, that hold the guards, the generators and the release tooling. The installer is a single 714,827-byte file. - docs/ and the planning record
- Ninety-eight decision records, 1,827 KB, with a 34 KB index; 183 feature notes; contributor, branching, versioning and testing guides; twenty written refusals in
.out-of-scope/; 464 archived changesets; and the whole documentation set translated underdocs/ja-JP/,docs/ko-KR/,docs/pt-BR/anddocs/zh-CN/. The 922 KB changelog is the release history written back out.
Choices, and what they beat
Measure the workflow budget in bytes over counting lines, as the agent budget convention did
Written into the architecture document: line count over-penalizes prose and under-catches token-dense tables and code blocks, while bytes are deterministic and match the unit the vendors themselves bound on. The size is taken from one vendor’s truncation limit as a unit but deliberately not as a number, because these are orchestrators loaded by an agent, not instruction documents pasted from a file.
Count extraction as a saving only when the extracted file is read lazily over treating a smaller measured file as a smaller context
The stated reason is that moving prose into a file that is still eagerly imported shrinks the measured file without shrinking the loaded context, which games the proxy rather than serving the goal. The MVP bodies of the planner and executor are therefore referenced by name and read only on the paths that need them.
Nested skills where the loader is non-recursive, a flat listing where it is not over one skill layout for every runtime
The two-stage layout is realized only on runtimes whose loaders do not recurse; Claude was reverted to flat because its skill tool hard-errors on an unknown name instead of routing through the router, and Antigravity was moved from nested to flat because it scans only the top level of the skills directory, which had made the nested sub-skills unreachable.
Ceilings that may only decrease over fixed limits that can be relaxed when a workflow grows
Under the tighten-only ratchet each ceiling tracks its tier’s current high-water mark inside a small grace band, so budgets can be reduced but not quietly raised. The same shape recurs across the repository as committed ceiling files for test counts, lint exceptions and mutation scores.
Let what is on disk decide, and report the disagreement over letting a ticked checkbox in the roadmap decide
The maintainer’s ruling on a reported mismatch kept the selectors disk-authoritative and added an always-present conflict field naming every phase where the roadmap checkbox and the phase directory disagree, together with its plan and summary counts. Seeing the disagreement was preferred to silently resolving it, and the one behaviour change inside a byte-for-byte refactor shipped with a documented changeset.
Read fromdocs/ARCHITECTURE.md (82,825 characters), including its workflow byte-budget and skill-routing sections; the complete 3,630-file tree with sizes and the two-level directory summary; the 99 files under docs/adr/; the 27 live and 464 archived files under .changeset/; the twenty refusals in .out-of-scope/; the 31 workflows under .github/workflows/; README.md; and the thirty newest issues and pull requests read in full.
Build log
6 stages- 01
A history older than the repository that holds it
The metadata says the repository was created on 2026-05-22 under the
open-gsdorganization, but the oldest commit it carries is dated 2025-12-14 and is titled “Initial commit: Get Shit Done - meta-prompting system for Claude Code”, so the project began under another name and the history came with it. The rename was not left as a repository setting: it survives in the tree as installer migration003-rename-get-shit-done-to-gsd-core.cts, an archived changeset file, anddocs/cleanup-get-shit-done-cc.md— which is what a rename has to become once users already have files installed on their disks. The default branch isnext, notmain. In ten months the repository took 6,046 commits: 101 in December 2025, then never fewer than 276 in a month, peaking at 1,028 in the month it was created. Around that sit 10,038 stars, 719 forks, 43 watchers and 186 open issues, on a tree of 3,630 files and 76 MB. The newest commit, on the last day of the record, is a fix that labels a user’s typed arguments in every argument-taking command template. - 02
The process is itself the largest artifact in the repository
The slogan is Git. Ship. Done., and the project runs that loop on itself, in public. There are 98 architecture decision records under
docs/adr/, each numbered after the issue that produced it, from a 1,405-byte note to a 126,874-byte design for one owner of the planning semantic model; the rest of the long ones cover enforcement by construction, an embeddable orchestration engine, and one owner per workflow verdict. The index above them is generated by a script and gated by an 87 KB test, so a new record that is not indexed fails the build. Above the records sit epics split into numbered phases: epic #5056 is designed by ADR-5057 and delivered as at least seven phases, and each phase arrives as its own issue, its own pull request, and a fail-first test written to fail against the current code and closed before the fix lands. The pull request bodies cite the ADR section and the phase entry by name, and a phase that changes behaviour ships with a changeset entry. The repository also keeps its own output:.gsd/phase/feat-3677-quick-batch-hardening-acceptance/holds a 29 KB design document, a test matrix and 12 KB of acceptance evidence. - 03
Thirty-one workflows, one for each rule that stopped being a convention
.github/workflows/holds 31 workflows, and their names read as a list of decisions about what to stop enforcing by hand.pr-target-validatorgreets a pull request opened againstmainfrom an ordinary fix branch with a comment saying most pull requests should targetnext, then lists the four exceptions: release branches, hotfix branches, critical-fix branches for a production outage, and back-merges.pr-template-formatrequires one of three typed templates and refuses a body without the matching heading; when a contributor’s patch was otherwise sound, the automated objection was answered inside the hour by retargeting the branch and adding the heading, with a note that the patch itself was unchanged.auto-branchcreates the branch and comments the two commands that check it out.require-issue-link,changeset-required,docs-required,duplicate-check,duplicate-sweep,dismiss-unauthorized-pr-approvals,auto-close-unsolicited-prs,branch-cleanup,staleand aversion-gatefill out the set. One of them was itself broken: the Dependabot auto-merge job declared onlypull-requests: write, sogh pr merge --autofailed on every dependency pull request until the job was grantedcontents: writeas well. - 04
The test suite is the biggest directory in the tree
tests/is 1,247 files and about 29 MB, larger than the source it tests, with 997 of those files sitting directly in the directory and the rest split between 146 fixtures, 35 helpers and a 55-file QA set. The runner is a 119 KB script. Beside the ordinary test files there are property tests, nineunit.test.cjsfiles and anadversarial/fixture directory with separate corpora for frontmatter, roadmap, security and TOML parsing. Static enforcement is the second wall: 32 custom ESLint rules live undereslint-rules/, among themno-adhoc-markdown-parsing,no-source-grep,no-tautological-assert,no-magic-sleep-in-tests,no-posix-mode-bit-assertandno-rendered-text-length-assert, rules that exist because each of those mistakes had already been made. A further 55lint-*scripts, 22gen-*generators and sevencheck-*guards sit inscripts/beside their committed allowlist files. Several are ratchets with a committed ceiling file, so the number may only fall: a mutation-score ratchet, an allowlist ratchet, a test-file-count ratchet, and a history file that records every CI job’s wall clock against its timeout and is republished by one rolling pull request. - 05
Budgets measured in bytes, and only ever tightened
The most distinctive engineering here is aimed at prompt text rather than at code. Workflow files are loaded verbatim into the agent’s context every time the matching command runs, so they are held to a byte budget enforced by a test, in tiers of 90,000 bytes for the three top-level orchestrators, 54,000 for large workflows and 38,000 for a focused one, and the ceiling may only fall. The reasoning is written down: bytes rather than lines, because line count over-penalizes prose and under-catches token-dense tables and code blocks; the unit adopted from a vendor’s truncation limit, but not that vendor’s number; and the caveat that extraction only helps when the extracted file is read at the step that needs it, since prose moved into a file that is still eagerly imported shrinks the measurement rather than the context. The same pressure produced
workflows/discuss-phase/, where the parent became a dispatcher and per-flag bodies moved into nine mode files, and.compact.mdvariants for 29 of the agent definitions. A second budget governs skills: a flat listing of 86 skills costs roughly 2,150 tokens on every single turn, so six namespace routers costing about 120 tokens were layered above them, using pipe-separated keyword tags on the strength of published routing research. - 06
What is forbidden, and what the users find anyway
An unusual share of the repository is about prohibition. Twenty files sit in
.out-of-scope/, each a written refusal — other runtimes in the core, a human-readable rendering for plan documents — with a 72 KB record arguing for enforcement by construction rather than by review. The managed hooks are where that lands: eleven registered hooks, including a 53 KB secret-read guard that blocks Read, Grep and Bash from touching secret files and replaced a deny rule the installer used to write, a prompt-injection guard, and guards that scan tool output for injected instructions and refuse edits to unread files. The bugs that reached users are the instructive ones. A cosmetic status line spawnedgit statuson every render without opting out of git’s optional index lock, so a decoration could fail another commit; the fix setsGIT_OPTIONAL_LOCKS=0at both seams that read. Command templates spliced typed arguments into instruction prose, so a flag was read as ordinary text and ignored; the same text went into a double-quoted shell snippet, so an argument containing a command substitution ran in the user’s shell. The community, meanwhile, reads the integration branch source: one contributor filed three bugs within seconds of one another — an include that resolves to prose on one runtime, three assumptions in a bootstrap document that fail on others, and a symlink that aborts a command.
Adjacent records
All records →No. 091
OKF Agent Memory
Keeps what a coding agent learns as plain Markdown inside the repository — an OKF v0.2 knowledge bundle searched in-process by BM25 — so the memory can be diffed and reviewed instead of living in a database.
No. 064
delegate-skills
A skills package in which every coding-agent CLI gets its own delegation skill: the orchestrating agent writes a self-contained brief, a separate CLI edits a real working tree, and the human keeps the review and the commit.
No. 061
Reticle
An MCP server and a dev-only SDK that let a coding agent read and drive a running web or desktop app from the inside, then answer with a verdict and the file and line to fix instead of a screenshot.