gstack
Twenty-three specialist roles and eight power tools for Claude Code, all written as Markdown slash commands, plus the evaluation harness their author uses to decide whether any of it is working.

What it is
A large collection of Claude Code skills that turn one session into something shaped like an engineering organisation: a chief executive who reconsiders the product, an engineering manager who fixes the architecture, a designer who removes AI slop, a reviewer who hunts production bugs, a QA lead who drives a real browser, a security officer who runs OWASP and STRIDE audits, and a release engineer who ships the pull request. Twenty-three specialists and eight power tools, all slash commands, all Markdown, installed by pasting a single instruction into Claude Code. The repository also carries the machinery behind it — an evaluation harness that reports what each run cost, a weekly benchmark lane, and a test suite that was audited down by two hundred and twenty-seven files.
Who built itWrites 370 of the 411 commits. The README introduces him as President and CEO of Y Combinator, and describes a career before it — among the first engineers and designers at Palantir, cofounder of Posterous, and the builder of Bookface, Y Combinator’s internal network. The productivity figures in this record are his own, published with a methodology document and a reproduction script rather than asserted.
Build log
8 stages- 01
The productivity claim arrives with a methodology
The README opens not with what the tool does but with a quote from Andrej Karpathy saying he does not think he has typed a line of code since December, and then with the author’s attempt to answer it. He reports three production services and more than forty shipped features in sixty days, done part time alongside running Y Combinator, and then does the harder thing: he states the metric. Not raw lines, which he agrees artificial intelligence inflates, but logical code change, on which he reports a two thousand and twenty-six rate of about eight hundred and ten times his own two thousand and thirteen pace, eleven thousand four hundred and seventeen logical lines a day against fourteen, measured across forty public and private repositories of his after excluding one demo repository. The claim is linked to a document giving the method, the caveats and a reproduction script, and the README concedes the critics’ real point before answering the rest: that they are right that raw counts inflate, and wrong that a normalised count shows him doing less. Two screenshots sit beside it, one of his contribution graph in two thousand and twenty-six and one from two thousand and thirteen when he built Bookface, with the caption that this is the same person and the difference is the tooling.
- 02
Twenty-three specialists, written as Markdown
The toolkit defines roles rather than utilities. A chief executive who reconsiders the product, an engineering manager who locks architecture, a designer who strips out what the README calls AI slop, a reviewer, a QA lead who drives an actual browser, a security officer running OWASP and STRIDE audits, a release engineer. Twenty-three of those and eight power tools, every one a slash command written in Markdown, under an MIT licence. The install instruction lists about thirty-seven of them by name, and one line in it is worth isolating: it tells Claude Code to use gstack’s own browsing skill for all web browsing and never to use the built-in browser integration. That is a toolkit taking a position on who owns the browser, and it matters because the browsing and QA skills are explicitly designed to drive the sessions the user is already logged into.
- 03
Eighty-eight per cent of the commits name a co-author, and not only one vendor
Three hundred and sixty-one of the 411 commits carry a co-author trailer — eighty-eight per cent, in the range this archive has been recording all round, and the count rather than the ratio is the remarkable part given the repository’s size. The distribution is where it gets interesting for a single developer’s toolchain: Claude Opus 4.6 on a hundred and five, the million-token Opus 4.6 on eighty-five, the million-token Opus 4.7 on sixty, Fable 5 on thirty, the million-token Opus 4.8 on twenty-one, plain Opus 4.7 on eighteen, Sonnet 4.6 and Haiku 4.5 once each, alongside Cursor on six, OpenAI’s Codex on five, and a single Hermes Agent. Human co-authors are named throughout as well. One person’s commit history, then, records a working method that moves between vendors rather than settling on one, and does so visibly enough that the proportions can be counted.
- 04
The bug reports cite file and line
With nine hundred and forty open issues on a hundred and thirty-four thousand stars, the tracker could easily be noise. The ones that matter read like an internal engineering log. One reports that a recent version changed a deployment skill to gate continuous integration on the check marked as required; on a repository that declares no required checks, the tool reads “no required checks reported” as “nothing to wait for”, skips the wait, and can merge over red or still-running work. The report gives the two changed commands, the file and line numbers, the commit that introduced the narrowing, and the observation that the changelog entry did not mention it. Another is a capacity report about a fixed snapshot limit blocking full-repository audits, and it opens by stating that it is not a claim of a security vulnerability and not a request to bypass anything, then pins the version, the commit and the date of the failing run. Precise, bounded, and specific about what is not being claimed.
- 05
The test suite was audited down by two hundred and twenty-seven files
One recent pull request applies a test-audit method borrowed from another project and deletes tests that protect nothing a user sees. The finding is concrete: about four in ten of the free test files never touched product code, because they replayed captured transcripts through evaluation graders that in several cases no longer had a caller. The before-and-after table is in the pull request: tracked test files down from eleven hundred and eighty-four to nine hundred and fifty-seven, test TypeScript lines from two hundred and seventy-four thousand to two hundred and twenty-seven thousand, helper lines from fifty-one thousand to thirty-nine thousand, fixtures from sixteen megabytes to nine. The plan that authorised the deletion was itself run through the toolkit’s own planning command, a chief-executive review followed by a developer-experience review with the design review skipped, and the maintainer approved it unchanged. A tool that deletes its own tests through its own review process is doing something more interesting than accumulating features.
- 06
The continuous integration reports what the run cost
On a merged pull request, the automated evaluation comment reports a hundred and twenty-three automated passes against a hundred and twenty-three final results, none failed, a hundred and seven cases executed and sixteen reused from cache, three cases that needed multiple attempts, the profile and judge counts, a hundred and six behaviours deferred to scheduled coverage with an explicit note that deferred checks earn no credit — and, on its own line, sixty-eight dollars and thirty-two cents of total cost. Pricing an evaluation run per pull request is rare, and it changes what a test is worth: a suite you can see the bill for is a suite somebody can decide to shrink, which is exactly what the test audit then did. The same repository runs a weekly automated benchmark lane and files an issue when its red lane needs triage, which at the time of this record it did.
- 07
The installer is a paragraph you paste into the agent
There is no installation script to run in the usual sense. The README tells you to open Claude Code and paste an instruction: clone this repository shallowly into the skills directory, run the setup, and add a section to your project instructions that lists every available command and redirects all browsing to gstack. The agent installs the toolkit into itself. That is a genuinely different distribution model from a package manager, and it comes with its own prerequisites stated plainly — a JavaScript runtime and its version, Git, and, for the security audit command, a specific runtime build with four compile flags plus a native toolchain on the platform. The README also says what happens when those are missing: setup installs everything else, removes stale helpers, and the security command reports that it did not run and why. Declining loudly is the right behaviour and not the common one.
- 08
Nine hundred and forty open issues and no releases
The version number is at one point nine one point eight, and there is not a single release or tag in the repository: versions advance inside commit messages, so there is no published artifact for a version number to point at. Twelve people have contributed and one of them wrote three hundred and seventy of the 411 commits. Nine hundred and forty issues are open against twenty thousand forks, which is the arithmetic of a project that got much more attention than it can process. And a fair description of the whole repository is that most of its machinery points inward — the evaluation harness, the test audit, the continuous integration gates and the review commands are largely aimed at developing gstack itself. What a reader cannot learn from it is what happens to the projects that install it, which is the question this archive would most like answered and the one a star count cannot answer for anybody.
Adjacent records
All records →No. 035
Nexus Agents
A control plane that sits above coding agents rather than being one: it admits work through a single entry point, puts every real fork to a multi-agent vote, records every action in a hash-chained audit log, and requires the repository owner to ratify any change to the rules that govern it.
No. 061
DeepSeek Harness
DeepSeek’s agent harness, built so that the model adapter, the tool registry, the session log and the agent loop itself are plugins — swapped from a configuration file rather than a fork.
No. 060
VibeGame
Describe a game in one sentence and a team of agents divides the work — an architect plans it, a programmer builds it, an auditor checks the code against the plan, and a player has to actually play it before the task is accepted.