Ponytail
A ruleset and skill pack installed into a coding agent so that it stops at the first of seven rungs before writing code — it does not need to exist, it is already in the codebase, the standard library does it, a native platform feature does it, an installed dependency does it, or it is one line — with six commands for setting the level, reviewing a diff, auditing a repository, collecting deferred shortcuts into a ledger, printing the benchmark scoreboard and listing the commands.

What it is
Rules for a coding agent, installed as a skill or plugin, that change what the agent writes: before any code it stops at the first of seven rungs that holds — it does not need to exist, it is already in the codebase, the standard library does it, a native platform feature does it, an installed dependency does it, it is one line — and only then writes the minimum that works. It ships as six skills and as plugins for the hosts that have a plugin system, with a per-session level of lite, full, ultra or off, default full, also reaching subagents. Installing takes one or two commands: /plugin marketplace add DietrichGebert/ponytail then /plugin install ponytail@ponytail in Claude Code; codex plugin marketplace add DietrichGebert/ponytail then codex plugin add ponytail@ponytail in Codex; node ponytail/scripts/cursor-hooks.js install for Cursor; or one line in opencode.json. Hosts without a plugin system copy one rules file. The default level can be pinned with PONYTAIL_DEFAULT_MODE or ~/.config/ponytail/config.json.
Who built itOf the repository’s 224 commits, 114 carry his own GitHub account and 8 more use a second email of his from another machine; the rest come from about sixty contributors, one of whom has 23 commits under the name Emeriko.
How it is put together
The parts · 6One document of rules, copied into whatever format each host reads, plus a thin layer of hooks that decides when to inject it. Nothing here is a service and nothing calls a model: the ruleset is text, the injection is a per-host lifecycle hook, and the state is a small file under the user’s home directory. That shape explains almost everything in the tree — the same rules checked in as eight files, six skills written once and rendered into a second copy for OpenClaw, one hook manifest per host, a test file per adapter, and scripts whose only job is to keep the copies identical and the versions consistent. It also explains where the project breaks: when a host changes its plugin API, as OpenCode 2 did, there is no server to update and the fix has to be made once per adapter.
- AGENTS.md and the host rule copies
- One ruleset, eight checked-in copies:
.agents/rules/ponytail.md,.clinerules/ponytail.md,.cursor/rules/ponytail.mdc(2,620 bytes),.github/copilot-instructions.md,.kiro/steering/ponytail.md(2,560),.qoder/rules/ponytail.md,.windsurf/rules/ponytail.mdandAGENTS.md(2,593 bytes), most of them 2,495 bytes, withscripts/check-rule-copies.js(2,981 bytes) to check they still match after an edit. - skills/ and .openclaw/skills/
- Six skills, one
SKILL.mdeach: ponytail at 6,637 bytes, help 2,796, review 2,383, gain 1,973, debt 1,703, audit 1,652 — and a generated second copy for OpenClaw (its ponytail skill is 5,957 bytes) built byscripts/build-openclaw-skills.js(2,873 bytes), which a test fails on when it goes stale. - hooks/
- Twelve files and 34 KB:
ponytail-activate.js4,601,ponytail-config.js5,881,ponytail-instructions.js5,487,ponytail-mode-tracker.js6,194,ponytail-runtime.js5,682,ponytail-subagent.js2,683, two statusline scripts (.sh696,.ps1885) and four hook manifests, one per host. - .opencode/, commands/ and pi-extension/
- Six commands written out twice — six
.opencode/command/*.mdfiles (510 to 1,141 bytes) and sixcommands/*.tomlfiles (515 to 1,152 bytes)..opencode/plugins/ponytail.mjs(3,773 bytes) is the V1 entrypoint that OpenCode 2 refuses to load, next toponytail-frontmatter.cjs(1,054);pi-extension/index.jsis 7,171 bytes. - ponytail-mcp/ and benchmarks/
- The MCP server named in the v4.8.0 title:
index.js1,922,instructions.js1,181,README.md1,507,package.json365. The measurement side is 17 files underbenchmarks/, nine results underbenchmarks/results/, and five files and 103 KB underbenchmarks/agentic/(tasks.py50,646,run.py26,677,judge.py9,987,complete.py7,944). - tests/, scripts/ and assets/
- Sixteen test files and 84 KB, one per adapter —
hooks.test.js21,189,cursor-hooks.test.js15,639,hermes-plugin.test.js9,613,correctness.test.js6,382.scripts/is six files and 20 KB: publish the OpenClaw skills, check the rule copies, check the versions, install and uninstall the Cursor hooks, uninstall.assets/is 14 files and 1,127 KB,logo.pngalone 567,655 bytes — the weight #939 took out of the npm package.
Choices, and what they beat
Keep a safety floor and never cut validation, error handling, security or accessibility over a bare “write one-liners” prompt
The README states the rule and then measures it: “The rule was never ‘fewest tokens.’ It is: write only what the task needs, and never cut validation, error handling, security, or accessibility. The code ends up small because it is necessary, not golfed.” In the benchmark table the “YAGNI + one-liners” arm comes out at 95% on the adversarial safety tier, while this one stays at 100%.
Fix the shared function once, at the root cause over patching the path the report names
Written into
AGENTS.mdas a rule rather than left to judgement: “a report names a symptom. Grep every caller of the function you touch and fix the shared function once — one guard there is a smaller diff than one per caller, and patching only the path the ticket names leaves a sibling caller still broken.” The same file concedes the cost of getting it wrong: “The smallest change in the wrong place isn’t lazy, it’s a second bug.”Cut the corner and label it with a
ponytail:comment naming the ceiling over doing it properly nowFrom
AGENTS.md: “Mark deliberate simplifications that cut a real corner with a known ceiling (global lock, O(n²) scan, naive heuristic) with aponytail:comment naming the ceiling and upgrade path.” Theponytail-debtskill exists to collect those markers into a ledger, “so ‘later’ doesn’t become ‘never’”.The Gemini adapter ships no root hooks/hooks.json over sending the same lifecycle hook file to every host
Stated in the README as intentional: “Gemini auto-loads that path, while Ponytail’s lifecycle hooks use Claude/Codex event names.” The adapter installs an extension instead and lets the skills ship separately.
Grok Build gets no lifecycle hooks over wiring it up like the other plugin hosts
The README gives the reason the mechanism fails rather than a preference: “Grok lifecycle hooks are not used because their SessionStart output cannot inject instructions.” Grok users enable the plugin and invoke the skills by name instead.
Read fromAGENTS.md (2,590 characters printed in full; 2,593 bytes on disk), README.md (23,529 characters, of which the report printed the first 6,000 and the remainder was read from the raw file on the default branch), the complete 166-file tree with sizes, the two-level directory summary, the 16 release titles, and the thirty issues and pull requests with their comment threads. docs/cursor-hooks.md (13,264 bytes), docs/platform-native.md (9,434) and docs/agent-portability.md (6,878) appear in the tree with their sizes but their contents were not read, and the 1,127 KB of assets/ was recorded by name and size only.
Build log
6 stages- 01
Seventeen days of releases, then thirty-nine days of silence
The repository was created on 2026-06-12; its first commit is
Initial commitat 00:52:37 and the last ischore: release v4.10.0 (#870)at 2026-09-14T14:34:42Z. That is 224 commits in 94 days, piled into the first two months: 155 in June, 51 in July, 4 in August, 14 in September. The 16 release titles carry the story.v1.0.0 — He ships.went out on 2026-06-12T02:43:54Z, under two hours after the first commit;v4.0.0: production grade, still lazyfollowed the same day, thenv4.1.0: three more agentsandv4.2.0: lazy in OpenCode nowon 2026-06-13, and on 2026-06-15v4.3.0: more agents, still lazy,v4.4.0: field-tested, still lazy,v4.5.0: lazy in Copilotand, four hours later,v4.6.0: help, reluctantly.v4.7.0: lazy in OpenClaw nowarrived on 2026-06-16; the four releases of 2026-06-23 and 2026-06-24 arev4.8.0: comprehension first, now with an MCP server,v4.8.1: consistent versioning,v4.8.2: now on npmandv4.8.3: lazy in subagents too; andv4.8.4: lazy in Hermes nowclosed the run on 2026-06-29 — 14 of the 16 releases inside seventeen days. Thirty-nine days then pass beforev4.9.0: 53 commits of doing lesson 2026-08-07 andv4.10.0on 2026-09-14. Around it: 148,962 stars, 8,009 forks, 356 watchers, 319 open issues against those 224 commits, about one commit per 665 stars, in a 2,666 KB JavaScript repository. - 02
The product is one document and six commands
What gets installed is small enough to quote.
AGENTS.md(2,593 bytes) opens: “You are a lazy senior developer. Lazy means efficient, not careless. The best code is the code never written.” Then seven rungs, each a question: does this need to exist at all (YAGNI); is it already in this codebase; does the standard library do it; does a native platform feature cover it; does an installed dependency solve it; can it be one line; only then, the minimum that works. The ladder runs “after you understand the problem, not instead of it”. The same file names what is not on the chopping block — input validation at trust boundaries, error handling that prevents data loss, security, accessibility — and requires that non-trivial logic “leaves ONE runnable check behind”, while “trivial one-liners need no test”. Around it sit six skills./ponytailsets the intensity tolite,full,ultraoroff;/ponytail-reviewreads the current diff for over-engineering and returns a delete-list;/ponytail-auditdoes the same for a whole repository;/ponytail-debtharvests the markers left on deliberate shortcuts into a ledger, “so ‘later’ doesn’t become ‘never’”;/ponytail-gainprints the measured scoreboard;/ponytail-helplists the rest. The default isfull, set byPONYTAIL_DEFAULT_MODEordefaultModein~/.config/ponytail/config.json. - 03
Half the commits name a model as co-author
Of the 224 commits, 112 carry a Co-authored-by trailer, and the biggest groups name a model rather than a person: Claude Opus 4.8 on 32 commits, Claude Opus 4.8 (1M context) on 31, Claude Fable 5 on 7, Claude Opus 5 on 3, and a bare Claude on 1. Four name a tool or a bot instead — Cursor, Copilot, Devin and
google-labs-jules[bot], one commit each. People appear in the same trailers: Emeriko on 7, Dietrich Gebert on 5, Admin on 2, then a long tail of ones. The account breakdown carries the rest: 215 of the 224 commits have a linked account, and 114 of them are the owner’s own. By email, 103 commits come fromdietrich.gebert@gmail.comand 8 fromdgebert@Dietrichs-MacBook-Pro.local, a second address of his. By author name there are 70 distinct entries, among them one — Emeriko, with 23 commits — that has no matching row in the by-account table. The contributor list holds 50 names, led by Lakshya77089 with 10, ousamabenyounes with 8, then dhedhialy, hamza-ali-shahjahan and salaamdev with 4 each. What makes this more than trivia is the last line ofAGENTS.md, which turns the ruleset on the repository itself: “(Yes, this file also applies to agents working on the ponytail repo itself. Especially to them.)” - 04
Twenty hosts, one ruleset, eight copies of it
The README gives install steps for Claude Code, Codex, GitHub Copilot CLI, the Pi agent harness, OpenCode, Gemini CLI, Qoder, Antigravity CLI, Hermes Agent, CodeWhale, Swival, Devin CLI, OpenClaw, Grok Build and Cursor, and names instruction-only hosts as well: Windsurf, Cline, Copilot Chat, Aider, Kiro, Zed. A badge at the top of the README says “works with 20 agents”. The tree shows what that surface costs: the same rule text is checked in eight times, once per host format —
.agents/rules/ponytail.md,.clinerules/ponytail.md,.cursor/rules/ponytail.mdc(2,620 bytes),.github/copilot-instructions.md,.kiro/steering/ponytail.md(2,560),.qoder/rules/ponytail.md,.windsurf/rules/ponytail.mdandAGENTS.md, most of them 2,495 bytes apiece — withscripts/check-rule-copies.js(2,981 bytes) andnode scripts/check-versions.jsto keep the copies and the version strings aligned. Hosts with a plugin system get code instead:hooks/holds twelve files and 34 KB, among themponytail-activate.js(4,601),ponytail-config.js(5,881),ponytail-instructions.js(5,487),ponytail-mode-tracker.js(6,194) andponytail-subagent.js(2,683), plus one hook manifest per host. The same six commands exist twice, as six.opencode/command/*.mdfiles (510 to 1,141 bytes) and sixcommands/*.tomlfiles (515 to 1,152 bytes). - 05
The OpenCode 2 port, settled in public
The busiest thread in the report is not about writing less code; it is about one host changing its plugin API. Issue #941, opened 2026-09-26, reports that
@dietrichgebert/ponytail@4.10.0cannot be loaded by OpenCode v2 at all: “Plugin must export a default definition with an id and an effect or setup function. (cause: SchemaError(Expected object at ["default"]))”. The same failure was reproduced on v2.0.18 on Windows 11 with Node v24.19.0, on Linux x64 with Node v22.23.2, and on macOS arm64; #930 had already been opened against v2.0.15 and closed by its own reporter as “filed prematurely”. What follows is a queue of competing fixes: #928, closed as a duplicate of #915 and #907, then #933, #940 and #943. Reviewing them side by side in #943, a commenter lists #907, #940, #933, #864, #729 and #734 as the alternatives and calls #943 the cleanest; the same thread reports that on 2.0.20 the dual{ id, setup, server }export registers 0 commands and 0 skills, while deleting theserverkey from the same file registers six — the log lines readcommands: 2, ponytail: 0andcommands: 8, ponytail: 6. One pull request takes the other direction on size: #939 dropsassets/and the pi-extension tests from the npm package because nothing that runs from it references them, taking the published package from 997 kB to 49 kB and from 54 files to 38. - 06
The benchmark and its own correction
The README claims “~54% less code (up to 94%) · ~20% cheaper · ~27% faster · 100% safe”, measured by a headless Claude Code session editing tiangolo’s full-stack-fastapi-template — a real FastAPI + React repository — and scored on the
git diffit leaves: twelve feature tickets, the same agent with and without the skill, n=4, Haiku 4.5. The table gives ponytail −54% lines of code, −22% tokens, −20% cost and −27% time against a no-skill baseline, at 100% on a separate adversarial safety tier; the terse-prose control, caveman, shows −20% / +7% / +3% / +2% and 100% safe, and a “YAGNI + one-liners” prompt shows −33% / −14% / −21% / −30% and 95% safe. The largest cuts land where an agent over-builds: a date picker going from 404 lines to 23, a color picker from 287 to 23, reaching for a native<input>instead of a component. The README also corrects its own earlier numbers: the older single-shot run, which reported “80-94% less code”, is kept in a collapsed section with the note that issue #126 “fairly pointed out that the bare-model baseline pads its answer with prose and options, so that gap is partly a conversational-baseline artifact”, and the badge says the flat figure is “the per-task ceiling, not the average”. Nine result files underbenchmarks/results/carry the dates, from 2026-06-12 to 2026-06-22, the largest a 12,226-byte agentic write-up.
Adjacent records
All records →No. 061
Reticle
An MCP server and a dev-only SDK that let a coding agent read and drive a running web or desktop app from the inside, then answer with a verdict and the file and line to fix instead of a screenshot.
No. 064
delegate-skills
A skills package in which every coding-agent CLI gets its own delegation skill: the orchestrating agent writes a self-contained brief, a separate CLI edits a real working tree, and the human keeps the review and the commit.
No. 067
tty7
A pure Rust terminal whose shells belong to a background server rather than to the window: quitting the app leaves them running, a restart brings the panes back with their layout and the last of what was on screen, four CLI commands let one coding agent drive another, and a remote workspace is that same server running on another machine.