Open Steps
Eight skills and two hooks that change what a coding agent says at named moments — a ten-line session report that ends on a verdict, steps a non-engineer can follow, a question that ends on one marked recommendation, a premortem run by a fresh agent, and one standing file for the whole product — and the same folders install on Claude Code, Codex CLI, Cursor CLI and Gemini CLI.

What it is
Eight skills and two hooks for coding agents, written for the person running the development who does not read code. Each skill fires at a named moment: os-done-or-not writes a ten-line session report that ends on a verdict, os-step-by-step numbers the steps a non-engineer has to take, os-ask-simple rewrites a question and ends on one marked recommendation, os-what-could-go-wrong has a fresh agent assume a hard-to-undo decision already failed, os-whats-next merges what is verified and says what is left, os-check-work re-measures another session’s claims at their source, os-say-simple restates any text without dropping bad news, and os-big-picture keeps one standing file of the whole product. Two hooks put the routing table and the last report in front of a new session and ask for a report when real work landed. The same folders install on Claude Code as a plugin and on Codex CLI, Cursor CLI and Gemini CLI by being copied into a shared skills directory. MIT licensed, with an eval harness whose numbers are written by a script rather than typed by hand.
Who built itHe introduces himself as not an engineer but a market-led builder: twenty years of building web and software products, always from the product side, and more than 50 developers at his company today. He wrote 50 of the repository’s 64 commits in its first five weeks, and six other accounts wrote the rest, the largest landing 8 and the next 2. He answers issues himself, reads a contribution by running it in a clean checkout, and has pushed the fix onto a contributor’s branch rather than asking for another round.
How it is put together
The parts · 6The organising idea is that what a coding agent tells a non-engineer is a document with a required shape, and that the shape has to be forced at named moments rather than hoped for in a general instruction. So the product is eight folders, each one a SKILL.md whose description is written as a command and whose body is a list of hard rules, each with one worked example beside it in references/; around them sit two hooks that manufacture the moment — one puts the routing table and the last report in front of a new session, the other compares the session’s repositories against a fingerprint taken at the start and asks for a report when real work landed. The second idea is that measurement and assumption must never share a sentence, and it is applied twice over: to the agent, through the rule that every yes names its proof while anything unchecked says “not checked”, and to the project’s own claims, through a README where every cell of the results table says how it is known and through evals whose numbers are written by a scorer rather than typed. What runs on a machine is plain shell — two short hooks, one shared fingerprint formula, an adapter for the two tools that answer in their own JSON, two small scripts the skills call, and a 39 KB read-only install check — and reports are written outside the repositories so they stay out of commits, survive an uninstall, and are shared by every tool pointing at the one folder.
- skills/
- The product: eight folders, each one
SKILL.mdbetween 5,180 and 7,644 bytes, from a description written as an instruction to the moment, through the steps, to a list of hard rules that the file says are the skill. Seven carry a worked example inreferences/;os-big-picturealso ships a 4,024-byte census script and a rationale file, andos-what-could-go-wrongships a prompt script, a handover prompt of 10,614 bytes and two worked examples, the larger of them 21,995 bytes. - hooks/
- Four scripts and their test suite: a session-start hook that prepends the routing table and the last report, a stop hook that decides whether real work landed, a 6,533-byte adapter that gives Cursor CLI and Gemini CLI the same two behaviours through their own JSON, and a 4,673-byte fingerprint script that both hooks share on purpose so the question “did real work happen” cannot be answered two ways. The suite beside them is 38,760 bytes, larger than everything it tests.
- evals/
- The measurement: a 9,255-byte runner, a 34,718-byte scorer, a 26,529-byte test file that drives both with stand-ins so no model is called, the phrase list, a small model registry that also fixes column order, the generated results files for Claude and for Codex, a second runner per tool, and 36 fixture streams in eight folders that keep a hand-made case for each bug the harness has had.
- docs/ and commands/
- What the instructions do not fit in: a 26,248-byte document on installing and running the pack on the three non-Claude tools, a 3,756-byte argument for the routing block, a 1,339-byte copy of the block itself, and the 3,430-byte page for the optional output style. One command file, 2,572 bytes, runs the install check and reads the result back for someone who will not open a terminal.
- doctor.sh and .claude-plugin/
- A 39,257-byte read-only install check that reports what is wired and says “not checked” for what it cannot look at, next to the plugin and marketplace manifests, the first pinned at version 0.4.6. The single workflow runs both test files on Linux and macOS, validates the plugin manifest, and checks the sign-off line, so a construct that works on one platform fails a check instead of somebody’s terminal.
- The root and its seven READMEs
- A 29,313-character README with Spanish, French, Russian, Ukrainian, Korean and Chinese versions beside it, 6,125 bytes of contribution rules, an MIT licence, two SVG assets that show before and after, and an issue template. The whole tree is 99 files and about 587 KB, of which the evals, the hooks and one skill account for most of the weight.
Choices, and what they beat
Write the skill descriptions as commands rather than summaries over descriptions that describe what the skill does
The README names this as one of three decisions that shape everything: the descriptions open with “ALWAYS invoke this skill…”, and with that form the skills switched on by themselves in the measured runs. It also states the limit of the evidence — the evals never compared this wording against another.
Put anything that must hold in every reply in the tool’s standing instructions file over asking a skill to enforce it
Stated as the pack’s second decision: a skill cannot force itself to run, so a rule that has to hold in every reply belongs in
CLAUDE.md,AGENTS.mdorGEMINI.md, or in a Claude Code output style, and the pack says which layer each piece belongs to. The routing block is the one piece the user adds by hand for exactly this reason.Keep what was measured apart from what was assumed over a confident verdict that reads the same either way
The third decision and the one that shows in every surface: a “yes” has to name its proof, anything unchecked says “not checked”, and that is also the stated reason the reports are short, because the agent stops narrating its checks and states the result. The same rule is applied to the project’s own README, where each results cell carries one of seven words for how it is known.
Leave the output style switched off, and say so over the one-line flag that would apply it to everybody at install time
The style’s own document records that the alternative also overrides the user’s own setting, and gives the reason for refusing it: someone who chose another style, or wrote their own, should not lose it silently by installing an unrelated pack, because a pack whose whole premise is not surprising the person using it does not get to surprise them at install time. A team can add the line, and is told to say that it did.
Write the reports outside the repositories, into one folder every tool shares over saving the report inside the project
The README gives three consequences: the reports stay out of the user’s commits, they survive uninstalling the pack, and pointing every tool at the same folder gives one history instead of four. It also notes that the
.claudein that path is only a name.Change what the agent does at a named moment, not only how its answer is worded over a style layer that reshapes every reply, or a display-only rewriter
This is the line between this pack and the neighbouring plain-language records in this archive, and the maintainer states it himself in the reply to issue 34: changing the agent’s behaviour at defined moments is the product, not a side effect, and the invariant is that only the wording changes while every measured number stays as it was. Attention Span, already in this archive, does the other thing — three output styles that change how the agent talks, benchmarked on the work and the output separately — and claudish-to-english redraws every assistant message on screen while the transcript keeps the original text.
Borrow the ideas of ASD-STE100 without claiming the standard over citing certification the pack does not have
The README and the rewriting skill both name the source of the plain-language rules — the simplified English written for aerospace manuals, short sentences, active voice, one idea per sentence — and both add the same qualifier: borrowing is the word, and nothing here is certified against the standard.
Read fromREADME.md (29,313 characters), docs/output-style.md, .claude-plugin/plugin.json, hooks/hooks.json, skills/os-done-or-not/SKILL.md, skills/os-ask-simple/SKILL.md, skills/os-say-simple/SKILL.md, CONTRIBUTING.md, and the complete 99-file tree with sizes.
Build log
6 stages- 01
Five weeks, eleven releases, and a pack that moves on its own
The repository was created on 2026-08-25T14:24:24Z and
v0.1.0was tagged about a minute later, at 14:25:44. Five weeks on it holds 64 commits, 52 of them in September and 12 in August, the last on 2026-09-30, and 11 releases with 11 matching tags, every one of them a full release and none a prerelease or a draft, so nothing here is a beta channel. The numbering skips an entire minor line:v0.1.0, thenv0.3.1on 2026-08-31,v0.3.2on 2026-09-06,v0.3.3on 2026-09-10,v0.4.0andv0.4.1twelve minutes apart on 2026-09-11,v0.4.2on 2026-09-13,v0.4.3on 2026-09-14,v0.4.4andv0.4.5an hour apart on 2026-09-28, andv0.4.6on 2026-09-30. Around that sit 1,072 stars, 99 forks and 40 watchers, with 4 issues open and 30 issues and pull requests in total. Seven accounts have committed: the author on 50, one contributor on 8, another on 2, and four with one each; every commit carries a linked account, and only 2 carry a co-author trailer, one of them naming Claude Opus 5. The plugin manifest is pinned at version 0.4.6, the tree is 99 files, and the README exists in seven languages. - 02
What an honest report is made of, rule by rule
os-done-or-notis the skill the pack is built around, and itsSKILL.mdstates the shape rather than describing it. Before writing anything the agent picks one of eight named outcomes — shipped and verified, one action left for you, did not work and was rolled back, something broke, research only, and three more — because otherwise the report says “fully done: yes” and “safe to close: no” in the same breath. Then a lead of one or two sentences with the best result first and never the chronology, a checkmark table of two to five rows, and a verdict table of four lines: fully done, anything needed from you, new debt, safe to close. What stops it from whitewashing is a set of hard rules the file says are the skill: every “yes” names its proof, and no proof means “not checked”; bad news gets its own warning row and is never buried inside another line; an unhandled security risk or data loss is the one exception to the ceiling, spelled out in full rather than by inflating a verdict cell. The gotchas argue the same thing from the other side — “new debt? no” gets written both when there is no debt and when nobody looked, so “no” comes only after checking, and a green check or a merge does not mean users have it. The outcomes where something failed or broke, the file says, are where reports start lying, and “the approach failed and was rolled back” is a complete result. - 03
Straight verdicts: the rules that forbid a menu
Where something has to end on a decision, the pack writes the rules that force one.
os-ask-simplefirst asks whether the question should reach the person at all — can the agent answer it by looking, is there a conventional default — then puts it in one plain sentence with why it matters, what changes later and whether it is easy to undo. A structural choice passes six checks shown as a table, and every row is answered: “not checked” is allowed and honest, silence is not, and a row that always answers “no” is, in the file’s words, decoration. The hard rules read like a list of ways to avoid a verdict: always weigh doing nothing and say why it lost; a recommendation is required, because “it depends” is not one; recommend against the user’s own idea when the screen says so; never recommend what you have not screened; and treat “this is the most interesting thing to build” as a reason for suspicion rather than for preference.os-say-simplecarries the matching rules for text, and its first one is the anti-whitewash clause at the sentence level: add nothing and drop no bad news, because a summary that loses the warning line is a lie by omission. Numbers stay exact; the rewrite is as true as its source and no truer, so it says “the report says the tests pass” rather than “the tests pass”; unclear stays unclear; code and commands are exact strings. - 04
Four hosts, three files to edit by hand, and a table of what was actually run
The skills are the same folders everywhere; only the wiring differs. On Claude Code one command installs the plugin;
hooks/hooks.jsonbinds the start hook toSessionStartand the report hook toStop. The other three tools read a shared skills directory, so one copy command installs the skills for all of them, and what stays per tool is a routing block — a short table naming when a skill must be used, placed by hand in~/.codex/AGENTS.md, the project’sAGENTS.mdon Cursor CLI, or~/.gemini/GEMINI.md— plus hook settings in~/.codex/config.tomland in~/.cursor/hooks.jsonor~/.gemini/settings.jsonthroughhooks/adapter.sh. The adapter runs the two hook bodies unchanged and translates only what goes in and out: the handover becomesadditional_context, the report request becomesfollowup_message. On Gemini the stop sits onAfterAgentand goes out as a denial carrying the request as its reason, which Gemini returns as the next prompt; Cursor can only ask, because a stop cannot be blocked there. What makes the README readable is its evidence vocabulary: each cell of the evidence table says whether the behaviour was measured, watched, checked by hand, set up from docs, in place, did not start, or not tried. Claude Code comes out at 85% to 100%, Codex CLI at the right skill read in 75 of 75 runs, and Cursor CLI and Gemini CLI at not measured. - 05
A scorer with no model in it, and the bugs the measuring found in itself
The harness asks 25 phrases a person would say, three per skill plus one that tests the line between two of them, three times each on three Claude models with nobody at the keyboard, and three off-topic questions;
score.pyreads tool calls and prints the results, which the author does not type. Measured on 2026-09-12 the reach was 85% on Haiku 4.5, 98% on Sonnet 5 and 100% on Opus 5, with no off-topic question pulling in a skill, and the write-up keeps the misses in view rather than the score: on Haiku two skills are unreliable, and four rounds of the same phrases put one skill at 50%, 33%, 44% and 44% — “I would rather say that than quote the friendliest round”. The measurements then found faults in the measuring. Headless runs denied every skill call, so the arm meant to measure quality compared the pack against itself: on 2026-08-29, 234 streams made 279 skill calls of which none ran; on 2026-09-01, 261 streams made 362 of which 12 ran. A sweep was also not sealed off from the machine around it — one skill lists live sessions, and three runs pinged the session that launched the sweep while a fourth pinged an unrelated session two hours into somebody else’s project. Both were fixed and re-measured: the skill tool is granted to the two quality arms only, and every run now starts with two cross-session tools denied, which sealed 270 of 270 streams on 2026-09-12. - 06
The community rounds, and one issue closed as not planned
Seven accounts contributed, and the largest run of work came from one of them: eight pull requests that added worked examples for three skills, the adapters that make both hooks work on Cursor CLI and Gemini CLI, and a command that runs an eval day, and a generator that writes the README’s numbers table. Another contributor added a standing map of the product, which the maintainer renamed to
os-big-pictureandBIG-PICTURE.md, cutting one trigger because “status” belongs to a ticket rather than to the project. A third contributed the install check, a script that prints facts and sets an exit code, accepted after two fixes, and the Codex runner arrived as a pull request that closes its issue, with 84 runs attached. Because a first-time contributor’s pull request cannot start the repository’s workflows without an approval its tooling cannot reach, the maintainer opened a mirror pull request at the same commit so the checks would run, then closed it unmerged. One issue was closed as not planned: a Chinese-language request to make term translation a rule that never enters the agent’s decisions, and to give Codex a full plugin like Claude Code’s, answered with the point that changing behaviour at named moments is the product rather than a side effect, and that Codex has no equivalent plugin format, so the install there is a copy of the skill files plus hooks in its config file.
Adjacent records
All records →No. 117
sepia
A portable de-AI writing skill: four operations over one canonical rules file, narrative architecture repaired before word choice on fiction, a thin rule file matched to the venue on professional prose, and every rule labelled as measured, consulted or the project’s own inference.
No. 066
makerskills
Twenty-one agent skills for running a one-person business — decide, unstuck, maker-council, deep-research, second-brain, company-brain, domain, jab-hook, pm, personal-cfo and the rest — each one a Markdown workflow document rather than a program, installed into Claude Code, Codex, Cursor or any host that reads the Agent Skill format, with every piece of personal state kept in a config directory the repository never touches.
No. 129
Agents Universe
An open-source agent platform that keeps one project context shared by everyone working in it. Agents read the whole knowledge base when a project is opened and write what they learn back into the same files while they work; knowledge is Markdown on disk with a database index behind it, and there is no embedding model or vector search.