Autoprompt Skill
An installable workflow for coding agents rather than a standalone program: it takes one goal, plus the constraints and the definition of done, and runs the execution loop — scope, plan, build, test, review, repair, verify — across eleven different agent tools. Coordination, execution and independent judgment are kept in separate layers, so the agent that writes a change is not the one that signs it off. Its headline claim, 45% fewer failures, comes from one Terminal-Bench 2.1 comparison the author ran, whose baseline verdict file is linked but not present in the repository.

What it is
Autoprompt Skill is a workflow you install into a coding agent rather than a program you run: one command hands it a goal together with the constraints and the definition of done, and it drives the execution loop itself — scoping, planning, building, testing, reviewing, repairing and verifying — instead of waiting to be prompted at every step. Eleven agent tools are supported, among them Claude Code, Codex, VS Code, OpenCode, Kilo Code, Prime Agent, Oh My Pi, DeepSeek Harness, Reasonix, Hermes Agent and Grok Build. Work is organised into three paths, direct, light and roadmap, chosen automatically or named before the goal, and carried out by a layered fleet that separates coordination, management, execution and independent judgment, so the author of a change is not its approver. Concurrency is capped, and model choice runs through a registry. It is MIT licensed JavaScript with 1,298 stars, 88 forks, eleven contributors and 1,235 files. The claim of 45% fewer failures rests on a single Terminal-Bench 2.1 run by the author.
Who built itThe owner account, credited with 22 of the repository’s 44 commits — 17 under the 77499101+Spielewoy address and 5 under a second Spielewoy address. Eight commits carry a co-author trailer naming Claude Opus 5, and three are attributed to an account called claude; the rest come from eight other accounts, most of them a single change, and the contributor list has eleven entries in total. The MIT licence is held as “Copyright 2026 Spielewoy”.
How it is put together
The parts · 6The project is a payload plus an installer rather than a service: an npm install, then an interactive installer that detects the chosen agent, writes that host’s native role files, records a receipt, and can be re-run later as a doctor or an uninstall. Everything downstream hangs off one idea — the workflow is the product, and it has to be expressed eleven times, once per host, without the eleven copies drifting apart. So the source of truth is seven JSON contracts (product, routes, state machine, roles, checks, providers, plain language); a generator turns them into each provider’s native projection; a manifest pins what is installable; and a parity test fails when a projection stops matching the contracts. Generation and admission are deliberately separate — opening all eleven projections does not mean a run may start, and each host must present capability evidence taken from a real installation first. The run itself is a deterministic control plane around a hierarchy of roles: an L0 run owner holds the goal and the final response, L1 to L4 carry scope, management, execution and fresh terminal judgments, a controller owns the physical launches and enforces ownership and independent checking, and only the roles the selected route needs are started. Benchmarking sits beside that as its own subsystem with its own schemas, leases, ledger and quality gate, and that gate is allowed to refuse the claim.
- agents/
- Eleven provider packages, 59 to 63 files each, holding the same 32 role profiles in that host’s native format — Markdown, TOML,
.agent.md, permission-scoped JSON or Python personas — plus 18 framework procedures each and aVERSIONfile. A 5 KB README explains the projection rules, andagents/manifests/pins eleven runtime payloads between 8.6 and 39 KB, the Codex one much the largest. - agents/contracts/
- 89 files and 619 KB of authority: the seven source-of-truth contracts, 32 schemas under
schemas/— twelve of them for benchmark evidence, from attempt evidence and mechanism evidence to run manifest, aggregate report, route holdout, pricing and upload spool — 25 persona texts, eight provider contracts, the gate registry and the route and role tables. - scripts/
- 46 files and 1.36 MB in three groups.
scripts/install/is 21 files and 947 KB and exists twice over, once in bash and once in PowerShell:install.ps1107 KB,install.sh74 KB,install-lib.ps1278 KB,doctor.ps1anddoctor.sh25 and 21 KB.scripts/benchmark-evidence/is 25 files and 488 KB. Theharness-v2-*family carries the v2 bridge, withharness-v2-transport.cjsat 124 KB andharness-v2-native.cjsat 104 KB. - agents/codex/workflow/
- 45 files and 3.5 MB, the largest single body of code in the repository: router, scheduler, runtime state, run records, mission lock, recovery checkpoint, request and context envelopes, a route decision module of 126 KB and a
phase-budget.jsof 1,612,186 bytes, plus a Windows AppContainer path written three times over in C#, JavaScript and PowerShell. - tests/
- 180 source tests and 4.9 MB —
codex-supervisor-integration-v2.test.cjsalone is 944,791 bytes, withcodex-runtime-state-v2at 270 KB,codex-live-conformanceat 142 KB and four more above 100 KB — alongside 47 fixtures, 12 helpers, a three-filetests/benchmarks/harness, and aharness-v1fixture set that preserves one role file per provider as the migration reference. - docs/, bin/ and assets/
bin/autoprompt.cjsis 77 KB of interactive installer and CLI with a 22 KB provider-root compatibility module beside it.docs/holds four translations of 11 to 14 KB each, five guides including an 11 KB Codex local-records document and the compatibility guide, seven FAQ pages, three benchmark pages and an 8.6 KB Lima runtime note.assets/is 2.8 MB across six files plus sixteen localised copies of the loop, hierarchy and leaderboard charts.
Choices, and what they beat
No default route over one fixed pipeline for every request
The source README says it plainly: “There is no default route.” DIRECT and LIGHT require no coordinators or managers, and ROADMAP may use the run coordinator and the work-group manager only when the canonical policy admits them. The work-paths page goes further — file count does not decide the path, a mechanical rename across twenty files can be
directwhile a three-file change across separately deployed services can needroadmap, and a failed attempt alone does not justify a larger path.Capability evidence can refuse a run over falling back to an unrestricted native launch
“Missing required capability evidence must block admission without an unrestricted native fallback,” the source README says, and the cost is visible in the tracker: the 2.0.0 bundle ships ten reviewed local records that are all Linux x64 with an empty trusted key list, so a native Windows host is refused with
PROVIDER_UNSUPPORTED. The support table carries the same warning — “These versions passed Linux runs” — and the FAQ says the verified execution path is Linux.The 25 v1 roles survive as read-only aliases over deleting them and shipping only the seven active roles
All eleven harnesses project the same 32 physical roles, described as seven active roles plus 25 inactive compatibility aliases. Leaves and aliases cannot dispatch, and the aliases are read-only redirects that cannot take new v2 work. Authority moves to the new roles while an existing installation keeps resolving its old ones instead of breaking on upgrade.
An explicit path stops instead of falling back over quietly selecting another route when the named one does not fit
From the work-paths page: an explicit path bypasses automatic route selection but “does not remove authorization, capability, budget, or verification requirements”, and invalid or incompatible choices stop with an error instead of silently selecting another path. Silent substitution would leave an unattended run’s plan impossible to audit after the fact.
The Oh My Pi projection omits the task tool over shipping the same tool list to every host
The agents README gives the specific reason: that host’s parser can infer unrestricted spawning from the task tool even when an empty child list is passed, so the projection disables prewalk and advisor handoffs and removes the tool. It is the clearest case in the repository of a host’s parsing behaviour changing the shape of the payload rather than the payload staying identical everywhere.
A canary that is not allowed to claim quality over letting a mechanics run stand in for a benchmark
The three-arm low-compute canary sets
qualityClaimEligibleandcomparisonClaimEligibleto false on every record it emits and states that its results cannot support quality, cost, latency or redesign-effect claims; every attempt is stampedsourceBinding=declared-existing-commit-not-executed; the route fixture is marked synthetic and readiness fails closed until human labels, two raters and a pre-tuning freeze exist. That is exactly the discipline the Terminal-Bench page is missing.
Read fromREADME.md (11,027 characters, fetched in full from the GitHub contents API), agents/README.md, agents/contracts/, docs/faq/work-paths.md, docs/faq/what-are-the-layers-for.md, docs/faq/which-coding-agents-are-supported.md, docs/benchmarks/terminal-bench-2.1.md, docs/benchmarks/codex-low-compute-mechanism.md, tests/benchmarks/autoprompt-benchmark.json, docs/CONTRIBUTORS.md, the pull request and issue bodies, and the complete 1,235-file tree with sizes.
Build log
6 stages- 01
Six releases in twenty-three days, then nineteen days of quiet
The repository was created on 2026-08-17 at 22:40 UTC and the first commit —
Autoprompt Skill 1.0.0— landed two minutes later. Six releases followed in twenty-three days:v1.0.0on 2026-08-18,v1.0.1andv1.0.2both on 2026-08-19,v1.0.3on 2026-08-20,v1.0.4on 2026-08-21, and then a nineteen-day gap beforev2.0.0on 2026-09-09. None of them is a prerelease or a draft. The 44 commits split 29 in August and 15 in September, and the newest one,Preserve community PR authorship for the released v2 integration, is timestamped 2026-09-09 at 22:05, while the repository’s pushed timestamp is 2026-09-28, three weeks later, so some of that later activity is not visible in the commit list. Around the code sit 1,298 stars, 88 forks and 4 watchers, three open issues, an MIT licence, 1,235 files and 375,291 KB of repository. - 02
The 45% number, and where its evidence stops
The claim is on the README’s first line — a coding-agent workflow that cuts failures by 45% by reviewing, fixing and rechecking its work — with a badge reading Terminal-Bench 2.1, plus 14.61 points. OpenCode alone solved 60 of 89 tasks for 67.42%, failing 29; OpenCode with Autoprompt solved 73 of 89 for 82.02%, failing 16 — thirteen more solves and 45% fewer failures. The method page adds the setup — DeepSeek V4 Flash, OpenCode 1.18.7, the same 89 tasks — and notes these are version 1 results, with version 2 benchmarks to follow. The evidence section is where it thins. The baseline is said to have all 89 retained task verdicts at a path under
benchmark/terminal-bench-2.1/; the Autoprompt side is called the recorded completed aggregate in a comparison report whose original per-task map was not retained, so that result cannot be rebuilt task by task. Thatbenchmark/directory is not in the published tree at all: the contents API returns 404 for it. The promised trade-off — three times the time, twice the tokens — is a planning estimate from user experience reports, because the timing and token logs were not retained. DeepSeek’s own 82.7% is called an external reference, not a comparable third run. So: no dataset in the material, no reproduction command, no surviving per-task output, no independent rerun. Read 45% as the author’s own single measurement. - 03
The repository’s own benchmark gate says no
Beside that headline sits a body of work that argues the other way.
scripts/benchmark-evidence/holds 25 files and 488 KB of harness —aggregate.cjsat 53 KB,manifest.cjsat 25 KB,route-holdout.cjsat 22 KB, a spool, a run lease, a trust registry and a quality gate — andagents/contracts/schemas/carries twelve more benchmark schemas covering attempt evidence, mechanism evidence, run manifests, aggregate reports, pricing and result bundles, exercised by an 82 KB source test. What it does with its own results is the interesting part. The three-arm low-compute canary records one task, one repetition,evidenceClass: "harness-mechanics-only", and stamps every record it emits withqualityClaimEligible: falseandcomparisonClaimEligible: false; its document says flatly that this is a mechanics check, not a quality benchmark, and that its results cannot support quality, cost, latency or redesign-effect claims. The real Codex observation is carried over rather than rerun, still blocked withPROVIDER_UNSUPPORTEDforcodex-command-sandbox-network-open, and a real three-arm comparison is still open. The 16-row route fixture is marked synthetic, and readiness fails closed until independent human labels, two raters, inter-rater agreement, adjudication, a development and test split and a pre-tuning freeze exist; its perfect confusion matrix is mechanics evidence only. - 04
Nineteen days of silence, then ten replies in one evening
Between 2026-08-19 and 2026-08-31 the repository collected a steady stream of outside work: Grok Build support (#5), a CRLF shebang and missing exec-bit fix (#7), an Oh My Pi adapter (#8), a supervisor and model-casting addition to it (#11), two independent attempts at the non-blocking stdin crash (#15 and #24), a Python 3 resolution fix for the bash installer (#16), Hermes Agent support (#17), five grammar corrections from one contributor (#18, #19, #21, #22, #23), a report that the DeepSeek preset denied two unregistered tool names (#20), three filings of the same Windows Git Bash path failure (#12, #13, #14) and a complaint that the compatibility guide told readers to replace two bracketed tokens without saying what they stood for (#4). On 2026-08-23 the author answered several of them with the same sentence — “Currently working on some major refactors” — and then went quiet. On 2026-09-09, within minutes of the
v2.0.0release at 21:25, ten of those threads received a reply saying the issue was fixed in v2 and the reporter credited in the README, and a final commit is titledPreserve community PR authorship for the released v2 integration. The credits page names seven contributors with PR numbers and adds a disclaimer that a credit does not mean an old pull request was merged unchanged, or that a provider passed release verification. - 05
Thirty-two roles, three paths, and a gate that fails closed
The v2 rewrite replaced the 25-role fleet with a fixed set of 32 physical roles described as seven active roles plus 25 inactive compatibility aliases, projected into all eleven harnesses along with 18 framework procedures. The source README states the rule that shapes all of it: “There is no default route.” DIRECT and LIGHT need no coordinators or managers, ROADMAP may use the run coordinator and the work-group manager only when the policy admits them, the run owner keeps the goal and the final response, and the controller enforces physical launches, assignment ownership, independent checking and recovery; leaves and aliases cannot dispatch. Responsibility is layered from L0 to L4 so that an author cannot approve its own work, and the work-paths document insists that file count does not decide the route — a twenty-file rename can be
direct, a three-file change across two deployed services can needroadmap. The same discipline applies to admission: “Missing required capability evidence must block admission without an unrestricted native fallback.” The bill arrives in the issue tracker: a Windows user reports every activation refused, because the bundle ships ten reviewed records that are all Linux x64 and the trusted key list is empty, against a support table that says the tested versions passed Linux runs. - 06
Two shell ports and the Windows bill
Maintaining eleven host projections is the visible cost; maintaining two operating-system ports is the quieter one.
scripts/install/exists twice over in 21 files and 947 KB —install.ps1at 107 KB andinstall.shat 74 KB, with shared libraries of 278 KB and 228 KB — and the bugs follow the seam between them.v1.0.2was published from a Windows release runner with no.gitattributes, so the LF-committed launcher was packed with CRLF endings and no executable bit, and every fresh macOS or Linux install failed with eitherpermission denied: autopromptorenv: node : No such file or directory; the fix adds.gitattributesand a packaging test that asserts each declared entrypoint carries an LF shebang, no carriage returns and the executable bit. A later pull request normalises Git Bash short paths through their long Windows form, which it says cleared three lifecycle failures in release-readiness run 32481402055 and unblocked thev1.0.4publication. Reports that Windows hosts resolvebashto a WSL stub arrive three times over the same day, native Windows support is requested and refused twice more, and the still-open pull request that fixes npm.cmdshims notes thatnpm run verifyremains blocked by six pre-existing Windows, macOS, Codex and cache-environment failures in the CLI test suite.
Adjacent records
All records →No. 095
HarnessRouter
The self-hosted, Apache-2.0 edition of HarnessRouter: it puts sixteen existing agent CLIs — Codex, Claude Code, Hermes, DeepSeek Harness and twelve more — behind one OpenAI Responses-compatible API, with sessions, streaming, files, cancellation and structured failures, and it carries the Unified Harness Protocol it implements together with the conformance suite that measures it.
No. 091
Lody
A workspace where a team shares the coding agents it already runs: connect a machine, bring Claude Code, Codex, Kimi or any other agent that speaks the protocol, and dispatch work from desktop, phone, web or terminal while sessions delegate to each other and the code stays on the machine its owner connected.
No. 086
agentacct
Reads the session files your coding agents already write, records what they claim they did over MCP and hooks, prices the tokens from a public price list, and shows the whole thing in a local dashboard that never leaves the machine — keeping an agent’s word for done and a passing check as two separate facts.