unlazy
An anti-laziness skill for coding agents: the agent writes an acceptance ledger before it starts, a Node checker runs the shell command behind each gate and accepts nothing but exit code 0 plus a matching expectation, and the Depth Tree method gives every leaf of the decomposed task the full time budget of the whole job.

What it is
A skill that makes an agent write down what finished means before it starts. Each gate in the ledger is a shell command plus the output it must produce, and a run counts as a pass only when the process exits 0 and the expectation matches the combined output, so a claim of completion has to survive a re-run. A checker script — Node 16 or newer, no third-party runtime packages — can inspect a ledger without executing anything, execute it after an explicit approval, or re-run every runnable gate including the ones already marked complete, and an optional Stop hook refuses to let a Claude Code session end while gates are unmet or launch waves are unfinished. Above the single gate sits the Depth Tree method it is built on: split the task N layers deep and give every leaf the full time budget of the whole task, so the effort spent rises with the depth of the tree instead of being divided by it, and orchestrated runs claim disjoint file ownership before independent leaves start together.
Who built itA developer in Munich writing as Leonxlnx, with an account opened on 2025-07-03, 47 public repositories, 2,167 followers and the bio “I build cool stuff”. He wrote 24 of this repository’s 49 commits; the remaining 25 are spread across nine other accounts plus two commits with no linked account. His other repositories include taste-skill, at 91,596 stars, and agentic-ai-prompt-research, at 2,550.
How it is put together
The parts · 6Two artifacts with different jobs: a skill document that tells an agent how to decompose work and how to write down what completion means, and a Node script with no third-party runtime dependencies that decides whether the written-down thing is true. Trust runs one way — the model writes the ledger and the checker judges it — which is why the checker may execute shell commands but only after an explicit approval that binds the exact command, expectation, working directory, shell, timeout and inherited PATH, and why the design is organised around the one thing a checker cannot do: infer that an English title and a piece of shell code mean the same thing. Everything above a single gate — the tree of leaves, the OWNS: claims, the dispatch waves, the parent re-verification — exists so that a hierarchy of agents can be verified by that same primitive, bottom-up, with each parent re-running its children instead of believing them.
- SKILL.md
- The skill itself, 10,520 bytes: core instructions and mode routing, the file an agent reads once the skill is invoked, and the only part that is prose rather than machinery.
- references/
- Six documents totalling 41 KB, one per subject:
gates.md(15,368 bytes) for the strict ledger format, approval, shell and authoring rules;orchestration.md(7,324) for leaf and branch states, rolling dispatch and the verification hierarchy;parallel.md(6,596) for the limits of scope and lease coordination;dispatch.md(5,901) for native launch waves, host adapters and recovery;method.md(3,853) for the Depth Tree decomposition itself; andtoken-economy.md(3,288) for attention and verification cost. - scripts/ and scripts/lib/
- Five executables over five library files — 72 KB and 59 KB.
gate-check.mjs(39,966 bytes) is the checker,gate-lint.mjs(9,540) an advisory non-executing linter,stop-hook.mjs(9,230) the optional hook,install-hooks.mjs(8,639) the installer anddispatch-check.mjs(6,250) the dispatch recorder; underneath sitgates.mjs(39,395),dispatch.mjs(13,299),process-tree.mjs(5,994),check-supervisor.mjs(1,462) and a 282-byte regex worker. - templates/
- Three ledgers, 10 KB:
PLAN.md(6,492 bytes),gates-node.md(2,139) andgates-leaf.md(2,067). After one community pull request the plan also carries a leaf dispatch table withOwns,Needs,TierandStatecolumns plus an explicit wave schedule, so what can run in parallel is visible at plan time rather than only in the driver’s head. - tests/
- Seven files, 199 KB, the largest surface in the repository and the one that grew fastest:
hardening-tests.mjs(69,156 bytes),dispatch-tests.mjs(34,447),stress-tests.mjs(34,116), the runner (30,761),self-check.mjs(16,171),lint-tests.mjs(14,435) andcontract-tests.mjs(4,656). They run throughnpm testand through a 1,491-byte workflow that spans three operating systems and three Node major versions. - SECURITY.md, research/validation-protocol.md and CHANGELOG.md
- The three documents that make claims auditable: a 12,742-byte threat model for
CHECK:, shell, approval, hook and lease behaviour; a 4,442-byte protocol setting out the historical limitations and what a defensible rerun would take; and a 16,636-byte changelog holding contributor history and pull-request links, which is where the credit for each integrated change lives.
Choices, and what they beat
A shell command the user approves, rather than a claim the agent makes over trusting a model’s account of what it verified
The README states the boundary the checker cannot cross — it can prove only the command oracle that was declared, and it cannot infer that an English title and arbitrary shell code mean the same thing — and the rest of the design follows from it: the ledger is written first,
--statusis the one mode that never executes, and a normal run on an unapproved oracle prints the command, expectation, working directory, shell and PATH rather than running them.Approval covers the declared command and environment, not the files it calls over hashing every script, fixture and dependency a check reads
Approval binds the exact
CHECK:andEXPECT:, the resolved working directory and shell, the timeout, the output and regex limits, the platform and the full inherited PATH; the documentation says outright that it does not hash called scripts, fixtures or other transitive inputs, so changed dependencies have to be reviewed and re-verified. A report asking for the stricter binding was closed as a documented boundary, on the grounds that a safe opt-in version needs a separate design for paths, storage, races, Windows and stop latency.Sequential checks by default, with parallelism as a number you opt into over running every gate that looks unrelated at the same time
Dispatch is rolling, but gate checks stay sequential unless
--jobs <N>is passed, and a documentation pull request defined independence as runtime independence — noCHECK:writes anything anotherCHECK:reads or writes, including working directories, generated files, caches, ports, databases and lock files — with the instruction to stay sequential when in doubt. A feature request for parallel agent workflows was answered in the same direction: launch every independent ready leaf before the first wait, and leave the checks themselves ordered.An abandoned gate is a handoff, not a pass over letting an impossible gate count as met
A valid abandonment is terminal handoff rather than success: the checker exits 1 with
HANDOFF REQUIREDand reports the qualified ids, and the Stop hook allows the session to end with an explicit non-completion message instead of a clean bill of health. This was tightened after a report showed an abandoned required gate printingALL METand giving a parent completion credit; the stated rule is thatMETmeans verified completion andABANDONEDmeans truthful non-completion, so it cannot promote a parent.Evidence is bound to the gate definition over reusing a previous pass as proof of the current command
A report showed
--approvereporting PASS from evidence recorded before theCHECKtext was edited, which is a false completion signal. Automatic evidence now begins with a versioned definition digest of the parsedCHECK:,EXPECT:and rawCWD:, the runtime approval identity is kept separate from it, checker,--statusand Stop share one stale-unmet model, and legacy or mismatched evidence fails closed.Launch waves seal before the first wait over a driver loop that starts one worker and waits for it
The dispatch contract requires each declared leaf in a ready set to receive a distinct host handle and the wave to be sealed before any result is collected, which makes the serial launch-then-wait pattern invalid rather than merely discouraged. Wave state is recorded under the scope, and a partial launch that cannot recover has to go through an audited abandonment transition — the instructions say never to invent a handle or delete state.
Read fromREADME.md, 17,238 characters, fetched and read in full through the GitHub API — including its repository map, its research section and its dated sources list — plus the complete 37-file tree with sizes, and the pull-request bodies and maintainer replies quoted in the recon report.
Build log
6 stages- 01
An anti-laziness skill with 3,782 stars in its first month
The repository was created on 2026-08-09T23:39:12Z and its oldest commit is stamped fifteen seconds earlier — 2026-08-09T23:38:57Z, “unlazy: anti-laziness skill built on the Depth Tree method” — which is what a local history pushed into a fresh repository looks like. The published history is 49 commits over twenty-five days: 48 in August and one in September, the newest dated 2026-09-03T09:31:08Z, “fix: bind gate evidence and harden Windows identity”. By 2026-10-01 it held 3,782 stars, 281 forks, 12 watchers and nine open items, eight of them pull requests, the newest of those dated 2026-09-22. Ten accounts appear in the contributor list; the owner wrote 24 of the 49 commits, 47 commits carry a linked account, and twenty carry co-author trailers, every one of them a Claude model — twelve Opus 5, five Fable 5, two Opus 4.8 with a million-token context and one more Opus 5 with the same context. This record files the project as maintained rather than active: four weeks without a push is a pause rather than an ending, nothing is archived, and the queue kept moving after the commits stopped — eight pull requests are still open, dated from 2026-09-11 to 2026-09-22, every one of them newer than the last commit on
main. - 02
The Depth Tree, and a research claim written out in full
The repository states its core in one line: the Depth Tree method, which splits a task N layers deep and gives every leaf the full time budget of the whole task, so effort multiplies with depth — grounded, it says, in 2025–2026 research on model laziness, underthinking and premature completion. Its research section refuses the strong reading of that grounding: research supports the failure modes motivating explicit structure, it says, and does not prove that unlazy produces a fixed improvement; it then argues from numbers. SlopCodeBench, where the best agent passed 14.8% of checkpoints; METR’s Time Horizon 1.1, reporting a 196.5-day overall P50 doubling-time fit against 130.8 days after 2023, with a warning not to quote the shorter figure as the all-years estimate; and s1’s budget forcing, which appends Wait when a model tries to stop, while denying that one token always improves work. Below that is a sources list ordered newest first, from 2025-02-18 to 2026-08-06 plus one undated page. Every entry carries a title, a link and a date, two of them a venue, and none an author name, so the citations are specific but not full references. The README also withdraws earlier evidence of its own: previous revisions cited a six-run internal comparison whose raw artifacts are not in the repository, to be read as historical design input rather than a benchmark guarantee.
- 03
What the agent has to write down before it is allowed to stop
Installation is one command through the skills CLI,
npx skills add Leonxlnx/unlazy, or a clone into~/.claude/skills/unlazyor~/.codex/skills/unlazy; it answers to/unlazy, to$unlazyin Codex, or to a natural-language trigger. The README’s example is/unlazy tree 5 refactor the payment module and verify every migration path; that number is the depth. What the agent owes is a ledger: each gate is a checkbox line with a title, aCHECK:shell command, anEXPECT:string, an optionalCWD:and anEVIDENCE:line. The checker has three states:--status, the only mode that never executes; a normal run, which on an unapproved oracle prints the resolved command, expectation, working directory, shell and inherited PATH instead of running it, and which the README warns is not a permanent dry run; and--approve, after the user has read every command and script it calls.--reverifyre-runs every runnable gate, including completed ones. A runnable gate passes only when the process exits 0 and the expectation matches the combined output, both capped at 1 MiB; a larger matcher string is never truncated into success. The parser rejects ledgers with no gates, duplicate ids, incomplete runnable gates, invalid expectations, or an abandonment with no reason or an unknown gate id. Abandonment is a terminal handoff rather than a pass: exit 1 with HANDOFF REQUIRED. - 04
The hook, the shell, and the machine where nothing could be read
The optional piece is a Claude Code Stop hook, installed only with consent through
install-hooks.mjs: it scans the session’s resolved ledger and dispatch state and returns Claude Code’s top-leveldecision: "block"while gates are unmet or waves are incomplete, without running any check itself. Its session-keyed guard releases after six consecutive blocks without semantic progress, and metadata-only edits do not reset it — a fix from a pull request where the earlier version hashed ledger bytes, so any edit rearmed the counter and trapped an agent that could not finish. The checker takes its shell from--shell, thenUNLAZY_SHELL, then the platform default, and checks inherit the launch environment including PATH, so the same checker started from Git Bash sees tools that one started from PowerShell does not;--shellchanges the interpreter, it does not installgreportail. On Windows a reader found every file read failing closed: on his machinefstat().devnever equalslstat().dev, so the checker could read no ledger and a clean clone scored 5 of 32 tests. The fix compares bigint descriptor identities through a second non-creating handle; the snapshot-based guarantee in SECURITY.md stays. An earlier Windows round removed orphaned processes: killing the shell left a check’s children running, so a timeout now callstaskkill /pid <PID> /f /t. - 05
Gates that could not fail, and the argument nobody read
The most useful rounds came from people running the skill on real work. One report showed that a gate whose
CHECK:never observes what its title claims still returns PASS with evidence and exit 0, because the leaf check,--reverifyand the Stop hook all consult the same oracle; a second reporter reached the same class from a different direction, trialling the skill on a sixteen-week course readiness audit with gates authored by another model under another harness and no hook installed, where the gate that worried him guarded the task’s one hard constraint — read only, no writes. The maintainer fixed the machine-detectable false greens and merged the authoring guidance, while saying plainly that the checker still cannot prove that prose and command mean the same thing and that the boundary is now explicit rather than hidden. A separate bug is worth recording because it was reported more than once:indexOf("--timeout")returns -1 when the option is absent, so the next index is 0, the filter drops the first positional argument, and naming a single ledger file either widened the run to every ledger in the directory or reported success without reading the file at all. It was fixed with explicit named-file targeting and a regression test, and one of the two reports was a pull request its own author closed as a duplicate of an earlier one after checking the open queue. - 06
Fifteen items of community work, and no release to put them in
Three reports went after the hierarchy rather than one gate: an abandoned required gate could print ALL MET and hand a parent completion credit; approval could stay current after the called script changed; and a plan could omit one of two requested deliverables while every declared gate passed. The first and third were fixed — abandonment can no longer promote a parent, and plans now inventory independently omittable requirements and reread the request before dispatch — while the second was closed as a documented boundary rather than a defect, because approval binds the declared command and environment, not transitive script bytes. There are no releases and no tags at all. The README says the current source targets
2.1.0, is not a tagged release, and tells readers to pin an exact commit; what would be release notes lives in a 16,636-byteCHANGELOG.mdthat also records contributor history and pull-request links. That unreleased2.1.0is where the community work landed, in fifteen listed items that integrate the useful parts of the pull requests while repairing their edge cases; the test count moved with it, from 137 at commit265fbd5dto 188, with the self-check at 15 of 15. Continuous integration runs across Ubuntu, Windows and macOS on Node 16, 20 and 24 — 9 of 9 jobs in August, 10 of 10 in September with a pinned Node 22.14.0 and libuv 1.49.2 Windows job.
Adjacent records
All records →No. 128
anything2explainer
A skill for Claude Code and Codex that turns a topic, or a document, into a 1280×720 narrated explainer film, with every frame drawn in code rather than generated and each agent in the parallel build writing one Remotion component. The same author’s video-shotcraft makes product promos out of 157 shot cards and a sound-design pass; this one is driven by the narration, ships no sound effects or music, and freezes the words into frame numbers before any shot is built.
No. 125
OKF Agent Memory
Keeps what a coding agent learns as plain Markdown inside the repository — an OKF v0.2 knowledge bundle searched in-process by BM25 — so the memory can be diffed and reviewed instead of living in a database.
No. 123
agent-memory
A long-term memory runtime for AI agents that keeps plain Markdown files as the single source of truth, ranks them locally without calling a model, answers recall with file paths the agent opens one level at a time, writes at conversation boundaries rather than on the agent’s initiative, and runs an independent sleep-time layer that may add and update on its own but can only ever file a deletion as a proposal — one store shared by Claude Code, Codex CLI and Hermes, with no API key.