Skip to content

Attention Span

Three drop-in output styles for coding agents — Attention-kind, Spartan and Rundown — that change how the agent talks to you rather than how it codes: answer first, plain English, built to be skimmed, with a committed benchmark that measures the work and the output separately.

Screenshot of Attention Span
Editor screenshot, 1 Oct 2026Attention Span ↗

What it is

A set of markdown files that change how a coding agent reports its work. Attention-kind is the flagship: the conclusion comes in the first line, each point is spaced out and marked with an arrow, the important words are bold so the bold alone carries the answer, and a rare technical term gets a five-word definition once. Spartan is the same shape with the warmth removed; Rundown opens with a TL;DR and shows state as a checklist. Around the three files sit a benchmark that measures the work and the output separately, a release workflow that fails when the version file, the markers inside the styles and the git tag drift apart, and four user-invoked skills generated from the style sources so they cannot drift from the wording. The styles install as a Claude Code plugin, work in the Claude apps as skills, and ship as plain markdown bodies that can be stripped of their frontmatter and dropped into another agent’s rules file.

Who built itThe repository is close to one person’s work: 61 of its 66 commits are the author’s, and the other five came from four accounts, each of which the changelog credits with a change that landed — the multi-agent install, the Spanish README, the style switcher command, and a rule about brevity. Nineteen commits carry a co-author trailer, eighteen of them for Claude Opus 4.8 (1M context) and one for Opus 5.

How it is put together

The parts · 6

The organising idea is that the interface between a coding agent and its user is a document rather than a program. Each style is one plain markdown file whose body carries no tool-specific behaviour, so the same text can be an output style inside one CLI, an appended block in another agent’s rules file, or the source a script turns into user-invoked skills. The only tool-specific part is the frontmatter at the top, and the installs strip it at a marker instead of maintaining a second copy. The second idea is that a style which makes a model say less has to be measured on two axes at once, because the obvious test — ask a model whether the answer is better — cannot separate a shorter answer from a worse one: the work is checked by hidden tests, where pass or fail is ground truth, and the output is measured from the shape of the text, deterministically, with the raw generations and the harness committed next to the numbers. Everything else in the repository is packaging around those two decisions: a plugin manifest, a release workflow that keeps the manifest, the version file, the in-file markers and the tag in agreement, and a styles directory whose three files are the product.

output-styles/
The product: three markdown files — attention-kind.md at 7,793 bytes, spartan.md at 3,676 and rundown.md at 2,631 — each one frontmatter, a version marker and the rule body, in that order.
skills/ and scripts/gen-skills.py
Four user-invoked skills for the Claude apps — attention-kind at 7,918 bytes, spartan at 3,774, rundown at 2,719 and tldr at 2,207 — produced from the style files by a 3,007-byte script so the two sets cannot drift apart.
benchmarks/
Five directories: code-eval (12 task definitions, their hidden tests, reference solutions, and the runners for work-equivalence and deliverable purity), harness (hermetic answer generation, the scannability metric, a blocking-placement checker), blocking (six side-effect scenarios with real runs and a non-regression probe), questions (the held-out question set) and results, which holds the 2026-08-11 writeup and a 295 KB file of raw generations.
commands/style.md
One 4,130-byte slash command that lists the installed styles, sets one, or restores the built-in, reading both the global and the project styles folder and writing the matching settings file as JSON so that clearing a style from a single-key file leaves valid JSON behind rather than an empty file.
.claude-plugin/ and .github/workflows/release.yml
The plugin manifest, pinned at version 0.8.0 with the author’s page and the AGPL-3.0 licence, the marketplace file beside it, and a 2,271-byte release job that fails when the manifest, the VERSION file, the in-file markers and the tag disagree.
Root documents and assets
README.md at 20,151 characters with Spanish and Chinese versions beside it, a 5,446-byte changelog, VERSION holding 0.8, the 34,523-byte AGPL-3.0 text, and 12 MB of assets led by a 3.2 MB hero image, a 3.0 MB illustration and two cat images of 2.9 and 2.5 MB.

Choices, and what they beat

  • Ship the style as one markdown file over a tool that wraps the agent or rewrites its prompts at runtime

    The README states that the styles change how the agent talks, not how it codes, and each file keeps the coding instructions intact. The hidden-test arm of the benchmark exists to check that claim, and both arms pass the same number of runs.

  • Measure the work with hidden tests and the output from the shape of the text over asking a model whether the answer is better

    The harness README says that question conflates completeness with quality and cannot fairly judge a style whose job is to say less. The writeup then lists what the method does not cover: open-ended design work with no ground truth, small samples, and readability metrics that are proxies for reading rather than a comprehension study.

  • No word ceiling anywhere in the styles over a hard maximum length

    The 0.2 changelog gives the reason: a hard cap makes the model optimize for the number over the answer. Length is made to signal importance instead, by expanding only where cutting the expansion would cost the reader.

  • Keep the style body provider-agnostic and strip the frontmatter at a marker over a separate copy of the style for each agent

    The pull request that added the installs for other agents argued that the body has no tool-specific behaviour and that only the YAML frontmatter does not travel, and the 0.3 changelog records the outcome as no duplication and no drift.

  • Generate the skills from the style files over maintaining a second set of prompts for the Claude apps

    The README says the skills come from the same sources by way of a script, so they never drift from the flagship wording, and the changelog ties them to an issue that asked for an official version updatable directly from the source.

  • Make brevity govern the reply rather than the work over letting a terse style also do less, or hand back an unverified conclusion

    The 0.8 changelog credits a contributor for a change it describes as “brevity governs the reply, not the work”, and the pull request behind it came from a session where the agent treated the write-up as the finish line and stopped short of verifying what it was about to post.

Read fromREADME.md (20,151 characters), CHANGELOG.md, benchmarks/results/2026-08-11-benchmark.md, benchmarks/harness/README.md, benchmarks/code-eval/README.md, .claude-plugin/plugin.json, VERSION, the issue and pull request threads recorded in the recon report, and the complete 55-file tree with sizes.

Build log

5 stages
  1. 01

    One markdown file becomes a plugin in a month

    The repository was created on 2026-08-04 and has 66 commits, 60 of them in August and six in the first six days of September. Seven releases came out of that month, one every few days at first: 0.2 on 2026-08-05, 0.3 the next day, 0.4 on 2026-08-10, 0.5 on 2026-08-11, 0.6 on 2026-08-15, 0.7 on 2026-08-21 and 0.8 on 2026-09-06. The first entry in that list is 0.2; there is no 0.1. The tags match the releases exactly, because the release workflow fails the release when the version file, the version marker inside each style and the tag disagree, and versioning is deliberately flat: the changelog says only the rightmost number bumps. Around that sit 1,150 stars, 37 forks and two open issues, and five accounts have committed — the author on 61, four others with one or two each. Nineteen commits carry a co-author trailer, eighteen of them naming Claude Opus 4.8 with a one-million-token context window.

  2. 02

    What the rewrite actually looks like

    The README argues with two side-by-side answers to the same question. Asked which database a new social app should use, the unstyled answer runs 430 words and reaches its recommendation after a paragraph of framing; the Attention-kind answer is 94 words in five arrow-marked points, the first of which is the recommendation. Spartan does the same to a question about cutting one of three priorities, 310 words down to 168, and Rundown turns a hiring summary into a TL;DR, a four-line checklist, one blocker and four numbered next moves. What the files ask for is narrower than “be brief”: answer first, say the least that fully answers, expand only where cutting the expansion would cost the reader, define a rare term once, never restate a point. The changelog records sharpening rather than invention. 0.2 split answers from deliverables, so an explanation or a decision stays lean while a document runs as long as the work needs, and stated that brevity trims the reply and never the internal reasoning. 0.5 added deliverable purity — ask for a commit message or an email and you get that and nothing around it — plus a rule that a warning is never trimmed to save space, and one that the bold and the TL;DR alone must carry the whole answer.

  3. 03

    The benchmark was rebuilt after somebody took it apart

    Five days after the first benchmark, a reader opened an issue arguing it measured compliance rather than quality: the model generating the answers was the model judging them, the sample was 12 questions times two runs, and the two metrics that moved — answer-first from 63% to 96%, skimmability from 2.7 to 4.8 — restated the style’s own instructions as a rubric, so of course they improved. The author agreed in the reply, said the benchmarks could have been better and rebuilt them. The version dated 2026-08-11 is built to be attacked: the work is checked by hidden test suites where pass or fail is ground truth, the solutions are generated hermetically with the ambient style off so the arms differ by one file, and the headline numbers use no model as a judge at all. On 12 coding tasks the two arms pass the same number of runs, 35 of 36; on 24 questions the styled answers are 43% shorter on average and 50 to 71% shorter on the verbose ones, while a two-line answer shrank 7% — the saving scales with how much the answer would otherwise over-explain. Instead of a reading-grade score, which sees only sentence length, it counts words before the first emphasized point: about 6 against about 40, with the answer in the first line 75% of the time against 3%. The writeup ends by listing its own limits, including that work-equivalence is shown only on tasks that have hidden tests.

  4. 04

    Every rule arrived with a user attached

    Most of the rules in these files came from people running them daily. Issue 6 reported a blocking question buried in paragraph four of seven, where a yes-or-no answer preceded a step with a side effect; 0.7 shipped the fix as a placement rule — a question you must wait on is the last block with nothing after it, and line one carries it when the reply has other content — and the same release added a placement checker and six side-effect scenarios in which real Claude lands the ask last six times out of six while a control that buries the question fails six times out of six. Issue 9 pointed out that the checklist folded unknown into not started, and 0.8 gave it a fourth state; issue 8 asked for the next-move choices to be numbered so they could be picked by number, also shipped in 0.8. A contributor’s pull request, titled “brevity governs the reply, not the work”, came out of a session on a slow data pipeline where the agent kept stopping early and offering to write up claims it had not verified, having turned the write-up into the finish line. Issue 5 asked for the styles as skills for the Claude apps, and 0.8 shipped those as well, generated from the style files. Two community threads are still open: a request to skip the whole TL;DR-plus-next-move scaffolding on a single-fact answer, opened the day of the last push, and a pull request carrying four more README translations.

  5. 05

    The distribution story, and what it declines to claim

    The shipped artifact is a markdown file, so distribution is mostly about where the file lands. The one-step route is a Claude Code plugin that brings the three styles, the switcher command and the four skills; the manual route curls one file into an output-styles directory and names it in settings. The plugin manifest is versioned in step with the styles, and the release workflow refuses to publish while manifest, version file, in-file markers and tag disagree. The trade-offs are written down. The style body costs about 650 tokens of input, loaded once per session and cached after the first request, and the README argues the measured output saving dwarfs it in a few replies. Styles apply to the main conversation only, because subagents run their own prompt, and the files keep the coding instructions intact — the claim the hidden-test arm checks. The skills are user-invoked only, so they cost no passive context until you type one, and a script generates them from the style sources so the two sets cannot drift. The README also declines a pitch a reader might expect: if cutting token spend is the goal, the bigger cost is the work the agent does rather than how it talks, and two sister tools by the same author are named for that job. One point in the benchmark thread is left uncovered by the material: AGPL-3.0 on a prompt written in markdown — the recorded reply stops before it.

Adjacent records

All records →