sepia
A portable de-AI writing skill: four operations over one canonical rules file, narrative architecture repaired before word choice on fiction, a thin rule file matched to the venue on professional prose, and every rule labelled as measured, consulted or the project’s own inference.

What it is
sepia is a writing skill for coding agents, and its argument is about where the tell actually lives: in structure rather than in word choice. Every popular humanizer edits vocabulary and syntax; this one starts from a study in which human editors rewrote the surface style of AI fiction and a classifier still caught it — detection fell only from 95.5% to 93.9% — so fiction goes through a narrative-architecture pass before anything touches the prose, and each professional document type gets a thin rule file matched to its venue on top of a shared checklist. Four operations — write, review, refactor, recreate — run on one canonical SKILL.md, installed by any agent that speaks the Agent Skills standard (the Skills CLI claims 77+) and packaged as a native plugin for Claude Code, Codex, Grok Build, Antigravity and QwenPaw. What it does not know is labelled rather than filled in: the evidence behind each rule sits in a research directory with fifteen primary sources, and vendors that publish no prompt guidance are recorded as consulted instead of guessed.
Who built itThe repository’s 303 commits come from six accounts: 283 of them from the maintainer’s own, under two commit names — Nanako Tsai on 214 and Nyanako on 69, both mapped to the same GitHub account — then AugustusW 13, CallMeHFK 2, shihyuho 2, Franky100-pig 1 and javaht 1, with one commit left unlinked. Two thirds of the history, 202 commits, carries a co-author trailer, and 201 of those name a Claude model: Fable 5.1 on 152, Fable 5 on 25, Opus 5 (1M context) on 19 and Opus 5.5 (1M context) on 5; the one remaining trailer names a person.
How it is put together
The parts · 6Everything here is text plus two small Python checkers, and the shape follows from one decision: the rules are the product, so the evidence for them has to travel with them. One canonical SKILL.md does the routing — which document type this is, which passes run, which calibration applies — and every other file is either a pass the router can call, a per-domain rule file, a per-model fingerprint, a language calibration, or a voice profile layered on top. The four operations are wrappers around that single skill rather than four skills, which is why the README says standalone wrapper installation is unsupported. Portability is bought by packaging rather than by forking: the same Markdown is checked in once and reached through five plugin manifests, one of them through a symlink that is also the origin of the package’s known install bug. Around the rules sit the two things the project trusts more than prose — a research directory where each claim carries its source and its limits, and validators that refuse a malformed contribution instead of parsing it. There is no service, no runtime and no model call; what runs at use time is the host agent reading these files.
- skills/sepia/ and the five wrapper skills
- The product itself: 21 files and 232 KB.
SKILL.mdis 13,642 bytes of routing, operations and guardrails;references/holds the three passes — narrative 12,995, discourse 5,393, style 15,365 — the 30-featurerubric.mdat 9,460,model-fingerprints.mdat 19,021, the professional checklist at 8,136, six per-domain rule files from 1,725 to 8,985 bytes,voice-skills.mdat 12,920, and the Chinese calibrationlanguages/zh.mdat 30,604. Five sibling skills are thin fixed-operation wrappers, 1,022 to 1,819 bytes each. - references/voices/
- The voice layer:
tw-journalism.mdat 34,825 bytes is the largest file in the repository — nine narrative shapes, each with its move, its source, the sepia check it answers to and its known cost — besidehemingway.md(12,731), the prose-firstPERSONA-TEMPLATE.md(7,084), the first built-in personapersonas/nyaneko.md(27,623) andregistry.md(6,381). - research/
- Eleven digests, 167 KB by the report’s own count, led by
sources.mdat 61,944 bytes — the ledger where each rule names its evidence and the limits of that evidence — thenhemingway.md16,909,newswriting-guides.md16,498,citations-style.md15,763,zh-news-corpus.md15,624,rhythm-syntax.md13,312,storyscope.md10,534,detectors.md7,852 andcitations-narrative.md6,011. - scripts/ and tests/
- Two validators and their test modules:
check_persona.py21,316 bytes withtests/test_check_persona.pyat 28,797, andcheck_versions.py16,378 withtests/test_check_versions.pyat 23,862 — 51 KB of tests against 37 KB of scripts. The suite held 88 tests when v0.11.0 shipped and 94 by 2026-09-20. - evals/ and .github/workflows/
- One behavioural eval case,
deaify-release-note: a 1,160-byte prompt and three graders —skill-fired.md104 bytes,no-slop-markers.md234 andreads-human.md889 — driven by a 3,996-byte workflow throughclaude plugin evalon every push, beside a 627-byte version-consistency check and a 3,518-byte.coderabbit.yamlthat raisesauto_pause_after_reviewed_commitsfrom 5 to 20. - The five packagings and the three READMEs
plugin.json(182 bytes) for Antigravity;.claude-plugin/(358 and 458),.codex-plugin/(358),.qwenpaw-plugin/(plugin.json916,plugin.py7,417 and a nine-byteskillssymlink) and.agents/(marketplace.json229,workflows/sepia.md843 and an 18-byteskills/sepiaentry). The documentation is three READMEs — 20,473 English, 20,333 Simplified Chinese, 19,639 Traditional — plusCONTRIBUTING.mdat 9,805 bytes.
Choices, and what they beat
Repair narrative architecture before touching word choice over a humanizer that edits vocabulary and syntax
The study the project is built on measured the alternative directly: when human editors rewrote the surface style of AI fiction, a structure-only classifier still detected it, 95.5% to 93.9% macro-F1. The tells that survive that kind of editing are architectural — an explained theme, a single causally tidy track, emotion rendered only as bodily sensation, no real-world reference — so the first pass works on those.
Calibrate toward the human distribution rather than invert the AI one over optimising the surface statistics detectors score
Stated as the governing principle in the README, and pushed to its conclusion in issue #268: separateness from burstiness and perplexity is deliberate, and the project says it does not aim to evade detectors. A first pull request adding optional diagnostics for exactly those axes was closed in favour of one with no detector wording anywhere in the title, body or branch name.
Record a vendor as consulted when the vendor says nothing over inferring how a model writes by default
The README states it as a rule — vendors that publish no prompt guidance are recorded as consulted, not guessed — and the pull requests apply it: Opus 5.5, GPT-6 Sol and GPT-6 Luna got no operative row, and where a page did yield statements, #280 replaced a claim of three with all five and explained why the conclusion held anyway.
Keep sentence-length spread and discard three other surface signals over checking everything that can be counted
The README’s table sorts four syntactic measures by what the studies say: within-passage variance is consistently higher in human text across English and Chinese corpora and is kept as a signal, while mean sentence length, punctuation counts and paragraph length are discarded because the measured directions contradict each other across corpora.
Refuse a malformed contribution rather than parse it over guessing at the cell boundaries of a bad table row
The persona validator only collected override rows that begin with a pipe, so a token the format forbids could hide in a row written without one. The fix rejects the line and names it, and the pull request gives the reason in one sentence: guessing the cells of a malformed row would be more guard than row.
Describe a persona in prose rather than in measurements over a body built from sentence-length shares, emoji density and counted moves
The first built-in persona was rewritten twice and settled by five runs on one fact list: the measured versions read as a format, while the writer’s own prose specification, which contains no distribution target, produced the output a reader recognised as her. The template became prose-first with one optional exemplars section, and the measurements stayed where their claims are bound to a corpus.
Read fromREADME.md (20,473 bytes) with README.zh-CN.md (20,333) and README.zh-TW.md (19,639), CONTRIBUTING.md, the complete 65-file tree with sizes, and the pull-request and issue bodies it is described from — #257, #258, #263, #264, #265, #266, #267, #268, #270, #271, #273, #274, #275, #277, #280, #281 and #283. The recon report’s architecture-document scan found no separate design document: the repository documents its shape in the README, in references/ and in research/.
Build log
6 stages- 01
Twenty-seven days, 303 commits, and what the trailers say
The repository was created on 2026-08-28 and its last push is 2026-09-23: twenty-seven days, 303 commits and fourteen releases, from v0.2.0 on the day it appeared to v0.12.2 on 2026-09-22. The pace is uneven — 45 commits in August, 258 in September — but the shape of the authorship is what makes the history worth reading. 283 of the 303 commits come from one account under two commit names, Nanako Tsai on 214 and Nyanako on 69; 202 carry a co-author trailer, and 201 of those name a Claude model — Fable 5.1 on 152, Fable 5 on 25, Opus 5 (1M context) on 19, Opus 5.5 (1M context) on 5 — against a single trailer naming a person. Five other accounts contribute the rest: AugustusW 13, CallMeHFK 2, shihyuho 2, Franky100-pig 1, javaht 1, and one commit carries no linked account at all. The release pull requests are written in the same dialect as the code: each is a version bump of four declarations checked by
python3 scripts/check_versions.py, each closes with a Claude Code session link, and each says outright that release notes live on the tag rather than in the body. The documentation was produced the same way, and by the project itself: the three READMEs were restructured by the Antigravity CLI running sepia’s ownrecreateoperation on the documentation route, with the current file as the only fact source, the style guide inlined, andlanguages/zh.mdloaded for the two translations. - 02
The research directory is the product, and it carries its own limits
What separates this from a prompt pack is one habit: every rule names its source, and every source names what it does not cover.
research/sources.mdis 61,944 bytes — the largest file in a repository that is mostly Markdown — and the README’s source table lists fifteen primary studies. The founding one is StoryScope: 61,608 stories by humans and five frontier models, where a classifier using narrative-structure features alone reached 93.2% macro-F1 on AI fiction. The number that decides the design is the next one: in the same study’s edited condition, where human editors rewrote the surface style, detection moved only from 95.5% to 93.9%. That gap is the argument for repairing architecture instead of vocabulary. The professional route then ran on that argument with no measurement of its own, and issue #263 says so plainly — on fiction the claim carries a figure, on professional prose it did not. A preprint published on 2026-09-14 closed part of the gap: 2,250 pre-ChatGPT human blog posts from 268 company domains against 11,250 AI mirrors from five frontier models, 214 features across 11 dimensions of which 187 are structural, reaching 98.0 macro-F1 from structural features alone on held-out companies and 98.1 after each model reworded its own post. It entered the ledger asSLOPSHAPE-2026with its limits attached: the features are LLM-scored, and the experiment never tested human editing. - 03
How the claim is checked, and the four days nobody noticed
The distance between what the skill claims and what the repository measures is visible, and reported. There is one behavioural eval —
evals/deaify-release-note/, a 1,160-byte prompt and three graders of 104, 234 and 889 bytes, asking whether the skill fired, whether the output reads human and whether it carries slop markers — and it runs on every push throughclaude plugin eval. That workflow was red for four days and the failure looked like nothing: from 2026-09-15, seven consecutive runs (34980491671, 35005656553, 35107215014, 35119480316, 35368600325, 35368661029, 35368703581) exited 1 before the first case ran, because the CLI had gained a first-run trust prompt that a runner cannot answer. The last green run was 34592097838 on 2026-09-11. Issue #259 lists the seven ids and notes that the exit is immediate, so the summarise step printed nothing; #260 adds--trust-pluginwith a comment stating what the flag asserts and what it does not. Almost everything else that checks this repository is text and Python: two validators, 88 tests when v0.11.0 shipped on 2026-09-18 and 94 by 2026-09-20; a CodeRabbit configuration whoseauto_pause_after_reviewed_commitswas raised from 5 to 20, because at the default a paused review is indistinguishable from a slow one; and live-model A/B runs, which the README lists as an ongoing cost of the project rather than a feature of it. - 04
The persona turn: measurements that made the writing read as a format
The sharpest reversal in the repository is about voice. Issue #257 opened as an RFC with evidence behind it: a private pilot from 2026-09-16 to 2026-09-18 had produced nineteen persona profiles, each written from close readings of ten to fifteen articles by one writer rather than from metrics, and every one of the nineteen needed at least one existing rule to yield. Step two shipped the interface, the template and
scripts/check_persona.py; step three ran it end to end, and the maintainer’s report lists what held — the route gate, the persona-cost block, the categories that never yield — beside five gaps, among them aConsent:field with no honest value for a privately held profile. Step four then rewrote the first built-in persona twice. Five runs on one fact list settled it: the bodies built on sentence-length shares, emoji density andkoverncounts read as a format rather than a person, and the output a reader recognised came from the writer’s own prose specification, which carries no distribution target. v0.12.0 carried it into the interface — the template became prose-first with fifteen fixed sections and one optional exemplars section, the validator now requires a blind-test record ofnone yetunderTested: untested, and a persona written to the previous template no longer validates, which is why a 0.x release took a minor bump rather than a patch. - 05
What other people sent, and what the project would not do
Thirty issues and pull requests sit in the report; the most instructive are the ones the project did not take. Issue #268, filed by a contributor, is a limitation stated against the project’s own interest: sepia’s output can still be flagged as mixed by detectors, because it deliberately does not optimise for burstiness or perplexity — the two axes those detectors lean on — and because it states that it is not trying to evade them. The same contributor then built the optional skills that would have covered those axes; the first pull request was closed in favour of a second whose title, body and branch name never mention a detector, on the argument that the wording inverts the project’s stated position. Issue #267 is a thank-you from someone maintaining manuscript-polishing skills for journal submission, and it is careful about the boundary: the framing was taken, no text, no rule files, no corpus. Issue #274 measures a shipped package — an install whose
skillssymlink has been flattened reports success and lands nothing — and its author posted a correction retracting the fix he had proposed first. Issue #283, the newest item in the report, is the least flattering: a 2,900-character Chinese popular-science essay falls through the routing table into the generic prose route, so the narrative and discourse passes, where its defects actually are, never run. - 06
A fingerprint table kept honest one vendor page at a time
A 19,021-byte reference file claims to know how each model writes by default, and the pull requests are the record of what it costs to keep that claim honest. When Claude Opus 5.5, GPT-6 Sol and GPT-6 Luna shipped, #275 recorded them as consulted with no prose statement and gave them no operative row, leaving the existing tables in place as priors, because their pages say nothing about default length, register, formatting or tone. #277 moved the Gemini row down from an upper bound the vendor never wrote to the scope the pages actually support — Gemini 3 and 3.1 operative, 3.5 and 3.8 Flash priors only — on pages re-read on 2026-09-23 after the vendor had updated them on 2026-09-17. #280 re-read the Opus 5.5 page, found two more statements about writing, and replaced a claim that there were only three with all five quoted, next to the reason the conclusion still holds: none of the five names a default length, register, formatting or tone. The newest is an audit of the skill’s own text by a vendor tool — #281 ran Anthropic’s
/claude-api prompt-audit, the spec bundled with Claude Code 2.1.280 and byte-identical to an open pull request inanthropics/skillsat a named head commit, applied six medium-confidence rewrites and left routing, operations and rules unchanged. One of the six makes a citation name GPT-5, DeepSeek-V3 and o3-mini and admit that no Claude model was measured.
Adjacent records
All records →No. 080
HarnessRouter
The self-hosted, Apache-2.0 edition of HarnessRouter: it puts sixteen existing agent CLIs — Codex, Claude Code, Hermes, DeepSeek Harness and twelve more — behind one OpenAI Responses-compatible API, with sessions, streaming, files, cancellation and structured failures, and it carries the Unified Harness Protocol it implements together with the conformance suite that measures it.
No. 070
OpenChatCut
A local-first video editor whose editing surface is a conversation: the built-in agent and external Codex or Claude Code sessions call the same editing tools the interface itself uses, so every change lands on a real multi-track timeline as a clip, transition, caption, effect or audio item that can still be dragged, undone and exported. Projects and media stay on the machine, and preview and final render both come out of Remotion.
No. 064
delegate-skills
A skills package in which every coding-agent CLI gets its own delegation skill: the orchestrating agent writes a self-contained brief, a separate CLI edits a real working tree, and the human keeps the review and the commit.