Skip to content

sepia

A portable de-AI writing skill: four operations over one canonical rules file, narrative architecture repaired before word choice on fiction, a thin rule file matched to the venue on professional prose, and every rule labelled as measured, consulted or the project’s own inference.

Screenshot of sepia
Editor screenshot, 30 Sep 2026sepia ↗

What it is

sepia is a writing skill for coding agents, and its argument is about where the tell actually lives: in structure rather than in word choice. Every popular humanizer edits vocabulary and syntax; this one starts from a study in which human editors rewrote the surface style of AI fiction and a classifier still caught it — detection fell only from 95.5% to 93.9% — so fiction goes through a narrative-architecture pass before anything touches the prose, and each professional document type gets a thin rule file matched to its venue on top of a shared checklist. Four operations — write, review, refactor, recreate — run on one canonical SKILL.md, installed by any agent that speaks the Agent Skills standard (the Skills CLI claims 77+) and packaged as a native plugin for Claude Code, Codex, Grok Build, Antigravity and QwenPaw. What it does not know is labelled rather than filled in: the evidence behind each rule sits in a research directory with fifteen primary sources, and vendors that publish no prompt guidance are recorded as consulted instead of guessed.

Who built itThe repository’s 303 commits come from six accounts: 283 of them from the maintainer’s own, under two commit names — Nanako Tsai on 214 and Nyanako on 69, both mapped to the same GitHub account — then AugustusW 13, CallMeHFK 2, shihyuho 2, Franky100-pig 1 and javaht 1, with one commit left unlinked. Two thirds of the history, 202 commits, carries a co-author trailer, and 201 of those name a Claude model: Fable 5.1 on 152, Fable 5 on 25, Opus 5 (1M context) on 19 and Opus 5.5 (1M context) on 5; the one remaining trailer names a person.

How it is put together

The parts · 6

Everything here is text plus two small Python checkers, and the shape follows from one decision: the rules are the product, so the evidence for them has to travel with them. One canonical SKILL.md does the routing — which document type this is, which passes run, which calibration applies — and every other file is either a pass the router can call, a per-domain rule file, a per-model fingerprint, a language calibration, or a voice profile layered on top. The four operations are wrappers around that single skill rather than four skills, which is why the README says standalone wrapper installation is unsupported. Portability is bought by packaging rather than by forking: the same Markdown is checked in once and reached through five plugin manifests, one of them through a symlink that is also the origin of the package’s known install bug. Around the rules sit the two things the project trusts more than prose — a research directory where each claim carries its source and its limits, and validators that refuse a malformed contribution instead of parsing it. There is no service, no runtime and no model call; what runs at use time is the host agent reading these files.

skills/sepia/ and the five wrapper skills
The product itself: 21 files and 232 KB. SKILL.md is 13,642 bytes of routing, operations and guardrails; references/ holds the three passes — narrative 12,995, discourse 5,393, style 15,365 — the 30-feature rubric.md at 9,460, model-fingerprints.md at 19,021, the professional checklist at 8,136, six per-domain rule files from 1,725 to 8,985 bytes, voice-skills.md at 12,920, and the Chinese calibration languages/zh.md at 30,604. Five sibling skills are thin fixed-operation wrappers, 1,022 to 1,819 bytes each.
references/voices/
The voice layer: tw-journalism.md at 34,825 bytes is the largest file in the repository — nine narrative shapes, each with its move, its source, the sepia check it answers to and its known cost — beside hemingway.md (12,731), the prose-first PERSONA-TEMPLATE.md (7,084), the first built-in persona personas/nyaneko.md (27,623) and registry.md (6,381).
research/
Eleven digests, 167 KB by the report’s own count, led by sources.md at 61,944 bytes — the ledger where each rule names its evidence and the limits of that evidence — then hemingway.md 16,909, newswriting-guides.md 16,498, citations-style.md 15,763, zh-news-corpus.md 15,624, rhythm-syntax.md 13,312, storyscope.md 10,534, detectors.md 7,852 and citations-narrative.md 6,011.
scripts/ and tests/
Two validators and their test modules: check_persona.py 21,316 bytes with tests/test_check_persona.py at 28,797, and check_versions.py 16,378 with tests/test_check_versions.py at 23,862 — 51 KB of tests against 37 KB of scripts. The suite held 88 tests when v0.11.0 shipped and 94 by 2026-09-20.
evals/ and .github/workflows/
One behavioural eval case, deaify-release-note: a 1,160-byte prompt and three graders — skill-fired.md 104 bytes, no-slop-markers.md 234 and reads-human.md 889 — driven by a 3,996-byte workflow through claude plugin eval on every push, beside a 627-byte version-consistency check and a 3,518-byte .coderabbit.yaml that raises auto_pause_after_reviewed_commits from 5 to 20.
The five packagings and the three READMEs
plugin.json (182 bytes) for Antigravity; .claude-plugin/ (358 and 458), .codex-plugin/ (358), .qwenpaw-plugin/ (plugin.json 916, plugin.py 7,417 and a nine-byte skills symlink) and .agents/ (marketplace.json 229, workflows/sepia.md 843 and an 18-byte skills/sepia entry). The documentation is three READMEs — 20,473 English, 20,333 Simplified Chinese, 19,639 Traditional — plus CONTRIBUTING.md at 9,805 bytes.

Choices, and what they beat

  • Repair narrative architecture before touching word choice over a humanizer that edits vocabulary and syntax

    The study the project is built on measured the alternative directly: when human editors rewrote the surface style of AI fiction, a structure-only classifier still detected it, 95.5% to 93.9% macro-F1. The tells that survive that kind of editing are architectural — an explained theme, a single causally tidy track, emotion rendered only as bodily sensation, no real-world reference — so the first pass works on those.

  • Calibrate toward the human distribution rather than invert the AI one over optimising the surface statistics detectors score

    Stated as the governing principle in the README, and pushed to its conclusion in issue #268: separateness from burstiness and perplexity is deliberate, and the project says it does not aim to evade detectors. A first pull request adding optional diagnostics for exactly those axes was closed in favour of one with no detector wording anywhere in the title, body or branch name.

  • Record a vendor as consulted when the vendor says nothing over inferring how a model writes by default

    The README states it as a rule — vendors that publish no prompt guidance are recorded as consulted, not guessed — and the pull requests apply it: Opus 5.5, GPT-6 Sol and GPT-6 Luna got no operative row, and where a page did yield statements, #280 replaced a claim of three with all five and explained why the conclusion held anyway.

  • Keep sentence-length spread and discard three other surface signals over checking everything that can be counted

    The README’s table sorts four syntactic measures by what the studies say: within-passage variance is consistently higher in human text across English and Chinese corpora and is kept as a signal, while mean sentence length, punctuation counts and paragraph length are discarded because the measured directions contradict each other across corpora.

  • Refuse a malformed contribution rather than parse it over guessing at the cell boundaries of a bad table row

    The persona validator only collected override rows that begin with a pipe, so a token the format forbids could hide in a row written without one. The fix rejects the line and names it, and the pull request gives the reason in one sentence: guessing the cells of a malformed row would be more guard than row.

  • Describe a persona in prose rather than in measurements over a body built from sentence-length shares, emoji density and counted moves

    The first built-in persona was rewritten twice and settled by five runs on one fact list: the measured versions read as a format, while the writer’s own prose specification, which contains no distribution target, produced the output a reader recognised as her. The template became prose-first with one optional exemplars section, and the measurements stayed where their claims are bound to a corpus.

Read fromREADME.md (20,473 bytes) with README.zh-CN.md (20,333) and README.zh-TW.md (19,639), CONTRIBUTING.md, the complete 65-file tree with sizes, and the pull-request and issue bodies it is described from — #257, #258, #263, #264, #265, #266, #267, #268, #270, #271, #273, #274, #275, #277, #280, #281 and #283. The recon report’s architecture-document scan found no separate design document: the repository documents its shape in the README, in references/ and in research/.

Build log

6 stages
  1. 01

    Twenty-seven days, 303 commits, and what the trailers say

    The repository was created on 2026-08-28 and its last push is 2026-09-23: twenty-seven days, 303 commits and fourteen releases, from v0.2.0 on the day it appeared to v0.12.2 on 2026-09-22. The pace is uneven — 45 commits in August, 258 in September — but the shape of the authorship is what makes the history worth reading. 283 of the 303 commits come from one account under two commit names, Nanako Tsai on 214 and Nyanako on 69; 202 carry a co-author trailer, and 201 of those name a Claude model — Fable 5.1 on 152, Fable 5 on 25, Opus 5 (1M context) on 19, Opus 5.5 (1M context) on 5 — against a single trailer naming a person. Five other accounts contribute the rest: AugustusW 13, CallMeHFK 2, shihyuho 2, Franky100-pig 1, javaht 1, and one commit carries no linked account at all. The release pull requests are written in the same dialect as the code: each is a version bump of four declarations checked by python3 scripts/check_versions.py, each closes with a Claude Code session link, and each says outright that release notes live on the tag rather than in the body. The documentation was produced the same way, and by the project itself: the three READMEs were restructured by the Antigravity CLI running sepia’s own recreate operation on the documentation route, with the current file as the only fact source, the style guide inlined, and languages/zh.md loaded for the two translations.

  2. 02

    The research directory is the product, and it carries its own limits

    What separates this from a prompt pack is one habit: every rule names its source, and every source names what it does not cover. research/sources.md is 61,944 bytes — the largest file in a repository that is mostly Markdown — and the README’s source table lists fifteen primary studies. The founding one is StoryScope: 61,608 stories by humans and five frontier models, where a classifier using narrative-structure features alone reached 93.2% macro-F1 on AI fiction. The number that decides the design is the next one: in the same study’s edited condition, where human editors rewrote the surface style, detection moved only from 95.5% to 93.9%. That gap is the argument for repairing architecture instead of vocabulary. The professional route then ran on that argument with no measurement of its own, and issue #263 says so plainly — on fiction the claim carries a figure, on professional prose it did not. A preprint published on 2026-09-14 closed part of the gap: 2,250 pre-ChatGPT human blog posts from 268 company domains against 11,250 AI mirrors from five frontier models, 214 features across 11 dimensions of which 187 are structural, reaching 98.0 macro-F1 from structural features alone on held-out companies and 98.1 after each model reworded its own post. It entered the ledger as SLOPSHAPE-2026 with its limits attached: the features are LLM-scored, and the experiment never tested human editing.

  3. 03

    How the claim is checked, and the four days nobody noticed

    The distance between what the skill claims and what the repository measures is visible, and reported. There is one behavioural eval — evals/deaify-release-note/, a 1,160-byte prompt and three graders of 104, 234 and 889 bytes, asking whether the skill fired, whether the output reads human and whether it carries slop markers — and it runs on every push through claude plugin eval. That workflow was red for four days and the failure looked like nothing: from 2026-09-15, seven consecutive runs (34980491671, 35005656553, 35107215014, 35119480316, 35368600325, 35368661029, 35368703581) exited 1 before the first case ran, because the CLI had gained a first-run trust prompt that a runner cannot answer. The last green run was 34592097838 on 2026-09-11. Issue #259 lists the seven ids and notes that the exit is immediate, so the summarise step printed nothing; #260 adds --trust-plugin with a comment stating what the flag asserts and what it does not. Almost everything else that checks this repository is text and Python: two validators, 88 tests when v0.11.0 shipped on 2026-09-18 and 94 by 2026-09-20; a CodeRabbit configuration whose auto_pause_after_reviewed_commits was raised from 5 to 20, because at the default a paused review is indistinguishable from a slow one; and live-model A/B runs, which the README lists as an ongoing cost of the project rather than a feature of it.

  4. 04

    The persona turn: measurements that made the writing read as a format

    The sharpest reversal in the repository is about voice. Issue #257 opened as an RFC with evidence behind it: a private pilot from 2026-09-16 to 2026-09-18 had produced nineteen persona profiles, each written from close readings of ten to fifteen articles by one writer rather than from metrics, and every one of the nineteen needed at least one existing rule to yield. Step two shipped the interface, the template and scripts/check_persona.py; step three ran it end to end, and the maintainer’s report lists what held — the route gate, the persona-cost block, the categories that never yield — beside five gaps, among them a Consent: field with no honest value for a privately held profile. Step four then rewrote the first built-in persona twice. Five runs on one fact list settled it: the bodies built on sentence-length shares, emoji density and k over n counts read as a format rather than a person, and the output a reader recognised came from the writer’s own prose specification, which carries no distribution target. v0.12.0 carried it into the interface — the template became prose-first with fifteen fixed sections and one optional exemplars section, the validator now requires a blind-test record of none yet under Tested: untested, and a persona written to the previous template no longer validates, which is why a 0.x release took a minor bump rather than a patch.

  5. 05

    What other people sent, and what the project would not do

    Thirty issues and pull requests sit in the report; the most instructive are the ones the project did not take. Issue #268, filed by a contributor, is a limitation stated against the project’s own interest: sepia’s output can still be flagged as mixed by detectors, because it deliberately does not optimise for burstiness or perplexity — the two axes those detectors lean on — and because it states that it is not trying to evade them. The same contributor then built the optional skills that would have covered those axes; the first pull request was closed in favour of a second whose title, body and branch name never mention a detector, on the argument that the wording inverts the project’s stated position. Issue #267 is a thank-you from someone maintaining manuscript-polishing skills for journal submission, and it is careful about the boundary: the framing was taken, no text, no rule files, no corpus. Issue #274 measures a shipped package — an install whose skills symlink has been flattened reports success and lands nothing — and its author posted a correction retracting the fix he had proposed first. Issue #283, the newest item in the report, is the least flattering: a 2,900-character Chinese popular-science essay falls through the routing table into the generic prose route, so the narrative and discourse passes, where its defects actually are, never run.

  6. 06

    A fingerprint table kept honest one vendor page at a time

    A 19,021-byte reference file claims to know how each model writes by default, and the pull requests are the record of what it costs to keep that claim honest. When Claude Opus 5.5, GPT-6 Sol and GPT-6 Luna shipped, #275 recorded them as consulted with no prose statement and gave them no operative row, leaving the existing tables in place as priors, because their pages say nothing about default length, register, formatting or tone. #277 moved the Gemini row down from an upper bound the vendor never wrote to the scope the pages actually support — Gemini 3 and 3.1 operative, 3.5 and 3.8 Flash priors only — on pages re-read on 2026-09-23 after the vendor had updated them on 2026-09-17. #280 re-read the Opus 5.5 page, found two more statements about writing, and replaced a claim that there were only three with all five quoted, next to the reason the conclusion still holds: none of the five names a default length, register, formatting or tone. The newest is an audit of the skill’s own text by a vendor tool — #281 ran Anthropic’s /claude-api prompt-audit, the spec bundled with Claude Code 2.1.280 and byte-identical to an open pull request in anthropics/skills at a named head commit, applied six medium-confidence rewrites and left routing, operations and rules unchanged. One of the six makes a citation name GPT-5, DeepSeek-V3 and o3-mini and admit that no Claude model was measured.

Adjacent records

All records →