Skip to content

Drama Skills

Eleven independently installable agent skills that take an AI short drama from a source novel to episode scripts, a storyboard, image and video prompts, confirmed generation, a rendered cut and a review, with five Markdown documents per episode as the only creator-facing record.

Screenshot of Drama Skills
Editor screenshot, 30 Sep 2026Drama Skills ↗

What it is

A suite of eleven agent skills for making AI short dramas and motion comics, installed into Claude Code, Codex or any assistant that reads the Agent Skill format, and driven by writing Markdown files in a project directory. One skill initialises projects and routes work; the other ten cover source-novel analysis and adaptation, story development, episode scripts, asset decisions, image prompts, storyboard and frozen keyframes, video prompts for a named model, production through provider adapters, editing into a finished film, and review. The defining constraint is that an episode keeps at most five Markdown documents and no parallel database: the screenplay, the visual bible, the storyboard, the image prompts and the video prompts are what the creator reads and edits, and everything else on disk is a record a script owns. Cross-shot consistency is written as checkable facts rather than hoped for. A look that has to survive across shots becomes a continuity lock whose phrase can be pasted into a prompt verbatim; every shot states which bible entries it relies on; and reference images are sorted into the ones that exist and have been checked, the ones planned, and the ones still missing. Generation costs money, so production prints exactly what a batch will contain and calls nothing until the creator confirms. A local dashboard serves the same files from a Python standard-library server bound to loopback.

Who built itThe organization account that owns and maintains the repository, with five sibling projects listed in the README. The commits come from one GitHub account: worldwonderer holds 329 of the 332, MrChenfafafa three, while the author fields alternate between pitechen, PiteChen and worldwonderer and 324 commits share one email address. 202 commits carry a co-author trailer, all of them naming a Claude model, and the README says the suite was distilled from more than a thousand short-drama projects the team has run since 2025.

How it is put together

The parts · 6

The organising idea is that the durable artefact of an AI video pipeline is a contract rather than a program. Every stage is an independently installable skill; the creator-facing state of an episode is at most five Markdown documents that a person owns and edits; each rule the suite holds is graded by who may enforce it, so mechanical facts are checked by standard-library Python while judgement is left to a reviewer that has to cite evidence; and anything that spends money stops at a preview and waits for an explicit confirmation. That split explains the shape of the repository: the prose lives in reference documents inside each skill, the enforceable part lives in checkers that ship with their own selftests, and the local dashboard reads the same five files through a loopback-only server instead of introducing a second source of truth. It also explains what is deliberately absent — no database, no hosted service, no GPU, and no second copy of project state for an interface to drift away from.

skills/ — eleven folders
Eleven installable skills, each shaped the same way: SKILL.md, references/, scripts/, assets/ and agents/openai.yaml. A twelfth skill sits outside the install path under maintainers/skills/short-drama-knowhow.
skills/short-drama/scripts/ and assets/dashboard/
The project layer. project_tool.py (101,685 bytes) initialises projects, routes stages and exports a delivery directory with a manifest and checksums; creator_markdown_check.py (76,402) validates the five documents against one another; creator_views.py (36,214) and dashboard_server.py (83,782) serve the local workbench, whose front end is a 133,200-byte app.js, a 65,157-byte stylesheet and a token file shared with the same team’s writing desk.
The reference corpus
Fifty-odd prose documents spread across the eleven skills, from script-craft.md at 49,092 bytes down to two-thousand-byte fragments: fifteen genre cards, six production-form cards, per-model dialects (minimax-h3.md 20,093, seedance-2.5.md 9,922, wan-3.0.md 7,120) and a target-model profile that says which craft rule applies when a model behaves unusually.
.agents/notes/ and AGENTS.md
Twenty-six decision notes in 152 KB, filed under proposed, implemented or rejected and under six categories, following the DeepSeek Harness agent-notes convention. AGENTS.md states four rules: search for an existing note before a non-trivial change, write Problem, Decision, Alternatives considered and Consequences, never rewrite a note into the opposite decision, and keep no index because the directory is the index — search with rg --hidden.
tests/
Thirty files and 850 KB, from test_dashboard_server.py at 125,415 bytes and test_creator_first_golden.py at 82,703 down to test_voice_direction.py at 2,400: 697 tests by v0.8.1 on Python 3.9 with ruff, mypy and node; a dedicated test_guards_bite.py that checks the guards can fail; shipping-boundary tests; and a release-gate test that keeps maintainer material out of the public tree.
evaluations/ and examples/
The evidence layer. An 18,319-byte probe log, a 55,135-byte content-quality gate with sixteen genre cases and a rubric, a reference run that carries a 147,010-byte source novel through twenty chapter extracts, and two kept projects: an eight-episode golden project of 136 files and 710 KB used as a regression fixture, and the public sample episode the README links to.

Choices, and what they beat

  • Grade every rule by who can enforce it over one undifferentiated list of good practice

    The note of 2026-07-16 defines four tiers — structural invariant, reviewed invariant, craft default, taste option — so a script may block on a missing reference while a numeric preference stays the creator’s. The story-development document states the consequence plainly: a numeric form constraint can be recorded as a creator choice but cannot prove quality.

  • Five Markdown documents as the only creator-facing state over a database or structured store that the interface reads

    The README says there is no parallel database and that changing a file is changing that layer of a decision. The workspace design document repeats it as a delivery constraint and adds that the dashboard classifies the same five file names and derives episode progress from the visible files.

  • Preview and confirmation ahead of any paid generation over calling provider interfaces as part of writing prompts

    The README says the confirmation was placed before production on purpose: the production skill shows the exact count, content, references, parameters, output and adapter for a batch, and provider credentials stay outside the project. Pull request 185 showed the same rule under pressure — when an adapter does not declare support for the audio binding it was given, the job fails before the paid call rather than being submitted with the binding silently dropped.

  • Keep model measurements in the repository as rows with their limits over publishing them as settled guidance

    The note of 2026-09-08 files the probe log in the tree, and the runs name their own limits — three runs per arm, one shot, one execution path. The README’s FAQ repeats the boundary: the sample is small, so the result has to be checked after generating.

  • Replace the single-page dashboard rule rather than stretch it over adding stage views to a design that forbade them

    The product principle of 2026-08-06 allowed one page, one always-visible reading area and no engineering mode. The note of 2026-09-27 retires that, and the pull request says the new note supersedes the earlier one and links to it, instead of the old text being edited into agreement.

Read fromDESIGN.md (Short Drama Creator Workspace Design, 5,266 bytes), AGENTS.md, skills/short-drama-develop/references/episode-design.md with its 2026-07-16 decision note, skills/short-drama/SKILL.md, evaluations/README.md, evaluations/model-behavior-probes.md, README.md, and the complete 558-file tree with sizes.

Build log

6 stages
  1. 01

    Ninety days, eighteen releases, and one name on nearly every commit

    The repository was created on 2026-07-16 and its first commit landed the same evening at 22:50: “Open-source release: short-drama Agent Skill suite”. By 2026-09-28 it had collected 332 commits — 40 in July, 208 in August, 84 in September — and eighteen releases, from v0.1.0 on 2026-07-26 to v0.8.1 on 2026-09-28, with a nineteenth tag, v0.3.0-rc.1, that never became a release. The titles carry the schedule: v0.4.0 shipped ten independently installable skills, confirmed production and an eight-episode golden sample; v0.6.0 was the creator-first five-document workflow; v0.7.0 added the editing skill. Around it sit 2,413 stars, 520 forks, eight watchers, one open issue and 558 files in about 74 MB, MIT licensed and counted by GitHub as Python. Two accounts contributed: worldwonderer holds 329 of the 332 commits and MrChenfafafa three, while the author fields alternate between pitechen, PiteChen and worldwonderer and 324 commits come from a single address. 202 commits carry a co-author trailer, and every one of them names a Claude model — 111 Opus 5, 65 Opus 5 with a million-token context, twelve Opus 5.5, nine Fable 5.1, two Opus 4.8, two Fable 5 and one Opus 5.5. The README says the suite was distilled from more than a thousand AI short-drama projects the team has run since 2025 and nearly 80,000 lines of an in-house tool that had become unmaintainable.

  2. 02

    Eleven skills, one shape, and a Python file behind every rule

    Everything under skills/ is one of eleven folders, and the folders look alike: an SKILL.md, a references/ directory of prose rules, a scripts/ directory of checkers, an assets/ directory of templates and example records, and a small agents/openai.yaml that is what makes the same folder visible to Codex. short-drama initialises projects, routes work and serves the dashboard; the other ten are the pipeline stages. Installation is one command — npx skills add zenstory-ai/drama-skills -y -g — or a symlink of the folders into ~/.claude/skills or ${CODEX_HOME:-$HOME/.codex}/skills, and the README insists that each folder is an independent unit, which a decision note dated 2026-07-26 records as a rule: a skill may not reference another skill’s files. The weight inside the folders is code rather than prose. edit_tool.py is 130,747 bytes, project_tool.py 101,685, production_tool.py 93,209, dashboard_server.py 83,782, creator_markdown_check.py 76,402 and provider_adapters.py 70,718, spread across thirty-one script files, each of which ships a selftest.py beside it. A twelfth skill sits outside the install path under maintainers/skills/short-drama-knowhow, carrying the internal method — blind forward evaluation, cards and coverage, a promotion ledger — that the public tree does not install.

  3. 03

    Five Markdown documents, and the records underneath them

    An episode keeps at most five Markdown documents and there is deliberately no database beside them: screenplay, visual bible, storyboard, image prompts and video prompts, plus a cutting sheet once the film is assembled. What that costs shows up in the checker. creator_markdown_check.py reads the five files and reports a cause rather than a verdict, and the README demonstrates it by breaking the public sample in two places — a frozen keyframe prompt that names a character the shot’s stated visual basis does not cover, and a storyboard duration of two seconds that the video prompt writes as three — producing two error lines that name the missing character and the mismatched seconds. Underneath the Markdown the run keeps structured records: the sample adaptation’s reference run holds shots.jsonl at 42,111 bytes, motion-specs.jsonl at 65,244, keyframes.jsonl at 43,960 and image-prompt-specs.jsonl at 18,586, alongside characters, locations, location views, looks, props, prop states, occurrences, continuity deltas and the creator’s recorded decisions. The consistency claim is a file fact rather than a hope: a look that must survive across shots becomes a continuity lock whose locked phrase can be pasted into a prompt verbatim, and reference images are sorted by state instead of by file name, so a planned slot is not mistaken for an image that exists.

  4. 04

    Every rule says who is allowed to enforce it

    Advice that cannot be argued with is not much use, so every rule carries one of four tiers: structural_invariant, which a script can check and block on; reviewed_invariant, where the reviewer has to cite evidence; craft_default, which usually strengthens an episode and can be overridden with a stated reason; and taste_option, where hook shape, arc shape and how much to leave unsaid belong to the creator. The note filed on 2026-07-16 calls this rule tiers and script boundary, and the reference states the line: a numeric form constraint may be recorded as a creator choice, but it can never prove that a story is good. Numbers therefore live in the project rather than in the suite. The develop stage proposes a rhythm profile and the creator accepts or rewrites it; only then is it written into the project file as accepted, and write, storyboard and review check arithmetic against accepted values only. The proposing table has nine fields for each production form — first hook by second five, no more than thirty seconds between emotional contact points, end on the peak, voice-over at most 0.3 of spoken characters, a target average shot of 2.5 seconds for live action or 3.0 for motion comics, close shots at least 45 percent, and the first major payoff by episode one — and several of those values are marked as unmeasured starting points rather than measurements.

  5. 05

    The model is the thing being measured, so they ran it

    Most of what this suite believes about video models is written as a row in evaluations/model-behavior-probes.md, an 18,319-byte log of real generations. Pull request 176 describes the method: about 170 videos across MiniMax H3, Seedance 2.5 and Wan 3.0, each group running a baseline arm first, the artifacts stripped of identity before blind judging, and seventeen rows appended before the conclusions went back into the dialect files — a nine-second container of three cuts with one keyframe per cut landed its cuts within 0.12 seconds and used a third of the generations; Wan 3.0 speaks at roughly 4.6 Chinese characters per second and truncates what does not fit; and a character whose inner monologue should be silent kept the mouth closed in only four of twelve runs, so the guidance moved to keeping the mouth out of frame instead of trusting a phrasing. Pull request 186 added the next round. A reaction-shot recipe scored 3.9 of 5 for clarity and 4.0 for naturalness; a first scene’s brightness spread fell from 18.3 to 6.6 and its warm-cool spread from 9.3 to 4.2 once each scene got one light sentence and a tone-setting frame; and binding a character’s other line as a voice reference lifted similarity from about 0.3 to 0.69 on Wan, though on Seedance it worked for only some characters. The limits travel with the numbers — three runs per arm, one shot, one execution path — a boundary the README’s FAQ repeats.

  6. 06

    What the community reported, and what review found

    The recon read thirty issues and pull requests in full, from 157 to 187 plus issue 168. The issue is a user reporting that a generated video gave the wrong on-screen face the dialogue; the answer measured nineteen clips, found every spoken segment in the 300 to 370 Hz child range — the voice had been right — and shipped two prompt rules instead of a promise. Pull request 157 runs the other way: a finished episode went to an external Codex review, which found three acceptance blind spots the author’s own measurements had missed. The opening seven shots had never been produced and nothing objected; in-frame text was on no acceptance list, so a line the model invented sat on the final beat of the episode; and the colour-matching pass overshot so far that a segment at minus 15.9 became plus 8.7, the bluest in the film, while the film-wide range looked better than before. The repairs became capabilities: unused shots must now be listed with their numbers and a reason, and since v0.7.1 a written range is rejected outright — pointed at the same cutting sheet, the rule named eight shots. Other pull requests repair the prose: one removes guidance that had hardened into a cross-model prohibition, and the README was rewritten around real outputs and ten questions taken from issues, checked by three rounds of adversarial review with sixteen agents. The tree keeps its disagreements too: a rejected dashboard proposal of 13,546 bytes is still there, and the notes forbid rewriting a note into its opposite decision — when stage views replaced the single-page dashboard on 2026-09-27, the new note says so and links back. Test counts over the same months ran 568 to 697, with a Windows job in continuous integration.

Adjacent records

All records →