anything2explainer
A skill for Claude Code and Codex that turns a topic, or a document, into a 1280×720 narrated explainer film, with every frame drawn in code rather than generated and each agent in the parallel build writing one Remotion component. The same author’s video-shotcraft makes product promos out of 157 shot cards and a sound-design pass; this one is driven by the narration, ships no sound effects or music, and freezes the words into frame numbers before any shot is built.

What it is
A skill for Claude Code and Codex that turns a topic, or a document, into a narrated explainer film. It is not a command-line tool: what ships is the method — a 21,395-byte SKILL.md with nine stages and four checkpoints, ten specification documents under reference/, a compilable Remotion 4 project with a primitives and lighting library, eleven scripts covering the pipeline and its quantitative QC, and one finished reference film with its whole paper trail. Output is 1280×720 H.264 at 30fps in Chinese or English, with chapter cards, a progress bar, a top HUD and subtitles aligned to word boundaries: edge-tts zh-CN-YunxiNeural for Chinese, kokoro-82m am_liam for English, or your own audio. A blank line in the narration is a paragraph and one shot, sentences inside it are 10 frames apart and a paragraph end is 30, so the pause lands where the picture changes and every shot holds 1–1.5 seconds after its last element lands. Every frame is drawn in code, and every number on screen has to trace to a source URL. 2,207 stars, 305 forks, PolyForm Noncommercial.
Who built itThe repository is published under the GitHub account Vincentwei1021, which also publishes video-shotcraft, and the README signs off as Vincent. Its 28 commits all fall between 2026-09-08 and 2026-09-18: 25 of them under this account — signed under three names, Wei Yihao on 17, Vincent Wei on 5 and Yihao on 3, but a single address — one by DHCatLaw, and two more signed mdipaolo1 from an address GitHub links to no account. Twenty-five of the 28 carry a co-author trailer and every one of them names a Claude model: Opus 5 with a million-token context on eighteen, Fable 5.1 on five and Opus 4.8 on two. The contributor list holds two entries: the author with 25 and DHCatLaw with one.
How it is put together
The parts · 6A skill wrapped around a template, with the shot as the unit of work and the narration as the source of time. The agent copies the Remotion project once per film, writes one component per shot against a fixed primitives and lighting library, and never touches the persistent layers in src/common/ and src/overlay/. Determinism is a requirement: every animation is a pure function of the frame number with seeded randomness, and text fitting is computed rather than measured in the DOM, so re-rendering produces identical frames. reference/ carries the style guide, the motion vocabulary with its frame counts, the composition and light rules, and the shot-pattern table. frame_metrics.py then samples a 320×180 content area and calls a frame static when greyscale changes by less than 0.35, with a budget of at most 40 percent static frames per shot; motion_check.py applies the same test per group and to the finished film, where an extra column separates true stillness from small-area motion. The group reading is not the verdict: the fourth film passed all 48 shots at group level and still had 11 over the limit in the finished cut. The look is one style on purpose with a single switch, the backdrop, because a theme system would be something the shot code could not rely on.
- SKILL.md and reference/
- The agent-facing half: the process document at 21,395 bytes and, beside it, about 123 KB of specification for the main session and the parallel agents —
lessons.md(32,865),narration-storyboard.md(21,950),agent-build-rules.md(13,343),composition-and-light.md(11,806),style-guide.md(10,762),motion-vocabulary.md(10,373),prompts.md(8,421) with six prompt templates,narration-guidance.md(7,121),agent-qc-rules.md(6,344) andresearch-brief.md(3,477). - template/
- The compilable Remotion 4 project, copied per film by
scripts/new_project.sh: 39 TypeScript files and 128 KB undersrc/, whereui.tsx(40,704) holds the primitives and palette,fx.tsx(17,209) the light, depth and camera work,config.ts(4,070) the language, backdrop and credit switches, andoverlay/Overlay.tsx(14,062) the title, chapter cards, HUD, pipeline rail and ending.src/common/adds fourteen files — fog, star field, dot-field wave, glitch, easings, subtitles, the progress bar, the footage layer, computed text fitting, and the timeline and subtitle tables — andsrc/shots/G1toG8are eight directories waiting for the shot components of the next film. - template/scripts/
- Eleven Python and shell scripts, 60 KB.
tts_build.py(29,185) parses the script, drives four TTS engines and writes the timeline and subtitle tables;frame_metrics.py(9,585) scores the finished film;motion_check.py(8,071) does the same per group and per film;selfcheck.py(6,310) andsheet.py(1,637) check the project and assemble frame sheets; andnew_project.sh(1,438),render_storyboard.py(1,311),preview.sh(1,141),still.sh(1,133),render.sh(989) andtest_render.sh(732) carry a film from script to frames. - examples/rag/
- One film’s entire paper trail, 119 files and 3,920 KB:
research.md(48,851) with a source for every number,storyboard_src.md(35,589) and a Chinese storyboard (33,529),narration.txt(4,486),timeline.md(6,367),ui_rag.tsx(19,711), delivery notes (4,664), nine shot directories from G0 to G8 holding 44 shot components with a build note of 12,401 to 15,707 bytes each, five QC reports of which the largest is 23,532, twenty reference frames and six overview frames. - examples/contrast/
- Six bad-and-good frame pairs with a 1,477-byte note and a 408,424-byte contact sheet: the yardstick for composition and light that the QC agents are told to measure against.
- Repository root
README.md(17,628 bytes) and its Chinese counterpart (16,380),SKILL.md, a 5,048-byte PolyForm Noncommercial 1.0.0 licence, a citation file (1,124) and a 591-byte sample narration for the template, with four bundled fonts — one of them a 17,772,300-byte Chinese face — under their own SIL OFL 1.1 licence.
Choices, and what they beat
Draw every frame in code on a deterministic canvas over generating the footage with a video model
The README’s comparison table draws the line in one sentence: generative models synthesize footage from a prompt, while here it is “deterministic code, not pixels”, every number on screen traces to a source URL, and any frame can be fixed by editing one shot file. The FAQ supplies the mechanism — every animation is a pure function of the frame number with seeded randomness, and text fitting is computed rather than measured in the DOM, so re-rendering produces identical frames. Why Remotion rather than another programmable canvas is the one question the record cannot answer: issue 14 asks it, with a HyperFrames experiment attached, and has no reply.
Ask the TTS engine for word boundaries explicitly over letting subtitle block starts fall back to interpolation by character count
Measured rather than argued: on the repository’s own sample narration the interpolated block starts were frames 0, 9, 45, 81 and 134, against real speech positions of 0, 6, 30, 76 and 130. The pin on
edge-tts==7.2.8carries its reason in a comment — it tracks a Microsoft endpoint and breaks across upgrades, and from 7.2.0 the boundary mode has to be requested — and asking for it explicitly is the fix that shipped in place of the contributor’s engine swap.Lock the narration before any shot is built over re-timing the film when the script changes
The second checkpoint states the cost: once the voiceover exists, the frame numbers are hard-coded into every shot file, so changing one word re-times the whole film. That is why the sign-off before voiceover is described as the cheapest place to intervene, and why the known limits call the words frozen once they are voiced.
Let the finished film run 5 to 8 percent longer than the speech over cutting the picture to the length of the voice track
The pause budget is deliberate: sentences inside a paragraph are 10 frames apart and a paragraph end is 30, so the silence lands where the picture changes, and every shot holds 1 to 1.5 seconds after its last element lands. The rule was extended after the fifth film, when viewers said the picture left the moment an animation arrived; the sentence gap went from 10 frames to 20 and the extra length was accepted as the price of the hold.
One visual style with a single switch over a theme layer the shot code would have been written against
The FAQ says there is one style on purpose, with the backdrop — star field with fog, or the dot-field wave ported from the sibling narration project — as the only switch, and that changing anything else means editing the style guide and the primitive library rather than the shots. The known limits restate it as a property of the project rather than a gap in it.
Strip the sample film out of the specification over shipping the skill tuned to the film that proved it
An audit pull request found one film’s habits written into the method: a fixed chapter formula that treated explaining a technology as the structure for every subject, a shot-pattern table whose left column was domain concepts such as vector space, index, prompt output and evaluation metric, and a pronunciation table full of one film’s CUDA and NIXL terms. The formula became “structure follows the main line”, the table’s column became topic-independent visual relations, and the pronunciation table was emptied.
Read fromREADME.md (17,628 bytes) — in particular the output-spec table, the length table, the four checkpoints, the comparison table, the FAQ, the originality rules and the known limits; the file sizes of SKILL.md (21,395), the ten documents under reference/ (3,477 to 32,865), template/scripts/tts_build.py (29,185), frame_metrics.py (9,585), motion_check.py (8,071), template/src/ui.tsx (40,704) and template/src/fx.tsx (17,209); the bodies of pull requests 1, 2, 3, 4, 7, 9, 10, 15 and 17 and the comment threads on pull requests 1, 10 and 15 and on issues 5, 6, 12 and 16; and the counted contents of examples/rag/ and examples/contrast/.
Build log
6 stages- 01
The same author, the same canvas, a different kind of film
Two records in this archive now come from the account
Vincentwei1021. video-shotcraft was created on 2026-07-19 and has 10,029 stars; it exists to make a product look good, and its centre of gravity is a library of 157 shot recipe cards plus a sound-design pass over 149 effects and five music beds. anything2explainer was created fifty-one days later, on 2026-09-08, and has 2,207 stars; it exists to explain a subject, and its centre of gravity is a script. The two share the machinery and almost none of the content: both are skills for Claude Code and Codex, both draw React components through Remotion onto a black canvas, and both are documented at the same length. What differs is what decides the picture. In the card library the shape of the film comes from the shots you pick and the sound you pin to them, and the product is the subject. Here the shape comes from the narration, which is researched, written, voiced and turned into a frame-accurate timeline before a single shot exists — and the shot files then hard-code those frame numbers. There is no sound-effect library here at all, no music bed, no card catalogue and no browser editor. The README draws the line itself: the output is motion-graphics diagrams that show the mechanism, with the narration driving the visuals. - 02
Nine stages, four stops, and a repository that is mostly method
The entire history is ten days: 28 commits between 2026-09-08 and 2026-09-18, and the last one adds the end credit
built by Anything2Explainer skill. Around that sit 2,207 stars, 305 forks, eight watchers, four open items and a 209-file tree. The repository is explicit about what it is not — “It is not a CLI” — and what it carries instead is a process:SKILL.mdat 21,395 bytes runs nine stages from scaffolding to QC, and stops at four checkpoints where the run waits for you. Ten documents underreference/add about 123 KB of specification, led bylessons.mdat 32,865 bytes andnarration-storyboard.mdat 21,950, down toresearch-brief.mdat 3,477. Undertemplate/sits the compilable Remotion project:src/ui.tsxis 40,704 bytes of primitives and palette,src/fx.tsx17,209 bytes of light, depth and camera, and eleven scripts of 60 KB whose centre istts_build.pyat 29,185. The heaviest file in the tree is not code:NotoSansSC.ttf, one of four bundled fonts, is 17,772,300 bytes, and the shell scripts are zsh plus Python 3, verified on macOS, with Windows untested. The reception includes one third-party test rather than only stars: the AutoClaw agent workspace reports installing the skill and completing an English explainer task end to end, with the reviewed shots meeting its stated visual criteria. - 03
The narration is the timeline, and a paragraph is a shot
Pacing here is not a matter of taste; it falls out of the voice track. A blank line in the narration marks a paragraph, and a paragraph is one shot, so the picture changes where the speaker pauses; sentences inside a paragraph are 10 frames apart and a paragraph end is 30. Subtitles are aligned to word boundaries, and the film is allowed to run 5–8% longer than the raw speech so that every shot can hold for 1–1.5 seconds after its last element lands. The second checkpoint is where this becomes expensive: once the narration is voiced, frame numbers are hard-coded into every shot file, so changing one word re-times the whole film. The rule set was rewritten once from real complaints. Viewers of the fifth film said the pacing was too fast and that the purple light sweep kept arriving, and the numbers agreed: of 47 sentences, nine ran under three seconds, and of 40 shots examined in the finished English cut, 26 left no stable frames at all at the end while only seven held for 30 frames or more, and 34 of the 47 shots pushed in. The fix, decided by the user: static time capped at 3 seconds instead of 30 frames, the last beat of a shot holding 30–45 frames before exit, shots at least 120 frames long so that a short sentence becomes a beat inside a neighbouring shot, sentence gaps raised from 10 frames to 20, and at most one highlight moment per chapter.
- 04
Where the word boundaries went
The sharpest piece of forensics in the repository arrived from a user. On 2026-09-09
DHCatLawfiled three problems found while running a four-and-a-half-minute Chinese film on macOS, two of which were real. One was trivial and total:test_render.shcreated its temporary directory with a name ending in a dot, and Remotion refuses an image-sequence output directory with an extension. The second looked like a change at Microsoft: across four voices and both languages,WordBoundaryevents came back as zero, leaving onlySentenceBoundary, and pinningedge-ttsat 7.2.8 changed nothing. The author reproduced all of it and found the cause on the client side instead: version 7.2.0, released on 2025-08-05, added aboundaryparameter toCommunicatewith a default of"SentenceBoundary", so the request went out sayingwordBoundaryEnabled:"false", where 7.0.2 and earlier had hard-coded it true. The failure was silent rather than loud: with no words, subtitle block starts fell back to interpolating by character count. On the repository’s own sample narration, read byzh-CN-YunxiNeuralat plus eight percent, the interpolated block starts were frames 0, 9, 45, 81 and 134 against real speech positions of 0, 6, 30, 76 and 130. The fix asks for word boundaries explicitly, and the contributor’s retry for empty audio was adopted with his authorship kept. - 05
What was merged, what was generalised, and what was turned down
The pull-request queue decided the shape of the project. A Linux and Raspberry Pi port merged after the contributor ran it all on a Pi 5: the default English engine will not install on ARM and Python 3.13, so two local engines were added behind
TTS_ENGINE, andREMOTION_BROWSER_EXECUTABLEpoints Remotion at the system Chromium. A follow-up then audited the skill for generality after finding the sample film written into the method — a fixed chapter formula that treated one subject’s shape as the structure of every film, a shot-pattern table of domain concepts rather than visual relations, and a pronunciation table of one film’s CUDA and NIXL terms — and all three came out. Windows arrived as a report rather than a patch: the pipeline does work, verified on Windows 11, but five shell scripts need replacing, and an open pull request now makes them portable. Two contributions were declined rather than ignored — a spoken title sentence at the head of the film, because the maintainer prefers the existing title shot and timeline, and Russian narration with its tests, as maintenance the main repository would have to carry. The question this record would most like answered — why Remotion rather than another programmable canvas, whether HyperFrames was evaluated, whether the render layer could become a swappable backend — is issue 14, filed with an experiment and still unanswered. - 06
One film shipped whole, and a licence that collects requests
The reference film is delivered with its entire paper trail, which is more than half the tree: 119 of the 209 files under
examples/rag/— a 48,851-byte research document with a URL for every item, a 35,589-byte storyboard source, nine shot groups holding 44 shot components with a build note of 12,401 to 15,707 bytes each, and five QC reports. That cut is the four-minute-35 Chinese version on a star-field backdrop, built by eight agents working in parallel for forty minutes through two QC rounds; its successor runs 4:54 on the dot-field backdrop, and the English cut runs 5:02 over 44 lines and 785 words voiced by kokoro. The visual language is credited rather than claimed: black canvas, white line art, purple accents and ultra-bold headlines were learned from a Douyin creator named in the acknowledgements, and none of that creator’s assets are used. Originality is stated as two rules: every frame drawn in code, and every fact sourced, with unverified material kept off the screen and out of the narration. Commercial use had to be answered three times: a small United States law firm, an education account publishing to Douyin, WeChat Channels, Xiaohongshu and WeChat, and a Turkish YouTube channel with paid Udemy courses. All three were sent to the author’s address with a request for scope and scale, and the README is clear that the videos made with it belong to whoever made them.
Adjacent records
All records →No. 112
video-talkcraft
An agent skill that turns Claude Code or Codex into a motion-design studio for narration videos: hand it a script and a finished voice track and it aligns the two word by word on your own machine, storyboards every semantic beat, then renders the film from 108 motion recipe cards under a camera system defined by subtraction. Where video-shotcraft makes product promos out of 157 shot cards and a sound-design pass, and anything2explainer turns a topic into an explainer whose narration becomes frame numbers, this one takes a voice track that already exists and treats it as the clock.
No. 117
sepia
A portable de-AI writing skill: four operations over one canonical rules file, narrative architecture repaired before word choice on fiction, a thin rule file matched to the venue on professional prose, and every rule labelled as measured, consulted or the project’s own inference.
No. 119
headcount
An agent organization shaped like a company — a chief executive over sixteen independently installable departments and 172 skills, where a skill is a folder of Markdown that loads itself when a request matches its description, one tree installs in both Claude Code and ChatGPT because only the manifests differ, and the 184 outside authorities that settle a question rather than decorate an answer — a regulator, a standards body, primary law — sit in a catalog beside the skills they answer for, each labeled with what an agent may do with it.