video-talkcraft
An agent skill that turns Claude Code or Codex into a motion-design studio for narration videos: hand it a script and a finished voice track and it aligns the two word by word on your own machine, storyboards every semantic beat, then renders the film from 108 motion recipe cards under a camera system defined by subtraction. Where video-shotcraft makes product promos out of 157 shot cards and a sound-design pass, and anything2explainer turns a topic into an explainer whose narration becomes frame numbers, this one takes a voice track that already exists and treats it as the clock.

What it is
An agent skill that turns Claude Code or Codex into a motion-design studio for narration videos. You bring a script and a finished voice track — any synthesised or recorded voice — and it aligns the two word by word on your own machine, with a 767 MB speech model behind sherpa-onnx as the default and a download-free recogniser as the alternative, then writes every semantic beat into a storyboard and renders the film with Remotion from 108 motion recipe cards. Each card is three artifacts: a recipe document, a self-contained component that imports nothing but React and Remotion, and a runnable browser preview that also carries a table of sound cues. A camera system defined by subtraction gives each scene one very slow push or pull and treats a static frame as a defect, and a browser workbench can reopen the delivered film as seven kinds of track. Personal, educational and research use is free under PolyForm Noncommercial 1.0.0, commercial use needs permission, and the videos you make are yours. 1,318 stars, 126 forks, one release that is a media drop rather than a version number.
Who built itPublished under the GitHub account Vincentwei1021, which also publishes video-shotcraft and anything2explainer, and the README signs off as Vincent. The repository was created on 2026-08-22 and its first commit landed six days later, on 2026-08-28. Of its 82 commits, 79 come from this account — 44 signed Vincentwei1021 from the account’s noreply address and 35 signed Yihao from a gmail address — leaving one each from three outside contributors: leksautomate, ltppp and scpcn01vision-oss. Fifty-seven commits carry a co-author trailer, and 49 of those name a Claude model: Fable 5.1 on 33, Opus 5 with a million-token context on 10 and Fable 5 on six; the remaining eight trailers name Vincentwei1021 himself.
How it is put together
The parts · 6An agent skill whose entry point is a method document and whose unit of reuse is a card. The method lives in one 62 KB SKILL.md and thirteen reference documents; the cards live three times over — as prose, as a deterministic component, and as a preview that is the visual truth — so an agent can be told to use a card by name, read what it is for and copy real parameters instead of inventing them. The timeline is not designed, it is measured: a local recogniser fits the script to the delivered voice track, every beat is anchored to a word, and the picture is built on top of those numbers, which is why the pipeline is Python around the audio and Node around the picture. Rendering follows the same logic — one segment per shot with a single continuous audio track, so changing one shot costs one segment plus its neighbours, and a frame-count assertion guards every assembly. Because a static frame is the failure the whole design is aimed at, the anti-slideshow rules are expressed as things a machine can check: a probe that measures frame differences, a signature test that detects a repeated source, a detector for the frame rate of host footage, and a lint that reads a delivered project’s structure before an editor opens it. The workbench is bolted on rather than built in, and it borrows the film’s own shape instead of demanding a second one: it reads the project through an adapter layer, falls back to a stand-in module wherever a project lacks one contract file, and writes parameters back through the same override the renderer reads.
- SKILL.md and references/
- The agent-facing half: the entry point at 62,221 bytes, and thirteen documents beside it in 306 KB —
taxonomy.md(60,834) as the card index,design-language.md(28,256),cinematography.md(22,231),shot-design.md(20,533),shotbook-example.md(16,020),demo-spec.md(15,549),review-protocol.md(13,996),broll-sources.md(13,370),layout.md(10,492),host-footage.md(9,485),schematic.md(8,527) andsemantic-annotation.md(8,155), plus a generatedcards-index.jsonof 85,527 bytes.references/cards/holds the 108 recipe documents in 1,124 KB. - template/
- What gets copied into a film: 108 self-contained components in 961 KB, a motion-system layer of 23 files in 108 KB —
backdrop.tsx(18,449) with twelve animated backgrounds,schematic.tsx(13,913),transitions.tsx(10,783),Subtitles.tsx(7,786),longtake.tsx(7,075) andcamera.tsx(4,452) — and six shared components in 42 KB led bytheme-frame.tsx(16,338), the eight video framings that deliberately stay outside the card count. - demos/ and gallery/
- The visual truth and its shop window: 108 single-file previews, each with a shared shell of 15,284 bytes, a GSAP and a Lottie bundle, a sound-cue table of 37,222 bytes and 1.76 MB of base64 samples, plus about 5 MB of credited stand-in footage and stills. The gallery is one 2.87 MB page with 109 thumbnails in 16 MB and 108 English card texts in 1,236 KB, published from a 2,272-byte workflow to GitHub Pages.
- scripts/
- The gates, in 32 files of 342 KB:
preflight.py(45,891) as the largest of the scripts,voice_trim.py(38,565),render_shots.mjs(24,378),semantic_annotate.py(21,539),timestamps_cpu.py(19,710),pipeline_state.mjs(17,727),tts_fishaudio.py(16,252),workbench_contract_lint.py(14,531),card_match.py(13,016),cards_index.py(12,107),motion_check.py(11,121),freeze_probe.py(8,570) andface_bbox.py(4,595), with two test files beside them. - workbench/
- A browser editor of 192 source files in 1,200 KB: an application shell of 12,601 bytes, a store of 10,834, the import path that reads a delivered film at 19,195, a 21,269-byte Vite configuration that also serves the live dashboard, 23,854 bytes of styles, a timeline, panels, a preview, fifteen stand-in contract modules and a 19,972-byte regression suite. Its guide (17,730) and readme (16,936) sit beside sixteen files under
docs/— a live-pipeline document and fifteen screenshots — in 1,653 KB. - runtime/ and .github/
- One dependency set for every film, in seven files of 198 KB: a 186,728-byte lockfile, a 4,207-byte check that installs, upgrades and smoke-renders one frame with a rollback on failure, a 4,051-byte link script, and a 5,723-byte Python test. A single 2,272-byte workflow deploys the gallery, and
LICENSE(4,895) plusTHIRD_PARTY_NOTICES.md(4,274) carry the terms for the code and for the embedded samples.
Choices, and what they beat
The preview is the visual truth, and the component must match it over writing the component first and letting the preview follow
Stated in
references/demo-spec.mdon the day the project began: the preview is the visual truth, and changing its timing or its picture obliges the same change in the component — the same discipline the sound-cue table carries, where re-timing a preview without re-timing its cues is treated as a defect.Neutral card skins, except when the card is a real product’s interface over one neutral treatment for every card
The visual baseline of 2026-08-23 asks for a white stage, no decorative gradients, grids, noise or vignetting, so a card can be carried into any style. The exception of 2026-08-25 is the interesting one: when the card shows somebody’s product screen, the skin is the content, so the colours, radii, spacing and marks are copied and only the copy is editable — otherwise the card stops reading as that product and loses the reason it exists. On 2026-09-05 the film side gained the matching rule, deriving a style from the domain of the script and writing a skinning line per shot.
Split the recognition input at measured silences over trying to raise what the model can swallow in one piece
The crash was in a fixed-length relative-position table inside the encoder, so the ceiling could not be argued away. The fix cuts inside a 20-to-75-second window at the midpoint of the longest measured silence and adds each segment’s offset back to its timestamps; the change was validated as a single variable on one 193-second recording, returning the same 975 tokens with 98.6 percent alignable, a median difference of 10 ms, and 34 seconds and 2.4 GB instead of 54 seconds and 8.6 GB.
Ask for confirmation after one shot, not after the film over implementing all thirteen shots and then showing the result
The confirmation point used to fall after the whole film was implemented, so a problem of the sample-shot kind — skinning, subtitles, camera amplitude — meant reworking everything. Now the skeleton is built with stand-ins, one sample shot is rendered with audio cut to its own interval, and the run stops and asks once: 47 seconds for a 6.22-second first shot on the test project, against the cost of a film rework.
Annotate the script before consulting the catalogue over filtering the cards by the kind of material a shot contains
The funnel had one filter and nothing had ever recorded what each sentence was doing, so a self-introduction received the lowest-energy text card while a card whose own notes said it was also used for self-introductions sat unused — and three independent reviews passed the shot, because the rubric checks what is on screen rather than what should be. The answer was a per-sentence annotation file with a closed vocabulary, a main-or-subordinate weight and the entities a sentence needs, plus a matcher that returns ranked candidates and, just as usefully, the reasons the others were rejected.
Rotate how a shot is presented rather than adding movement to it over keeping the layout that happened to work
The rule came from an eleven-shot film in which eight shots used the same split, the same badge and the same text column. Presentation now rotates across seven ways of showing a single piece of footage, adjacent shots must differ, no style may exceed a third of the film, and the note attached to the rule says the rotation is a change of composition rather than more motion — every style still holds only the one slow camera push.
Read fromREADME.md (12,055 characters; the reconnaissance report printed the first 6,000 and the remainder was fetched from the default branch on 2026-10-01), the two architecture documents printed in full in that report — references/demo-spec.md (8,104 characters) and references/shot-design.md (10,225) — the complete 928-file tree with sizes and the two-level directory summary, and the bodies of the pull requests and issues cited above, chiefly 14, 16 to 22, 26 to 31, 33 to 36 and 39 to 42.
Build log
6 stages- 01
Three skills, one account, and three different things deciding the picture
Three records in this archive now come from the account
Vincentwei1021. video-shotcraft, created on 2026-07-19, makes product promos out of 157 shot recipe cards and a sound-design pass over 149 effects and five music beds. anything2explainer, created on 2026-09-08, starts from a topic or a document and ends at a 1280×720 explainer in which the narration is frozen into frame numbers before any shot is built; it ships no card catalogue and no sound library. video-talkcraft, created on 2026-08-22 between the two, starts from a voice track that already exists and treats it as the clock. The assets differ accordingly: shotcraft’s motions were read from twelve named product films and de-branded; anything2explainer draws its diagrams in code; this one takes a script, a recording and optionally host footage, B-roll or screenshots, and its card library was partly grown from shotcraft — the batch of 2026-09-05 translated 19 of that library’s cards into this one’s contract, 89 cards becoming 108. Machinery moves between them as well: the browser workbench here was carried over to video-shotcraft on 2026-09-04, and they share dependencies through oneruntime/directory. The licences differ, and a user of two of them said so plainly: shotcraft is Apache-2.0 and raises no question, while this one is PolyForm Noncommercial and prompted a written request for commercial terms. - 02
Twenty-six days, eighty-two commits, and a release that is not a version
The repository was created on 2026-08-22 and its first commit arrived six days later, on 2026-08-28, describing the whole idea in one short line. Six commits fall in August and 76 in September, and the last of the 82, on 2026-09-22, is the merge of pull request 42. The single release,
gallery-media, is dated 2026-08-28 — the same day as that first commit — and tags the gallery’s motion previews rather than a version, so the project has never carried a number. Around it sit 1,318 stars, 126 forks, two watchers and one open item, on about 30 MB whose language GitHub reports as HTML because the gallery is a single 2.87 MB page. The issue and pull-request record runs from 14 to 43 and 30 of those entries are on file: for a project this young the queue was the project, because nearly every change arrived with a body that says what was wrong before it says what was done. Reception came in two forms rather than as coverage. The team behind a commercial agent workspace installed the skill, turned a meeting’s notes into a storyboard, beats and voiceover timing data, and published the deliverable with its report; and a developer who uses both of this author’s libraries asked in writing whether personal work for small businesses could be licensed, and was answered with an email address. - 03
The voice track is the clock, and the clock runs on your own machine
Word-level alignment is the first claim in the README.
scripts/timestamps_cpu.pyfits the script to the audio with a 767 MB speech model behind sherpa-onnx by default, or a whisper backend that needs no download, and the measured figure is a median deviation of 20 to 40 ms per character against a GPU forced aligner, so every motion beat can be pinned to an exact word. The ceiling: one 193-second narration ran in 54 seconds at 8.6 GB of peak memory, while longer inputs all crashed in the same place inside the encoder, where the relative-position table has a fixed length. The fix splits the audio at measured silences inside a 20-to-75-second window and adds each segment’s offset back; re-measured on the same recording, chunking at 75 seconds returned the same 975 tokens with 98.6 percent alignable, a median start-time difference of 10 ms and a worst case of 150 ms, and the run fell to 34 seconds and 2.4 GB.scripts/voice_trim.py, added on 2026-09-10, runs first: it cuts filler words, stutters and over-long pauses out of a real recording and emits an edit list that can cut the same-take camera footage with it; the script is the truth, so a character the recogniser misheard inside it is never cut. The one cloud path is opt-in: a community pull request on 2026-09-15 added a free streaming speech service that returns audio and word-level timestamps in one request. - 04
A hundred and eight cards, three artifacts each, and an audit against the code
The library is the centre of gravity and every motion is stored three times: a recipe under
references/cards/, a self-contained component undertemplate/cards/importing nothing but React and Remotion, and a runnable HTML preview — with the written rule that the preview is the visual truth and the component must match it frame for frame. A fourth artifact is easy to miss: a sound-cue table per card drawn from thirteen timbres, ten of them real samples; a re-pinning on 2026-08-26 cut the cue count from 439 to 269 and capped volume at 0.65. The count has two answers: the account-page description says 109 cards while the README banner and both folders say 108 — the extra thumbnail belongs to the video-frame component. On 2026-09-21 nine subagents compared every recipe document against the component that implements it, which the repository treats as the truth, and returned 29 accurate, 66 with minor deviations and 13 inaccurate: ten cards claiming footage their code never reads, three in the wrong input type. The next day a user found the deeper fault: the funnel had one filter — the kind of material a shot contains — and nothing recorded what each sentence was doing, so a self-introduction got the lowest-energy text card while a nameplate sat unused. Three independent reviews passed that shot, because the rubric checks what is on screen rather than what should be. - 05
An anti-slideshow system defined by subtraction
What the project description calls the anti-slideshow camera system is a set of prohibitions. Each scene gets one continuous camera curve and only two are allowed — a very slow push from 1.00 to between 1.04 and 1.06, or a very slow pull — and on 2026-09-04 ambient breathing, idle micro-motion and the exposure pulse were switched off by default. A static frame is treated as a defect that should be structurally impossible, and it is measured:
scripts/freeze_probe.pyrenders three points per shot, compares frames 0.8 seconds apart with an encoder’s frame-difference test, and flags a shot as frozen below a calibrated threshold, agreeing with the master on 98 percent of 45 sample points. The other half of the discipline is variety, learned from a complaint: after an eleven-shot film came back with eight shots in the same 60/40 split, a single B-roll shot must now be presented one of seven ways, adjacent shots must differ, no style may cover more than a third of the film, and rotation is a change of composition rather than more movement. A text-only shot must carry a line drawing, and a shot holding one video must be wrapped in one of eight framing devices. Pages are shot rather than pasted: a long screenshot is scrolled, toured or magnified while the camera stays still, and screen recording is refused because Remotion on a still is seek-safe and can anchor to a word. - 06
What the gates catch, and what one user’s bill caught
The checking machinery was written after a bad week. One narration film took 7 hours and 21 minutes, produced six full masters, threw five away and lost three reviews; its two worst mistakes, host footage at the wrong frame rate and an empty B-roll folder reaching delivery, had both been documented already.
scripts/preflight.pyturns the entrance rules into assertions that run before work starts: it checks host footage against the film, looks for the repeated-frame signature that betrays a mismatched source, and requires every shot’s material line to exist. Delivery is guarded three ways: a frozen-frame test on the picture, a per-cue energy check on the sound-effects stem, and one independent review. The order of work changed too: instead of implementing all thirteen shots and then showing the result, the skeleton is built with stand-ins, one sample shot is rendered with its own audio, and one question is asked — 47 seconds for a 6.22-second first shot. The sharpest entry is the one nobody has closed: a user reported nine whole-film renders in one night, about 2.5 GB and roughly 10,000 credits, and the investigation concluded that rendering itself was free because it runs locally — the expense was the loop around it, since each export is another round of agent turns that re-pays the conversation, and those renders had bypassed the iteration discipline written into the skill.
Adjacent records
All records →No. 083
video-shotcraft
An agent skill that turns Claude Code or Codex into a motion-design studio: 157 shot recipe cards carrying the real easing and timing values, 214 motion previews, a validated 36.2-second template film, 149 sound effects filed by scene, and a browser workbench that reopens the delivered film for editing.
No. 128
anything2explainer
A skill for Claude Code and Codex that turns a topic, or a document, into a 1280×720 narrated explainer film, with every frame drawn in code rather than generated and each agent in the parallel build writing one Remotion component. The same author’s video-shotcraft makes product promos out of 157 shot cards and a sound-design pass; this one is driven by the narration, ships no sound effects or music, and freezes the words into frame numbers before any shot is built.
No. 119
headcount
An agent organization shaped like a company — a chief executive over sixteen independently installable departments and 172 skills, where a skill is a folder of Markdown that loads itself when a request matches its description, one tree installs in both Claude Code and ChatGPT because only the manifests differ, and the 184 outside authorities that settle a question rather than decorate an answer — a regulator, a standards body, primary law — sit in a catalog beside the skills they answer for, each labeled with what an agent may do with it.