Skip to content

Skill Recorder

A macOS-first Electron desktop app that records a real work session — screen video, app and window switches, browser URLs, clipboard previews, an optional session-owned terminal and spoken narration — then hands the timeline to the bundled GitHub Copilot CLI, which reconstructs one overall intent plus an ordered list of steps and turns the approved version into a SKILL.md procedure or a scheduled automation for Microsoft Scout, Microsoft 365 Copilot Cowork or Copilot Studio.

Screenshot of Skill Recorder
Editor screenshot, 30 Sep 2026Skill Recorder ↗

What it is

A desktop recorder that turns a performed task into a reusable agent skill. On macOS and Windows 11 (x64 and ARM64) it captures low-frame-rate screen video, app and window switches, browser URLs, clipboard previews, an optional session-owned terminal with full command output, and optional spoken narration transcribed on the machine by Whisper. Nothing leaves the computer until the user presses Analyze, which sends the timeline to the bundled GitHub Copilot CLI to reconstruct one overall intent and an ordered list of steps. From an approved analysis it builds a Skill — a SKILL.md procedure rendered from a reviewable plan — or a scheduled Automation, for Microsoft Scout, Microsoft 365 Copilot Cowork or a generic skill-capable agent, with Copilot Studio still disabled. A secret-and-PII detection pass runs on device before anything is sent, and skill creation prefers the target agent’s native tools, such as the gh CLI, over replaying the recorded clicks. It ships as a source release: an installer builds the exact tagged commit on the machine.

Who built itA Microsoft repository rather than a personal one, released under Microsoft’s MIT license and CLA process. Its 164 commits come from two people: Giorgio Ughini, 78, and Adi Leibowitz, 81 spread over three email addresses and two linked accounts; a colleague added two and Dependabot two. 120 commits carry a co-author trailer and 100 of those say only “Copilot App”, so most of the history was written with Microsoft’s own coding agent in the loop.

How it is put together

The parts · 6

A local-first desktop application with two agent stages inside it. Electron hosts the recorder and a React renderer; capture, session storage, frame extraction and narration transcription all happen on the machine, and a detection pass for secrets and structured PII sits on the one seam that leaves it, so outgoing text can be masked and on-screen values blurred before any Analyze call. The captured signals are cheap operating-system events first — app switches, window titles, URLs, clipboard changes, terminal commands and full terminal transcripts — with low-frame-rate video as opportunistic enrichment that the describer pulls only where those events are ambiguous. That describer and the two builders are GitHub Copilot CLI agents whose briefs are plain, human-editable string exports in the repository rather than packaged skill directories, and each brief is paired with a Zod contract in common/ that fixes what the agent may submit. What a target can actually do is declared per architecture in a manifest with its own validator: Scout has skills and automations, Cowork has skills only because its catalogue has no browser automation, a generic agent skill is export-only, and Copilot Studio is present but disabled. The hard parts of the program are therefore boundaries — what stays local, which tool a generated step may reach for, and whether the file the user gets matches the plan they approved.

common/
Twenty-eight shared contract and logic files, 140 KB: the analysis and skill-plan types, the value substitution used by {{id}} tokens, the automation and bundle models, a 25 KB IPC surface, a 12 KB sensitive-content module, correlation and description helpers, the architecture manifest — and a unit test beside almost every one of them.
electron/
The main process and the two agents. Recording: a 24 KB recorder controller, a 31 KB audio recorder with its own capture preload, a 13 KB video recorder, a 16 KB frame extractor and its correlator, and window, URL and clipboard collectors. Analysis: the describer (a 13 KB brief, 22 KB of tools, a 17 KB implementation) and the two builders, skill (19 KB) and automation. Also terminal (node-pty, shell discovery and integration), five copilot-signin* modules each with its own test, and three architecture catalogues for Scout, Cowork and agent skills.
src/
Twenty renderer files, 261 KB, of which 68 KB is one stylesheet: a 65 KB session Library, a 30 KB Recorder, a 29 KB plan editor, 19 KB recording controls, plus the SKILL.md review modal, the sensitive-review screen, the recorded terminal, a what-gets-captured panel, and small state modules for placement and review that carry their own tests.
evals/
Four harnesses against the parts with real variance — describer scenarios and their mock pages, the automation builder, the skill builder and the on-device redaction pipeline — each with a rubric, most with a deterministic scorer and one optional model judge, plus an evals/tsconfig.json that the root TypeScript configuration excludes on purpose.
scripts/, install.ps1 and install.sh
The packaging and compliance layer: a 79 KB licence and redistribution checker with a 39 KB test, the reviewed-Electron download and launcher, Windows dependency installation with its own tests, lockfile portability, packaged-artifact verification, three small assertion guards for bundle boundaries and architectures, and the two commit-pinned source installers (34 KB of PowerShell, 20 KB of shell).
.github/
Three workflows — Windows on x64 and ARM64, Non-Windows on macOS and Ubuntu, and a hand-dispatched portable build — with every action pinned to a full commit SHA and Dependabot configured weekly with a seven-day cooldown. Beside them, copilot-instructions.md redirects agent pull requests to the current release branch.

Choices, and what they beat

  • Source-only releases as the default channel over attaching installers, portable binaries or built output to every tag

    RELEASING.md is explicit that dist/, dist-electron/, node_modules/ and anything assembled by the installers must never be attached, and that the release archives are not installers. Binary publication is a separate decision with its own obligations: a version-matched compliance bundle, a build on each native platform and architecture, published SHA-256 values, and a statement of whether the binary is signed at all.

  • Prefer the target agent’s native tools over the recorded clicks over replaying the user interface the recording contains

    Both output kinds are built to reach for gh, web_fetch or a device CLI before simulating clicks, and the instruction document orders it directly: prefer a first-class CLI on the device, above all GitHub via gh, and fall back to browser automation only for genuine UI-only steps. The builder eval suite exists because that preference was violated once — the builder chose Playwright over gh — and the catalogue was fixed so the suite would pass.

  • A reviewable plan before any file is written over generating SKILL.md in a single turn

    propose_plan stops for review and submit_skill follows only after approval, one proposal per turn. Release 0.7.0 extended the same principle to the artifact: the optional review renders the finished file read-only, writes nothing, and lets Add or Export install those exact bytes without another model turn, while an edit to the plan invalidates the preview instead of silently rendering stale text.

  • Fixed values as tokens substituted at render time over inlining the literals the recording happened to contain

    A canonical URL, repo slug or path that never moves is declared once with a label and referenced as {{id}}, so the user edits it in one place and it substitutes everywhere. The brief draws the boundary from the other side too: something that varies between runs must not become a value, and one machine’s path must not be pinned just because the recording used it once.

  • Pin the Copilot CLI and disable cached auto-update during sign-in over accepting whatever version the auto-update cache holds

    The bundled 1.0.71 rejects --web-flow, so the dependency was pinned to a reviewed 1.0.78 and run with --no-auto-update login --web-flow to make browser OAuth deterministic. The accompanying regression test drives the real PKCE and loopback flow with a fake browser that denies authorization, which tests the handshake without contacting GitHub or signing anything in.

Read fromREADME.md (13,195 characters), RELEASING.md, evals/README.md, electron/skillbuilder/instructions.ts, common/skill.ts, common/architecture-registry.ts, common/analysis.ts, electron/describer/instructions.ts, .github/copilot-instructions.md, .github/dependabot.yml, package.json, the three files under .github/workflows/, CONTRIBUTING.md, docs/future-features.md, and the complete 254-file tree with sizes.

Build log

6 stages
  1. 01

    Ten weeks, thirteen releases, and work that lands on a release branch

    The repository was created on 2026-07-29 — its first commit is dated 2026-07-24 — and has taken 164 commits: 107 in July, 34 in August and 23 in September, spread over thirteen releases, from v0.1.0, v0.2.0 and v0.2.1 published one after another at 19:50 on the same evening to v0.7.0 on 2026-09-21. Three of them landed on a single day, 2026-09-16 — v0.6.0 at 08:49, v0.6.1 at 12:06, v0.6.2 at 13:27 — which is the cadence of a team that publishes a patch as soon as a fix is verified instead of batching work. The unusual part is where a change lands: .github/copilot-instructions.md tells coding agents that pull requests must target the active release/<semver> branch and not main, because main trails the release branch and only advances when a version is cut and merged. That file still names release/0.3.0 as current, which is what a document edited once per release looks like when nobody revisits it. RELEASING.md then spells out the mechanics: a release pull request bumps package.json and package-lock.json without creating a tag, the annotated tag is created only after the resulting main commit passes every workflow, and the release notes must carry the full commit SHA plus the SHA-256 of install.ps1 and install.sh. Around all of it sit roughly 4,168 stars, 440 forks and 41 open issues.

  2. 02

    Two names in the history, and a trailer that reads Copilot App

    The contributor list has six entries, but the work is really two people: Giorgio Ughini, 78 commits across two email addresses, and Adi Leibowitz, 81 commits across three addresses and two GitHub accounts (adilei and adilei-powerapps). A colleague, Dan Fiedler, contributed the commit that pinned the GitHub Actions to full-length SHAs, Ramakrishnan Raman added two, and Dependabot two. What makes the attribution worth reading is the trailer census: 120 of the 164 commits carry a co-author line, and 100 of those say only “Copilot App”, with 17 for adilei, two for the dependency bot and one for Ughini himself. So the repository declares, per commit, that most of this history was written with Microsoft’s own coding agent in the loop. Outside contributions arrive and then wait: two pull requests adding Simplified Chinese localization (#72 and #94, different authors) are both still open, as are the skill-runtime eval proposal (#70, open since 2026-08-24, where a maintainer asks “so, what is this exactly doing?”), an audio stop-flush fix (#93) and a set of Windows launch helpers (#67). Each attracts the same CLA bot comment, and the author of #68 writes that the Windows and Non-Windows workflow runs sit at action_required until a maintainer approves them.

  3. 03

    A plan first, a file second

    The Skill Builder is a two-phase conversation, and its brief says so in a heading: “Two phases — never skip the plan.” The agent reads the approved analysis, calls propose_plan with its generalization, the fixed values it intends to hard-code and the ordered steps, then stops; the user answers in natural language, the plan is revised, one proposal per turn. submit_skill, the call that writes the artifact, follows only after a message says the plan is approved. The plan is typed rather than prose: common/skill.ts keeps a Zod schema in which every step is a calculation (reads, derives, decides) or an action (submits, sends, creates, deletes), and a literal that is the same on every run becomes a value with an id, a label and the exact string, referenced from step text as {{id}}. Rendering is deliberately literal: renderSkillMarkdown writes YAML frontmatter with the kebab-case name, a description emitted through JSON.stringify because that yields a valid YAML double-quoted scalar and a colon or comma therefore cannot break the file, an optional allowed-tools list, then the body with every token substituted. Release 0.7.0 added the other half of that bargain: an optional read-only Review SKILL.md renders the complete file without writing anything, serves those bytes to Add or Export without another model turn, and is discarded when the plan is edited.

  4. 04

    The regression that produced an eval suite

    evals/ exists because the interesting failures are not in the recorder. The describer harness materializes a synthetic session (session.json and events.jsonl) into a temporary sessions root, runs the real pipeline and the real Describer, and scores the result against a rubric, with no video and none of the flakiness of live capture — fifteen to twenty-five seconds a scenario. Scoring is deterministic and model-free: the intent must name the right subject, the step count must sit in a range, the expected applications must appear, the key actions must form an ordered subsequence, and recorder bracketing, permission dialogs and tracking-parameter hops must not become steps; a forbidden hit fails outright, otherwise eighty per cent of checks pass. An optional --judge flag adds a Copilot grader rating faithfulness from zero to five, off by default. The builder harness was created after a real regression its README states plainly: when generalizing GitHub work, the builder preferred driving the browser with Playwright over the gh CLI, even though Scout runs on the user’s own machine where gh is installed and authenticated. Ten scenarios now pin the native capability, two of them the gh-versus-browser case directly, and that suite drove the fix in scout-catalogue.ts. A third harness scores the shape of the proposed skill plan, a fourth guards the on-device redaction pipeline, and that one’s real-image OCR variant self-skips with exit code zero when fonts, weights or network are missing.

  5. 05

    One npm test, three workflows and a 79 KB compliance script

    Tests run on Node’s own runner with --experimental-transform-types plus a small hook that maps extensionless imports to .ts and swaps the single electron import for a headless stub, so the suite exercises real application source without a bundler. The npm test line names 43 files one by one, and the counts quoted in pull requests climb: 176 tests on 2026-08-23, 225 on 2026-09-15, 237 the next day and 267 on 2026-09-21. Continuous integration is three workflows. The Windows one runs on windows-latest and windows-11-arm, first proving install.ps1 rejects a mutable source reference, then performing a real commit-pinned source installation with the launch suppressed, checking lockfile portability, generating the license inventory, running the tests, building, packaging a native installer and verifying its architecture. Non-Windows does the same on macOS and Ubuntu, where macOS additionally prepares the complete corresponding-source bundle and asserts the platform’s libvips licence text is present. Portable preview builds run only when dispatched by hand, because the comment above them says an ordinary merge should not spend ten to fifteen minutes packaging. Around that sits a compliance apparatus larger than most of the application: scripts/compliance.mjs is 79 KB with a 39 KB test, package.json carries exact-version allowScripts approvals, and RELEASING.md requires the terms of the Copilot SDK, Electron, sharp-libvips, ONNX Runtime and Tesseract to be re-read before any of them moves.

  6. 06

    Pinning the CLI, and rewriting the install instructions after a 404

    Two fixes show how much of this project is spent making somebody else’s tool predictable. The first pinned @github/copilot to 1.0.78 and runs it with --no-auto-update login --web-flow, because the previously bundled 1.0.71 rejects --web-flow and an auto-updated CLI in the cache made browser sign-in depend on whatever version a machine happened to hold. Its regression test is unusual: it drives the bundled CLI’s OAuth URL, PKCE challenge and loopback callback with a fake browser that denies authorization, so it never contacts GitHub and never signs anything in. The second is a Windows-only recovery path. On a managed machine where the public registry is blocked and the approved mirror needs no interactive sign-in, npm ci is retried exactly once against the approved registry — only when there is no explicit npm configuration and local device metadata supplies a Microsoft Entra tenant hint, and only for that child process, never by saving npm settings, weakening TLS, skipping integrity checks or changing script approvals. The documentation got the same treatment from the other direction: after users copied a README command containing <40-character-release-commit> unchanged and hit a 404, two pull requests replaced the copyable template commands with a pointer to the release page and three plain-language steps, on the reasoning that the people installing it are not necessarily developers.

Adjacent records

All records →