AI Employees
Eight open source AI Employees, each a folder of scheduled routines that covers a whole business role, running on your own machine on the agent you already use: they drive the browser you are signed in to the way you do, draft and stage the work and leave the last click to you, get better at your business on every run, and every file they touch is yours.

What it is
In this project an AI Employee is not a chat window and not a skill: it is a folder of scheduled routines that together cover one business role, and eight of them ship, sixty routines in all. Each kit is about a megabyte of plain markdown plus three small Node scripts, with no framework, no runtime and no service behind it. A member extracts one folder onto their own machine, points the agent they already use at it, and from then on the routines fire at their own times, read pages in a browser already signed in as the member, append dated ledger lines, and reconcile into a single brief each morning. Everything outward-facing is held by default: the draft is written, the form is filled and left open, and the last click is the member's until a row they write themselves in RELEASES.md hands a channel over. The routines name capabilities such as page.read rather than tools, and one file per kit maps each capability to a concrete route on whatever harness the member actually runs, which is what lets the same instruction files run on all of them.
Who built itThe founder of Reinventing.AI, and of the Facebook group Vibe Coding is Life, which he describes as 340,000 members strong. All 52 commits in the repository are his, 50 of them carrying a co-author trailer naming a Claude model, and it had reached 480 stars, 141 forks and 6 watchers by 2026-09-30. He writes that he has run his own business every weekday since 2026-08-27 on routines of this kind.
How it is put together
The parts · 6The product is instruction text, so the architecture is a file system and a set of writing rights rather than a program. A kit is one role split into eight or so routines, each a folder with one SKILL.md, each owning named files and appending to named ledgers, with a single-writer rule over everything rewritten and a named-appender list over everything appended, so that no two routines can fight over one file and the run record stays one line per run. Above the routines sit four operator controls, a pause file, the corrections section every file ends with, a changelog of the kit's own amendments and the morning brief, and below them sits one capability file that maps abstract capabilities onto concrete routes per harness, which is what keeps the routine bodies free of tool names and lets thirteen different agents run the same text. Two design consequences shape the rest: because a scheduled agent runs unattended, every routine decides for itself whether it should work at all from the clock and its own row rather than from the scheduler, and because the member is often the only actor who can finish a job, every deliverable is built to be one click from done and to leave a durable file behind it so a closed tab loses nothing.
- employees/
- The eight kits, one folder each: GTM Engineer, SEO/AEO, Web Dev, Social Media, Ad Manager, Sales, Customer Satisfaction and Chief of Staff. 36 to 43 files and 1.0 to 1.1 MB of markdown each, around 8.5 MB in total, led by CONTRACT.md at 93 KB to 121 KB, the spine that carries the routine roster, the file map, the two guardrails and the run record schema.
- routines/
- Sixty folders across the eight kits, one SKILL.md each, from a 7 KB answer-visibility routine to a 104 KB inbox intake. The folder name, the YAML name key and the routine id in SCHEDULE.md are always the same string, and every one carries metadata marking it internal so a skills registry never offers a scheduled routine as an on-demand skill.
- scripts/
- The small deterministic helpers, all dependency-free and all with a self test: runlog.mjs at 33 KB to 37 KB is the only sanctioned way to append a run record, guard.mjs runs the pause, window and once-per-period checks before a single document is read, copy-check.mjs polices the copy rules, and the Ad Manager adds a 24 KB review page. selftests.mjs in CI runs every one of them.
- recipes/
- The browser craft. BROWSER-RECIPES.md ships with each kit, 45 KB to 51 KB of named technique that routine bodies reference instead of restating, and the per-site flow files never ship: recipes/flow-name.json is learned on the member's own screens by the routine that owns it and is classified as member data, so an upgrade never touches it.
- installer/
- Two files, 22 KB together. cli.mjs implements hire, list, upgrade and contribute with no dependencies on Node 18 or newer and refuses a destination folder under cloud sync; upgrade.mjs classifies every file in a kit as kit, member or merge from employee.json, compares hashes against the receipt written at hire time, and writes a file the member edited beside the original rather than over it.
- docs/, .github/scripts/ and .claude-plugin/
- Ten reader documents led by the 20 KB Agent Employee Standard, the harness matrix, the cost method and the upgrade notes; two CI gates, one running every scripts self test and one failing the build on a single em dash or en dash by code point; and the plugin and marketplace manifests that make the repository root installable as a Claude Code plugin at the same version as the package.
Choices, and what they beat
Outbound actions held by default and released channel by channel in RELEASES.md over a hard rule that an Employee never sends anything
Written up in the 1.4.0 release notes as a correction: the two stops had been written as law, but in practice one of them was a default people wanted to move, channel by channel, once an employee had earned it, and the only honest gate was always the harness's own permission layer. The release file ships empty, so every channel is held exactly as before, and only the member writes in it.
Self-improvement with no approval gate inside the kit over pending, approved and applied folders, proposal files and tick-to-approve
Stated with the reversal in full: that scaffolding was built once and removed. The harness already decides whether an agent may write a file, that control is enforced by software rather than prose and is in the right place, while a second gate invented in a markdown file adds no safety and puts friction in front of the one loop that compounds. What replaced it is a changelog line carrying the replaced text, and a brief that reports.
The save test, judged on what a control commits over the older rule that banned advancing on any control labelled Save
Named as over-blocking, and for a concrete reason: the old rule would have prevented saving a mail draft, which is exactly the deliverable this kit wants, and would have thrown away every long form it filled. The test now proceeds on draft, saved, unpublished or unlisted, stops on published, live, submitted, sent, active, ordered or visible to anyone else, and stops on every save inside an account that can spend.
Routines name capabilities and one file maps them to routes over naming the tool, the extension or the selector inside the routine body
Called a defect rather than a feature when a tool name appears in a routine or a recipe: the split is what makes a kit portable, lets a hosted route slot in later without a routine changing by one word, and keeps a harness-specific fact in one row of one file instead of scattered through sixty instruction files.
Per-site flow files learned on the member's own machine over shipping the flows for known sites
No flow file ships with a kit and none is the member's to supply, so a first run on a real account is the normal case and the absence of the file is a job rather than a blocker. A learned file carries only steps the routine read back from the live page, with an honest last_failed step number when it could not, on the argument that an invented selector is worse than a failing step because a failing step is visible.
No evasion techniques anywhere in the browser layer over looking less like automation with proxies, spoofing, jitter or backoff
Everything runs inside the member's own logged-in browser, as the member, on the member's own machine, reading pages the member can already see, so there is nothing to evade and building evasion into a member kit would put the member's accounts at risk for no gain. The recipes also refuse jittered delays, on the grounds that randomising them makes a failure impossible to reproduce without making anything safer.
Read fromdocs/STANDARD.md (the Agent Employee Standard 1.4, read in full), docs/HOW-EMPLOYEES-WORK.md, docs/HARNESSES.md, docs/GUARDRAILS.md, docs/WHAT-SETS-THEM-APART.md, CHANGELOG.md (every release section plus the unreleased one), README.md, AGENTS.md at the repository root, employees/gtm-engineer/SCHEDULE.md, employees/gtm-engineer/CAPABILITIES.md, employees/gtm-engineer/recipes/BROWSER-RECIPES.md, employees/gtm-engineer/routines/gtm-signal-sweep/SKILL.md, employees/gtm-engineer/employee.json, employees/gtm-engineer/RELEASES.md, employees/gtm-engineer/scripts/runlog.mjs, installer/cli.mjs, installer/upgrade.mjs, package.json and .claude-plugin/plugin.json, with the file counts and byte sizes taken from the recon report tree.
Build log
6 stages- 01
Eight roles, sixty routines, and no release to point at
The tree holds 359 files and roughly 8.5 MB of instructions: eight kits of 1.0 to 1.1 MB each, 36 to 43 files apiece, in which CONTRACT.md runs from 93 KB to 121 KB and a single routine's SKILL.md from 7 KB to 104 KB. None of it is a runtime. The only executable code is a dependency-free Node script or three per kit, a 12 KB installer and a 10 KB upgrade script, and roughly 60 to 70 percent of a kit by bytes is shared standard text with the role name substituted, which is why guard.mjs is byte-identical at 20,074 bytes in all eight of them. A member installs with
npx ai-employees hire gtm-engineer, which refuses a folder under OneDrive, Dropbox, Google Drive or iCloud, copies the kit, runs its self tests and prints the one prompt to paste into an agent session; on Claude Code the same eight kits also arrive as a plugin carrying one skill,hire. The repository was created on 2026-09-02 and carried 52 commits by 2026-09-30, every one of them by the author and all of them inside that single month, with 50 of them carrying a co-author trailer that names a Claude model. There are no tags and no GitHub releases. The version lives in package.json at 1.8.0, in each kit's VERSION file and in CHANGELOG.md, and the release gate is written down as the maintainer runningnpm publish --access public, becausenpxserves the kits from npm rather than from the repository. - 02
One table holds every time, and a repeat is a defect
SCHEDULE.md is the only file in a kit allowed to carry a cadence, a fire time, a window, a period key, a budget or a browser lane, and AGENTS.md calls a routine that repeats one a defect rather than a second source, because a time that lives in two places will eventually disagree with itself. Every routine opens on the same five items: the pause switch, the window guard, the once-per-period guard written before the work rather than after it, the wall-clock budget and the browser lane, and scripts/guard.mjs runs the first three before a single document is read. The GTM Engineer's rows begin at 06:45 with a 25 minute signal sweep in the heavy browser lane, then a 12 minute standup that never takes the browser, then drafting at 08:15 and a launch step runner at 09:15 on 30 and 35 minute budgets. The staggering is arithmetic rather than taste: a browser-capable routine takes the first free minute at or after the previous browser-capable fire plus that routine's full budget plus twenty minutes, which leaves the tightest gap in the week at 30 minutes. Sunday is deliberately absent from the day vocabulary, because it belongs to the ISO week that just ended and a weekly routine there would silently lose one of its two runs. The once-per-period guard turns a late or duplicated fire into nothing rather than a second copy of the same day.
- 03
The browser is inherited, and clicks are never coordinates
A routine never names a tool, only a capability: page.read, element.click, field.set, browser.tab.open. CAPABILITIES.md is the only file that maps a capability to a route on a concrete harness, and its confidence column reads
confirmedfor exactly one harness, Claude Code, the one this was built and run on, whose browser route is a Chrome extension bridge into the Chrome the member is already signed in to. Everything else isexpectedorunknown, since a row that saysunknownbeats one that says yes and is wrong at 06:45 on a Tuesday. Nothing in a kit ever signs in: it inherits a session, so a harness that launches a clean automated browser recordsblocked-loginevery morning, and the install tells the member to ask one thing first, whether the harness attaches to the profile they are signed in to or starts a fresh one. Clicks go by element reference taken from a structural read and never by screenshot coordinate, because a coordinate click silently does nothing when the page renders at a pixel ratio that does not match the frame; typing has a five-rung ladder; and a read taken straight after a navigation can return the previous view with no error, so verdicts come off a capture. Two routines are never the same signed-in identity at once: the lock file names platforms rather than the whole browser, is stale after 45 minutes, and is deleted on every exit path. - 04
Three loops make the next run better, and none of them asks
The Standard describes three self-improvement loops and says none of them stops for approval. An in-run repair is fixed in the run that hit it and written nowhere else. Site drift, a moved selector or a changed confirmation string, is written to recipes/flow-name.json, one site and one flow per file, carrying an owner, a version, a last_verified date and a last_failed step number. A standing instruction that turned out to be wrong goes into the routine's own SKILL.md, edited surgically rather than rewritten whole, and the amendment appends one line to improvements/CHANGELOG.md carrying the full text it replaced, which is the undo. No flow file ships with a kit: a routine that needs one learns it by driving the flow once and writing down only what it read back from the live page, so an absent file is a job rather than a blocker, and a page-level discovery is edited into BROWSER-RECIPES.md the same day. This was built the wrong way first: an earlier version had pending, approved and applied folders, proposal files and a tick-to-approve step, all of it removed, because the harness already decides whether an agent may write a file, that gate is software rather than prose, and a second prose gate only puts friction in front of the one loop that compounds. One property survived: a self-edit may make allowed work better and can never widen what is allowed.
- 05
What is held, and why the button label is not the question
Two guardrails carry the safety story, and only one is the member's to move. Outbound actions, sending, posting, submitting, publishing and spending, ship held on every channel, and RELEASES.md, shipped empty and classified as the member's file so no upgrade touches it, is where a channel is handed over one row at a time; no routine ever writes a row in it. The guardrail on credentials has no release at all: an Employee never creates an account, enters or generates a password, completes a captcha, enters payment details, accepts terms, or writes a credential into a file. The save test decides everything inside that, and it replaced an older rule that banned advancing on any control labelled Save, which over-blocked, since it would have prevented saving a mail draft, the deliverable. What matters is what the control commits, not what it says, so seven labels are barred by name whatever the page claims, Submit, Publish, Post, Send, Activate, Enable and Create account. The same instinct runs through the browser work: LinkedIn is read-only without exception, because that platform flags automated activity and the account is the asset; an email address is never constructed from a pattern; and one campaign per person holds forever. Every run ends in one line appended through scripts/runlog.mjs, which refuses a record carrying a secret, a draft, a URL or a person's name.
- 06
Seven versions in twenty days, and an unpublished matrix
Seven version entries in twenty days, from 1.2.0 on 2026-09-03 to 1.8.0 on 2026-09-23, with the Agent Employee Standard moving to 1.4. Three weekly compatibility matrix pull requests, the last dated 2026-09-26, are open and each supersedes the last, and each ends by saying that the maintainer decides whether the matrix goes into the public repository. What is inside them is the honest part: one row tested across all four columns, the GTM Engineer on Claude Code on Windows, seven more kits tested for install only, and every other harness, macOS and Linux recorded as documented rather than tested, with the live install's log at 104 records and 0 failed in seven days. The community has produced a localized pull request opened upstream in error and closed by its author, a routine request for discovering machine-readable paid work, and an unanswered reliability report from 2026-09-19 claiming a confirmed data-loss path where a second upgrade can overwrite a local customization. The author also writes down the failures behind his rules: the GTM Engineer reported his own already-cleared launch gates as blockers for two days while his launch was visibly live, which became the law that a routine observes a gate for three minutes before reporting the member as the blocker, and the standup once quarantined an edit made by an interactive session, which the contract calls a third actor.
Adjacent records
All records →No. 108
agy-staff
A plugin that hires Google’s Antigravity CLI as a member of staff: the host agent — Claude Code, Codex or Pi — keeps the decisions and hands the surveys, reviews and scoped edits to a Gemini 3.8 Flash worker that runs in the background and answers through a job id.
No. 073
Engram
A learning engine that installs into a coding agent: a curriculum architect breaks a topic into a first-principles concept map, a tutor makes you predict, attempt and explain before it explains, a blind assessor grades your verbatim free recall and writes a receipt for every verdict, and a deterministic FSRS-4.5 core in one Python file decides when each concept comes back — with explorable HTML built only for the concepts whose content rewards manipulation.
No. 117
sepia
A portable de-AI writing skill: four operations over one canonical rules file, narrative architecture repaired before word choice on fiction, a thin rule file matched to the venue on professional prose, and every rule labelled as measured, consulted or the project’s own inference.