ORCH
A runtime for running several coding agents on one project at once: you define a team, give it a goal, and a CTO agent decomposes the work while the rest pick up tasks — across Claude, Codex, Cursor, Grok and a plain shell — with all the state kept in files rather than a database.

What it is
An agent runtime shipped as a command-line tool. You deploy a team — a CTO, backend agents, a QA agent, a reviewer, each bound to whichever CLI you want it to run — give the team a goal, and an orchestrator loop reconciles processes, dispatches idle agents and collects finished runs on a tick. It drives Claude, Codex, Cursor, Grok, Antigravity, Pi, OpenCode or any shell command through one adapter interface, exposes a full-screen terminal dashboard and a headless mode for CI, and keeps every task, agent, run and message as files under a single hidden directory. The engine is also published as a library, so the same runtime can be embedded in another Node application.
Who built itA single account opened in September 2024, with 50 public repositories, 11 followers, and a bio that says “AI agents creator”. He wrote 581 of the repository’s 585 commits, under an iCloud address, with the remainder from two other contributors and one commit attributed to Claude.
How it is put together
The parts · 6A hand-wired engine with a thin shell around it. The core is a small layered application — domain models and a state machine, application services and the orchestrator, infrastructure for storage, processes, templates and adapters — and the command line and the terminal dashboard are clients of it rather than the thing itself; the package re-exports the engine so another program can use the same runtime without a terminal. The orchestrator is a tick loop with three phases, reconcile, dispatch and collect, which is what makes the tool survivable: a process that died is noticed, an agent sitting idle is given work, and a finished run is read back and turned into a task transition. Everything an agent CLI can be is reduced to one adapter shape, so adding a new one is a file rather than a branch in the loop. State is files, chosen so that a person can read and repair it with a text editor, with atomic writes on the way in and tail-reads on the way out. And the parts that grow without bound — event streams, dashboards, logs — all have explicit ceilings.
- src/ and its four layers
domain/holds the models, the transition table and the error types;application/holds the orchestrator, its tick loop and the event bus;infrastructure/holds the adapters, the file stores, the process manager, the template engine, workspaces and the skill loader;cli/holds the commands and the headless serve mode, andtui/holds an Ink-based dashboard of 26 files.container.tsis the two-tier dependency graph, andbin/cli.tsdecides which tier a subcommand needs.- .orchestry/
- The entire runtime state, beside the user’s project: YAML files for tasks, agents, goals and teams, a JSON record plus a JSONL event stream per run, per-task workspaces, and a single global state file. It is the whole database, and it is plain text.
- skills/library/
- Twenty-six Markdown skills, from a 32 KB design-review workflow and a 24 KB land-and-deploy runbook down to three-line guards, including one that teaches an agent how to drive Codex. They are prompt material loaded per agent at dispatch rather than code paths in the orchestrator.
- docs/
- Four specifications, each opening with a status line that says what it is: the terminal interface design, a 28 KB MCP specification, a 19 KB specification for the headless serve mode, and a 27 KB memory specification marked as research rather than roadmap.
CLAUDE.mdsits above them as the working guide to the codebase. - landing/ and assets/
- The product site, written by hand and committed: a 106 KB front page, a 38 KB comparison page, a 21 KB rationale page, a demo page, a 624 KB video, an Open Graph card, plus robots and sitemap files and a Vercel configuration. The release script edits the version shown here, which is why it is in the repository at all.
- test/ and .github/workflows/
- 125 unit test files mirroring the source tree, with mock factories and a dependency builder for orchestrator tests, plus integration, end-to-end and a single “battle” test. Three workflows: a 338-byte continuous integration file that checks out, installs, type-checks and tests; a publish workflow that re-runs the checks on a tag, skips an already-published version and extracts release notes from the changelog; and a welcome workflow for new contributors.
Choices, and what they beat
All state in files rather than a database over an embedded database or a server
Stated as a design line — everything under
.orchestry/, no database — and backed by the two utilities that make it safe: writes go through a temporary file and a rename so a crash cannot corrupt a record, and event streams are read from the tail so a long run does not have to be loaded to be displayed.A light dependency container for read-only commands over one container for everything
The two tiers are described with their contents: stores and services for commands that only read, plus the orchestrator, process manager, adapters and template engine for the three commands that run work. The entry point builds the tier the subcommand needs and imports the command lazily, so listing tasks does not pay for starting processes.
Skills are Markdown injected into the prompt, not code in the orchestrator over hard-coding each capability as a branch
The loader reads plain files, injects the content at dispatch, passes colon-namespaced skills through to the external CLI’s own system untouched, and validates every name against a pattern so a crafted file name cannot escape the skills directory. An agent’s abilities are listed in its own definition.
One adapter interface for every agent CLI over special-casing each vendor in the loop
An adapter spawns a process and returns a process id with an async stream of events, which is the only shape the orchestrator knows; eight adapters are registered and resolved by name from the agent’s own definition, including a plain shell so anything that runs from a command line can be an agent.
The release script owns the version in three files over bumping the package manifest and leaving the rest
The script updates the manifest, the version string printed by the command line, and the number on the hand-written landing page, then tags. The documentation adds a list of what it deliberately does not cover, so the manual steps are named rather than discovered later.
Specifications carry a status line over documents that read as though they were built
Each document under
docs/opens by saying what it is: a design reference, a specification, or research that is exploratory and not committed to the roadmap. The memory specification in particular sets out the problem and the current limits before proposing anything.
Read fromCLAUDE.md (6,937 characters), docs/CLI_UI_DESIGN.md, docs/AGENT_MEMORY_SPEC.md, README.md (25,731), the two workflows, and the complete 346-file tree with sizes.
Build log
6 stages- 01
One goal, five agents, and 85% of the commits that say who wrote them
The repository was created on 2026-03-10 and reached 585 commits by 2026-08-01, of which 497 landed in March alone — then the pace fell to 32, 20, 4, 27 and 5 in the months that followed. One author wrote all but four of them. The build signature is unusually clean: 498 commits carry a co-author trailer, 85% of the total, and they name Claude models almost exclusively — Opus 4.6 on 283, Opus 4.6 with a million-token context on 155, Sonnet 4.6 on 45, Opus 4.7 on 14. In five months the project cut 34 releases across 45 tags, from
v0.1.0on 2026-03-12 tov1.0.34on 2026-08-01. It has 167 stars, 17 forks and 7 open issues, and its README advertises 1,954 passing tests on the badge — the author’s own count, and the kind that a repository can keep honest or not. Nothing has been pushed since 1 August, and five pull requests are still open. - 02
A terminal interface with a written specification, in Russian
docs/CLI_UI_DESIGN.mdis 35 KB of interface design for a command-line tool, and it is written in Russian. Its opening line sets the aesthetic — instrument-grade terminal, in which every character counts and nothing is drawn for beauty’s sake, because the information should read like a pilot’s panel: glance, understand, act. It then defines three modes as three levels of immersion — a one-shot command that prints and exits, a full-screen dashboard, and a watch daemon with a live log and a compact status line — and treats each as its own design problem. What follows is the part most terminal tools never write down: the status block is specified to fit 80 columns so it works in any terminal, colour is ANSI 256 with a fallback to 16, task rows carry a fixed glyph vocabulary (●running,○waiting,◈review,✓done,✕failed,↻retrying), priorities have fixed colours from P1 red to P4 dim, rows sort in one defined order, the footer is always a single summary line, and every command has a--jsonflag so a script does not have to parse the picture. - 03
No database: all the state is files, and every path that could grow has a ceiling
The storage decision is stated in one line — everything lives under
.orchestry/, there is no database — and the consequences are spelled out. Tasks, agents, goals and teams are individual YAML files; a run is a JSON record plus an append-only JSONL event stream; global state is a single JSON file. Two utilities carry the weight:atomicWrite()writes to a temporary file and renames it, so a crash mid-write cannot leave a half-written record, andreadJsonlTail()reads the last N lines of an event stream without loading the whole file, which is the difference between a dashboard and an out-of-memory kill after a long run. The same reflex appears in the interface and the daemon: the terminal dashboard batches incoming messages on an 80-millisecond flush, caps any single detail string at 2 KB, and holds run identifiers in a map with an LRU ceiling of 500; the headless mode logs only every sixth idle tick so a quiet night does not fill a disk. This is what a program looks like when its author expects it to run unattended for hours beside processes that emit events faster than a person can read them. - 04
Layers that load only what the command needs
The architecture is layered domain, application, infrastructure and interface, with dependency injection written by hand — no framework and no decorators. The interesting part is the container, which comes in two tiers: a light one holding only stores and services, used by every read-only command such as listing tasks, reading logs or showing status, and a full one that adds the orchestrator, the process manager, the template engine and the adapters, used by the three commands that actually run work. The command-line entry point decides which tier to build and then lazily imports only the module for that subcommand. For a tool whose job is to start other processes, startup latency is the first thing a user feels, and this is a design that treats it as a feature rather than an accident. Underneath, an adapter interface turns each external agent CLI into the same shape — a spawned process, a process id, and an async stream of events — with registered adapters for Claude, Codex, Cursor, Grok, Antigravity, Pi, OpenCode and a plain shell. Around the tasks sits a small state machine that validates every transition with one function, with a promise-chain mutex serialising state changes and a tick loop that reconciles process liveness, dispatches to idle agents and collects results.
- 05
Skills as plain Markdown, injected at dispatch
Twenty-six skills ship in the repository as ordinary Markdown files, and the loader treats them as two different species. Ones named with letters and hyphens are read at dispatch time and their content is injected into the agent’s prompt; ones named with a colon are passed through untouched, because they belong to the external CLI’s own extension system. The loader caches them for the life of the process, reads them in parallel, and validates every name against a pattern — a file name that could climb out of the skills directory is rejected before it is opened. Each agent names the skills it wants in its own definition, so what an agent knows how to do is data in a file rather than behaviour in the orchestrator. The same instinct produced the memory specification: a 22 KB Russian technical document, marked Research — exploratory design, not committed to the roadmap, which states the problem before proposing anything — each invocation of the underlying CLI is a fresh subprocess with no history, and the only memory an agent has today is what the orchestrator assembles by hand into the prompt, a key-value store with no structure or search, one retry’s worth of previous failure, and a final summary capped at 2 KB.
- 06
A release script that keeps three files in step, and a checklist that names the ones it does not
The version number lives in three places, and the release script owns all of them: it bumps
package.json, rewrites the version string in the command-line entry point, and updates the copy on the landing page, which is a hand-written static site in the repository rather than a generated one. The publishing workflow runs the type checker and the tests again on the tag, checks whether that version is already on npm and skips the publish if it is, and pulls the matching section out of the changelog to use as the release body. The documentation is honest about the gap: a section headed “manual updates not covered by the release script” lists the changelog, the test-count badge in the readme, and the statistics and feature copy on the landing page. For a repository that has shipped thirty-four versions, that list is the difference between a release process and a hope. The pull-request history is the quieter half of the same story: eleven requests in total, five merged — including a Pi RPC adapter from one contributor and a light theme palette from another — five still open, and three from a single contributor sitting since April.
Adjacent records
All records →No. 039
Vibe Projects
One person’s monorepo of six AI-generated projects — an Android agent harness, a per-hunk diff reviewer for VS Code, and four small Android apps — kept deliberately as a running measure of what the models could build.
No. 027
CCManager
A terminal menu that keeps several coding agents running at once, one per git worktree: it shows which are busy, which are waiting on you and which are idle, creates and merges the worktrees, and can bring the whole set back after a crash.
No. 061
DeepSeek Harness
DeepSeek’s agent harness, built so that the model adapter, the tool registry, the session log and the agent loop itself are plugins — swapped from a configuration file rather than a fork.