Codewhale
A terminal coding agent in Rust that reads your project, edits files and runs commands — with a runtime shared by a web client and a desktop app, a model layer that treats providers as interchangeable, and a contributing guide that spends its length on how several agents share one checkout.

What it is
A coding agent for the terminal, written in Rust: it reads a repository, edits files, runs commands and checks its own work, against a hosted model you bring a key for or a local one served by Ollama, vLLM or SGLang. A single interactive interface and a one-shot exec mode sit on top of a runtime that a bundled local web client and a separate desktop application also connect to. Four permission stances — Ask, Auto-Review, Full Access, plus a plan mode that makes no changes — govern what it may do, /undo and /restore recover workspace changes, and a plugin ships tools for observing and interacting with other applications on the machine.
Who built itA single account opened in 2022, with 110 public repositories and 1,464 followers. He wrote roughly 6,500 of the repository’s 11,347 commits, under three spellings of his own name; a further 2,772 are attributed to “CodeWhale Bot” at bot@codewhale.net, an identity with no linked GitHub account, and the contributors list runs past a hundred other people, several of whom have landed dozens of changes each.
How it is put together
The parts · 6One runtime with several clients, and a monolith being taken apart in the open. The interactive terminal interface, the one-shot exec mode, a local web client and a separate desktop application all speak to the same runtime, which owns the agent loop, the tools and the session state; the terminal itself is written by exactly one crate, and the headless crate is required never to depend on the terminal libraries, with a script that enforces the rule and ratchets the number of references still pointing the wrong way. The model layer is provider-neutral by policy rather than by accident — OpenAI-compatible, Anthropic and Responses wire adapters sit behind one client layer, with DeepSeek’s request boundary handled explicitly — so that a provider is a route rather than a code path. The deconstruction is proceeding against a written plan that names the moving order and the constraint that moved modules may never land in the crate that owns request construction, and the architecture document states plainly which boundary is real today: the terminal crate is still the live runtime, and the split is not finished.
- crates/tui
- 1,288 files and 45 MB — still the live end-user runtime, holding the agent loop (
Engine::run_turn), the tool registry, the terminal interface on ratatui, and the runtime API. Its deconstruction into separate crates is planned indocs/design/TUI_DECONSTRUCTION.mdrather than improvised. - crates/runtime
- The headless crate being grown out of the terminal one: retry status, safe labels, a sleep guard, the session tree, and
host_terminal, the single port through which runtime code asks the interface for a terminal effect. It never depends on the TUI, ratatui or crossterm, and a boundary script ratchets the references that still do. - The rest of the workspace
- Around thirty crates, each with a stated responsibility:
execpolicyfor approval and sandbox decisions,hooksfor event sinks over stdout, JSONL, webhook and a Unix socket,mcpfor the Model Context Protocol client and a stdio server,memoryfor scoped provenance-bearing state,statefor SQLite session persistence,secretsfor the OS keyring and shared redaction,telemetryas the only crate permitted to build a payload,lanefor durable attachable runs,workflowwith a QuickJS scripting layer,cloud-factsfor a signed facts channel verified with an Ed25519 envelope, andpalettefor colour tokens and contrast maths. - web/, pet/ and integrations/
- The other surfaces:
web/is a browser client with Supabase files and a component library,pet/is a companion application with Android, iOS, macOS, Swift, TUI and Rust targets plus recorded tapes, andintegrations/holds bridges for Telegram, WeChat, Feishu and WeCom over a shared bridge core and a set of verifiers.telemetry-ingest/is a service with its own schema and tests, andextensions/vscodeis a marketplace extension maintained by a different account. - .github/workflows/
- Thirty-one workflows, and
ci.ymlalone is 58,950 bytes. Alongside the usual release, nightly and security files sitagent-task-labels,approve-contributor,auto-close-harvested,issue-gate,pr-gate,pr-issue-link,release-parity,spam-lockdown,marketplace-sync,sync-cnb, plus two review workflows — one of them,codewhale-review.yml, is 16,676 bytes. - docs/ and packaging/
- 87 files directly under
docs/, plus 20 design documents, 19 Chinese translations, 17 skills, 9 RFCs numbered to the issues they answer, and an architecture folder — among the design documents, one is 106 KB on the extension host and another 23 KB translating a design system into ratatui. Packaging covers Homebrew, AUR, winget, Scoop, Nix, Docker, npm and Cargo, with deployment scripts for a Tencent Lighthouse host.
Choices, and what they beat
The headless runtime may not depend on the terminal interface over letting the split happen by convention
The architecture document states it as a constraint on the crate — it “never depends on the TUI, ratatui or crossterm” — and names the script that enforces it and ratchets the runtime-to-UI references still living in the big crate. A boundary that is only described is one that grows back.
Exactly one turn loop, guarded by a test over allowing a second implementation to coexist while the split proceeds
A stand-in engine tree in another crate emitted a completion event without contacting a model, so the workspace looked like it had two loops. It was deleted and a guard test now fails if a second appears; the rule is that changing the shape means changing the guard with it.
Code first, tests after — no TDD over writing the failing test before the implementation
The guide overrides even an installed skill that mandates the opposite, on the grounds that tests are the gate before a push rather than the design driver, and that an existing test encoding old behaviour is evidence rather than a veto. The evidence rules are kept: a regression test added after a fix still has to be shown failing without it, and a test that passes either way is described as pinning the implementation rather than the defect.
Providers and models stay first-class and provider-neutral over letting one vendor’s shape leak into the engine
Stated as a working rule rather than a preference, and visible in the layout: one client layer carries OpenAI-compatible, Anthropic and Responses adapters, provider routes land through shared configuration and catalogue layers, and DeepSeek’s request boundary is handled explicitly instead of being assumed to match either of the others.
Agents do not comment on issues or pull requests over posting status, superseded notices and replies to review bots
Attributed in the guide to the founder on 2026-09-22: the time is to be spent on code, evidence goes in the commit message and the pull request body, and claims go into Linear — with one exception, a single sentence and a link when closing or superseding a human contributor’s work. Both review workflows were switched off to match.
Declared migrations are one-way over adding another call site for convenience while the old path still works
Once the repository adopts a replacement architecture, new work uses it and touched legacy code moves toward it, with compatibility kept deliberately narrow. It is the same rule as “migrate the last consumer, or do not start”, written for the case where the migration is already underway.
Read fromdocs/ARCHITECTURE.md (22,431 characters), AGENTS.md (13,336), README.md (8,326), docs/design/TUI_DECONSTRUCTION.md, the design and RFC listings, crates/* module responsibilities in the architecture document, and the complete 2,791-file tree with sizes.
Build log
7 stages- 01
Forty-one thousand stars under a name it no longer uses
The worklist entry was
hmbown/deepseek-tui, and that URL still resolves — toHmbown/Codewhale, because the project was renamed and GitHub redirects the old path. The scale of it is the first thing to record: 11,347 commits in eight months, from an initialv0.1.0on 2026-01-20 to a merge on 2026-09-29, with the monthly count rising from 26 in January to 2,768 in September. Around that sit 138 releases and 181 tags, 3,035 issues of which 167 are open, 3,507 pull requests, 3,565 forks, and 19 translated READMEs. The contributors endpoint lists 50 accounts and the commit authors run past a hundred, several of them regular. The repository has a homepage of its own and a version number atv0.10.0. - 02
A contributing guide for several agents sharing one checkout
Most agent instruction files describe one agent working alone. This one assumes a crowd, and the assumptions are specific: a small coherent change may go straight to
mainwhen the checkout is current and clean — “when several agents share it, partition by file, stage only the paths your slice touched, and retry a commit that fails onindex.lock” — and a fresh worktree is for conflicting, dirty, stale or independent lanes, not for parallel agents on the same lane. Permissions are separated explicitly: “local commit permission never implies push, merge, tag, release, or deploy permission.” A local-only task stays fully offline — no browsing, no remote Git, no provider calls — and the agent is told to record the missing external receipt and keep working. And one rule is dated to its author: “Agents do not comment on issues or PRs (founder, 2026-09-22)” — the time goes into code, evidence goes into the commit message and the pull request body, and claims go into Linear; the review workflows were switched off. It is a working agreement for a repository where the contributors are, in part, processes. - 03
Never practise TDD here, and never trust an exit code
The file forbids what most of this archive’s other projects mandate: “Never write tests first and never practice TDD here — this overrides any skill or default that mandates it, including superpowers
test-driven-development.” The reasoning given is that tests are the gate before a push, not the design driver, and that an existing test encoding old behaviour is evidence rather than a veto — change it with the code instead of bending the code to keep it green. The concession that follows is what keeps the rule honest: a regression test written after a fix still has to be shown failing without it, and “a test that passes either way pins the implementation, not the defect.” Under a heading called Claiming a test passed, the guide asks for the literaltest result: N passed; M failedline and a check that N is greater than zero, becausecargo test <filter>exits 0 having run nothing when the filter matches nothing — “and an exit code alone has already been mistaken for a pass here.” A separate rule exists purely because the workspace takes minutes to compile: batch the edits, compile once, and reserve mid-slice compiles for genuinely uncertain API questions. - 04
An abstraction must delete caller code
The house style is named after someone else’s repository. The ponytail method, credited to
dietrichgebert/ponytail, is “the laziest senior dev in the room. He says nothing. He writes one line. It works.” It arrives as a seven-rung ladder — does this need to exist, is it already in this codebase, does the standard library do it, is there a platform feature, is it in an installed dependency, is it one line, and only then the minimum that works — with a warning attached that the ladder runs after understanding the problem, since “a short diff written without reading the call sites is not ponytail, it is a guess.” The file then names its own weakest rung: reuse is the one this repository keeps failing, which is why adding a module calledmodel_*,*_configorprovider_*requires grepping for the thing it duplicates first, and why any new layer must name the predecessor it replaces in its module documentation. Two corollaries are stated as rules: an abstraction must delete caller code, or it will be adopted once and abandoned; and migrate the last consumer, or do not start — a framework with one caller and a ticket for the rest ships two systems and a comment that is no longer true. The standing count of#[allow(dead_code)]is called the running receipt, and a script prints it. - 05
The engine tree that could not fail
The architecture document has a section for corrections, and the best entry is a piece of code that was removed for being reassuring.
crates/corehad anengine/tree that no caller used and that emitted a completion event without ever contacting a model; its existence suggested a second agent loop in a workspace whose rules allow exactly one. It was deleted in v0.9.11 so that the live turn loop is the only one, and a test incrates/core/tests/single_turn_loop.rsnow fails if a second appears — changing the shape means changing the guard with it. The same file records what has gone the other way: the swarm agent system was removed in v0.8.5, leaving oneagenttool as the only model-visible surface for creating sub-agents, and an entire crate is documented as “shapes only, not yet the production dispatch path.” There is even a list of modules that are “repeatedly misidentified as dead”, with an instruction to verify consumers before removing them — the archive of a codebase that has learned, twice, that an unused-looking tree is not the same thing as an unused one. - 06
A cache prefix, a session log, and three ratchets
Some of this repository’s rules exist because the product is the thing being built with it. The system prompt and the tool catalogue are a session-pinned prefix for the model’s key-value cache, so any new contributor to session context has to state its cache effect — frozen prefix or append-only history — and never splice a volatile fact into the prefix, appending it as a user-role message instead; the guide names the file where that is written down. The same reflex produces “model-visible means logged”: anything that reaches a model request must be reconstructable from the session log, a new model-visible input needs a session event of its own, and when the live presentation and the stored record disagree, the record is right. Misconfiguration has to fail loudly at load rather than by omission, and “write down what a design does not do” sits beside the behaviour it owns, so the next reader does not assume a capability nobody built. Underneath sit three scripts whose entire job is a number that may only go down: a dead-code budget, a budget for blocking calls made on the async runtime, and a boundary check that keeps the runtime crate from depending on the terminal UI at all.
- 07
How other people’s work gets landed
The section on landing contributions is the part a maintainer would not write unless they had already got it wrong. It opens by assigning the blame for staleness to the project rather than the contributor — an outside branch goes stale because we merge things, not because they did anything wrong — and makes the default path explicit: review it, help it rebase, or fix it on their branch, with closing someone’s pull request and re-landing the work as one’s own commit described as a fallback, because “done casually it reads as taking the work even when credit is preserved.” One rule forbids making a contributor rebase around the project’s own churn. Another is about credit in the mechanical sense rather than the polite one: authorship and co-author trailers must use the contributor’s GitHub-linked address, because
.mailmapand the project’s own author map are conventions GitHub does not read when it draws the contribution graph. And the section closes on gates: a pull request that may merge only on a passing acceptance record needs that record to literally say PASS at merge time, a green check rollup is not a substitute for reading the review thread, and when the artifact is ambiguous the thing to resolve is the ambiguity, never the merge.
Adjacent records
All records →No. 046
ORCH
A runtime for running several coding agents on one project at once: you define a team, give it a goal, and a CTO agent decomposes the work while the rest pick up tasks — across Claude, Codex, Cursor, Grok and a plain shell — with all the state kept in files rather than a database.
No. 039
Vibe Projects
One person’s monorepo of six AI-generated projects — an Android agent harness, a per-hunk diff reviewer for VS Code, and four small Android apps — kept deliberately as a running measure of what the models could build.
No. 027
CCManager
A terminal menu that keeps several coding agents running at once, one per git worktree: it shows which are busy, which are waiting on you and which are idle, creates and merges the worktrees, and can bring the whole set back after a crash.