TrueForge
A self-hosted runtime that takes over the agent loop — model calls, MCP tool servers, git-backed skills, an optional sandbox, human approvals, context compaction and session state — and hands it back three ways: a chat UI, an HTTP API with a TypeScript SDK, and an embeddable UI SDK, on SQLite in a single process or on Postgres and Redis when it is hosted.

What it is
The open-source agent harness from TrueFoundry: the runtime layer that owns the agent loop — model calls, MCP tool servers, git-backed skills, an optional sandbox, human approvals, context compaction and session state — and exposes it as a chat UI, an HTTP API with a TypeScript SDK, and an embeddable UI SDK. The repository was created on 2026-07-23 with a first commit that extracts the harness into a pnpm workspace, and in the ten weeks since it has taken 778 commits from 38 accounts to produce 2,938 files, six packages published to npm or PyPI, a Helm chart and 6,033 stars. One codebase runs two ways: a single process on SQLite with no login for local use, or Postgres and Redis behind Docker Compose, Helm or Railway. Its vocabulary is agents, sessions, turns, events and deltas, and both the TypeScript and the Python SDK are generated from one OpenAPI specification.
Who built itA company repository rather than one person’s project: the 778 commits come from 38 linked accounts, 776 of them tied to a GitHub account. The largest single author is Chirag Jain (chiragjn) with 119, followed by debajyoti-truefoundry at 77, bhaveshpatel640 at 75, govindavashishtha at 67, sr07asthana at 59 and thesujai at 54; three bot accounts add 62 more and twelve accounts have exactly one commit each. Of the 386 co-author trailers, 155 say Cursor and 19 name Claude Opus 4.8. The README’s only contact route is the two founders’ addresses at truefoundry.com.
How it is put together
The parts · 6One pnpm workspace, four layers, and a contract file that all of them have to agree on. packages/trueforge-core is the harness itself — the agent loop, capabilities, MCP clients, sandbox providers, the event schema, the session and turn handles — carrying no HTTP and no interface. packages/trueforge is the server: handlers, routes, two parallel database trees, catalogs loaded from YAML, authentication and one 54 KB configuration module. packages/trueforge-ui is the interface as a library — atoms, containers, layouts, themes and overridable slots — and packages/frontend is the thin shell that ships inside the server package. packages/assistant-ui-runtime is the adapter joining a React chat runtime to the server’s event stream, and it owns the port types everybody else is forbidden to redeclare. The generated artifacts hang off the same contract: an OpenAPI document of 333,161 bytes, regenerated in CI and kept in two byte-identical copies, produces a TypeScript SDK of 1,010 files and a Python SDK of 403, and no human is allowed to edit either. Two consequences follow the layering around. Because the store is an interface with three implementations, local and hosted deployments differ only in which one gets constructed; and because the sandbox, the model providers, the skill sources and the web-search providers are all catalogs loaded from shipped YAML, the shipped defaults are configuration rather than code.
- packages/trueforge-core/ (188 files, 1,165 KB)
- The harness without a transport:
AgentThread.ts(53,368 bytes) andAgentThreadOrchestrator.ts(19,894) run the loop;core/capabilities/builtins/holds the subagents, compaction, large tool responses, OpenUI, web search and ask-user pieces;core/llm/is a 53,436-byte adapter, with five test files of its own;core/mcp/the remote and local tool servers;core/sandbox/the sandbox, its providers and two Python helper scripts of 11,943 and 29,098 bytes;agent-session/the session and turn handles, the store interface and the 128,154-byte contract suite. - packages/trueforge/ (459 files, 2,315 KB)
- The server.
src/apis/holds the handlers —turns.tsat 35,765 bytes,sessions.tsat 22,989,mcpServers.tsat 19,634,schedules.tsat 19,299 — oversrc/routes/;src/db/mirrors every store for Postgres and SQLite;src/sandbox/local/implements a sandbox provider on the host, including a 27,969-byte provider and a 21,197-byte runner;src/truefoundry/bridges the same interfaces to the company’s own platform through a 26,581-byte client; five YAML catalogs supply the shipped presets. - packages/trueforge-ui/ (559 files, 2,776 KB) and packages/frontend/ (24 files)
- The embeddable interface: 58 files directly under
src/atoms/, four layout modes (dock, drawer, sidebar, widget), a slot and theme system with a 26,577-byte theming document, and a changelog of 68,869 bytes beside a 25,328-byte README and its own contributor and security files — the shape of a package that used to live somewhere else. The frontend shell next to it is a small Vite and React application with login and logout screens, and it ships inside the server package. - packages/assistant-ui-runtime/, packages/trueforge-sdk/ and python/trueforge_sdk/
- The adapter and the two generated clients: 75 files of React runtime glue whose
server/types.tsis 39,681 bytes and is the single definition of the server ports, plus a 42,394-byte hook file and a 100,655-byte test; then a TypeScript SDK of 1,010 files and a Python SDK of 403, emitted from.github/fern/openapi/openapi.json(333,161 bytes) with wire tests for every resource. - docs/ and benchmark/
- Fourteen files at the top of
docs/—docs.json, a 333,161-byteopenapi.jsonand twelve pages — with three API pages, 22 UI SDK pages, five capability pages, one page each for agent creation, authentication and harness setup, 54 screenshots totalling 21,702 KB and 11 more assets at 1,960 KB.benchmark/is a Python harness of 11 files whose largest script is 16,205 bytes, and it is what the README’s comparison with Claude Managed Agents and deepagents is built from. - Root files, workflows and deployment
- A root
AGENTS.mdof 5,822 bytes with nine nested counterparts, nineCLAUDE.mdfiles of 11 bytes each,.cursor/BUGBOT.mdplus three rule files and two package-level bugbot documents, eleven pending changesets, nine workflows led by a 20,583-byte release file and a 12,319-byte CI file,CONTRIBUTING.mdat 16,775 bytes andRELEASING.mdat 13,607, a two-stage Dockerfile with an npm variant, two compose files, a Helm chart whose values file is 15,704 bytes and whose helpers template is 33,412, and a.railwaydirectory with a 3,927-byte infrastructure file.
Choices, and what they beat
Sessions, turns and a streamed event log as the public contract over a prompt-in, response-out call
The concepts page defines the hierarchy as one agent to many sessions to many turns to many events to some deltas, and then draws the consequences: turns chain automatically because
previous_turn_iddefaults toauto, so the caller never resends history; only one turn runs in a session at a time; every event carries an id, a thread id and a sequence number, which is what makes resuming after a disconnect possible; and deltas exist only on the live stream, because a turn’s events are already merged when they are listed afterwards.One codebase, SQLite locally and Postgres when hosted over picking a single deployment target
The README’s mode table pairs local use with one process, SQLite and no extra infrastructure against teams and multi-replica hosting on Postgres plus Redis; the code keeps two migration trees and one store interface with contract tests run against both, so the difference is a constructor rather than a fork. The documentation is equally direct about the price: standalone mode ignores OIDC even when it is configured, the chart’s dev defaults include a well-known Postgres password and an unauthenticated Redis, and local mode is described as something to keep on localhost.
Generated SDKs and an OpenAPI document nobody hand-edits over hand-written clients for each language
Both copies of the specification are regenerated in CI and required to stay identical, the TypeScript and Python SDK trees are marked as generated in the pull request checklist, and fork pull requests are told to change source only while maintainers regenerate the SDK after merging. Regeneration carries its own changeset, so a generated change is still a versioned change — the same discipline that makes the SDK the contract surface rather than a copy of it.
A sandbox that is off by default and provisioned on demand over running agent code inside the server process
The README states it as a feature and a boundary at once: an isolated environment for code, files and shell commands, off by default, provisioned only when needed, with secrets staying in the harness. It is required for skills and Code Mode, and it is written against a provider interface with three implementations and a shared contract suite, which is the part that makes an outside request for a fourth provider a configuration question rather than a rewrite.
Approval by label rather than by gating every tool call over stopping on every write-capable tool
The default policy names two labels,
@writeand@destructive, and the page then records the failure mode itself: those labels only match tools the MCP server has labelled, many servers skip labels, and a tool that changes data will therefore run without asking unless it is named or the policy is set to@all. Persistent policies are keyed by tool name, so a rename on the MCP side silently voids one, and the open pull request on paused turns shows the other edge — a stream that stays open, and resources that stay paid for, while it waits for a decision.
Read fromREADME.md (9,423 characters, fetched in full from the GitHub API), AGENTS.md (5,822 characters), docs/api/overview.mdx (13,586), docs/authentication/overview.mdx (12,546), docs/create-agent/overview.mdx (25,194), .github/fern/generators.yml, .changeset/, the nine workflows under .github/workflows/, the two migration trees under packages/trueforge/src/db/, and the complete 2,938-file tree with sizes.
Build log
6 stages- 01
The unit of work is a session, a turn and a stream of typed events
What this project actually sells is a vocabulary. An agent is a saved definition — model, instructions, MCP servers, config — and is explicitly not a running process; a session holds one conversation’s context; a turn is one request inside it; a turn emits events onto a stream as JSON objects, and model output arrives as deltas that the client merges into the base event by id. The documented event set is small enough to enumerate:
turn.created,mcp.initialize,model.message,tool.response,tool.approval_requiredandturn.done, withthread.createdandthread.doneannouncing subagent threads; every event carries anid, athread_id(mainfor the root agent, a generated id for a subagent,nullfor run-level events) and a sequence number the docs say exists to resume after a disconnect. Turns chain automatically, becauseprevious_turn_iddefaults to the stringauto, so an application stores a session id and never resends history. One turn runs in a session at a time, and a turn can stop in three places — an approval, a clarifying question (tool.response_required) or MCP OAuth (mcp.auth_required) — each of them resumed by sending a new turn. The code behind that vocabulary is where the weight sits:SessionHandle.tsis 19,965 bytes,TurnHandle.ts18,924, the store interface 17,814, and the shared store contract suite 128,154. - 02
Ten weeks, 778 commits, and a version line that keeps being restarted
The first commit is dated 2026-07-23 and reads "Initial commit: extract agent harness into pnpm workspace." — which is the whole origin story, a harness lifted out of a larger product and given its own repository. The commit curve is 57 in July, 412 in August and 309 in September, 778 in total, with the newest commit on 2026-09-30. Around it sit 6,033 stars, 486 forks, 17 watchers, 99 open issues and a repository of 44,151 KB holding 2,938 files. Releases are less tidy. Six things are published at once —
@truefoundry/trueforge,trueforge-core,trueforge-ui,trueforge-sdk,assistant-ui-runtime, a Pythontrueforge_sdk— plus the Helm chart, and their versions do not move together: at the end of September the server was at 0.3.1, the UI at 0.4.1, the SDK at 0.2.1 and the chart at 0.3.0, each with-rclines behind it. On 2026-09-29 a pull request titled "Reset versions to 0.0.0" set every package and both copies of the OpenAPI document back to 0.0.0, and the Changesets bot then opened release pull requests moving them to 0.0.1 and 0.4.0-rc.1. The visible tags are a third line again —v0.1.1throughv0.1.10-rc.1, next to per-package tags such as@truefoundry/trueforge-ui@0.3.0-rc.11— and a release-workflow test plan still asks for a re-run on the branchrelease-v0.176.0for0.176.0-rc.1, a version the current packages never reach. - 03
Three bugs in the release pipeline, and a body of written law
The most useful pull requests here are about the pipeline rather than the agent. One found that the release workflow wrote
skip=falseand passed it between jobs, and that GitHub Actions drops a job output whose value is the stringfalse— so both image builds and the Helm chart publish were silently skipped; the flag becamebuildorskip. Two more describe a job skipped because an earlier one was skipped, since a dispatch skipsselect-modeand a push skipsversion-release, leaving downstream conditions false; image builds, Helm, npm and PyPI publish now use!cancelled()and name the jobs they depend on. The fourth is a production image that had quietly grown:pnpm fetchin the store stage filled/pnpm/storebut also left a fullnode_modules/.pnpmtree that later stages carried into the runtime image together with build-only packages such as esbuild and TypeScript, so the step becamepnpm fetch && rm -rf node_modules. Beside those fixes sits an unusual amount of written law.AGENTS.mdis 5,822 bytes of mandatory rules: no assertion escapes (as T,as unknown as T, non-null!,as never); an error thrown from a catch must carry{ cause: caught }; every type, schema and helper has exactly one canonical owner, with no forwarding shims; server-port types exist only inpackages/assistant-ui-runtime/src/server/types.tsand the UI package may re-export aliases and nothing else; environment reads go through a 54,846-byteconfig.ts; UI lengths use rem, not px; comments explain intent and must not carry tracker ids. Every nestedAGENTS.mdalso needs a siblingCLAUDE.mdcontaining only the line@AGENTS.md, so that Cursor and Claude Code load the same scoped rules — nine of each, and every one of the nineCLAUDE.mdfiles is 11 bytes. - 04
Two databases, one contract, and a login mode that is deliberately absent
Every store exists twice.
packages/trueforge/src/db/postgres/holds 38 migrations running from 2026-07-27 to 2026-09-29;packages/trueforge/src/db/sqlite/holds 36 from 2026-07-30 to the same week, and the two trees stay in step by subject — sessions, agents, MCP OAuth, skills, sandbox providers, schedules, session metrics, sandbox environments, with the one Postgres-only GIN index on session metadata explaining part of the difference. What keeps them honest is a set of contract tests written once and run against the in-memory store, then Postgres, then SQLite:storeContractSuite.tsalone is 128,154 bytes, with sibling suites for the agent, MCP server, model provider and OAuth token stores, six jest configurations in the server package, and two of those six reserved for a local sandbox. That split is the product’s main deployment decision, and the documentation states the cost. Local mode is one process on SQLite with no login at all; standalone mode ignores OIDC even when the variables are set; and the README says local mode is not a production or internet-facing setup, asks that it be kept on localhost, and disclaims responsibility for data loss or unauthorized access beyond that. The authentication document adds that the Helm chart defaults to no login, a well-known bundled Postgres password and Redis without authentication, then lists two holes instead of hiding them: session history is owner-only, so admins are not global superusers today, and agents created by anyone are visible to everyone on the instance. - 05
A sandbox that is off by default, and approvals that trust a label
Sandboxing is opt-in and lazy: the docs say it is off by default, provisioned only when an agent needs it, and that secrets stay in the harness rather than the sandbox. Skills and Code Mode both require it, and it is written against a provider interface —
DaytonaProvider.tsat 21,556 bytes, a second provider backed by TrueFoundry itself at 11,818, and a local provider at 27,969 with a Lima configuration, a loopback probe and a smoke script — which is why an outside request to support NVidia’s OpenShell could be filed as an alternative to Daytona and answered as an extension point rather than a rewrite. Skills are git-backedSKILL.mdpacks mounted into the sandbox by a 29,098-byte downloader, and Code Mode talks to the sandbox over NATS. The human checkpoints have the more interesting documentation, because it records its own hole: tool approval defaults torequire_approval_for_toolsof["@write", "@destructive"], and the page warns in a note that those labels only match tools an MCP server has labelled — many servers skip labels, so a tool that changes data can run without asking, and the remedy is to name the tool or set["@all"]. A recent pull request adds persistent approval policies keyed by server and tool, records that "Approve once" is never stored, and admits that renaming a tool on the MCP side leaves an existing policy matching nothing. An open one would change the shape of a paused turn: instead of finishing with a result, the stream would emitpaused, hold the connection and its resources open until an abort, and stay paused rather than closing if the consumer simply walks away. - 06
A company repository, a bot-heavy queue, and two outsiders
The work belongs to a company and is spread over 38 accounts. Chirag Jain (
chiragjn) has the most commits at 119 of 778, ahead ofdebajyoti-truefoundryat 77,bhaveshpatel640at 75,govindavashishthaat 67,sr07asthanaat 59 andthesujaiat 54; three of the accounts are bots and twelve more have exactly one commit each. 386 commits carry a co-author trailer — 155 reading only "Cursor", 49 from the release bot, 43 from a project bot namedtrueforge-dev-bot, 19 naming Claude Opus 4.8 — the paper trail of a history written with agents in the loop. The queue is heavily instrumented: the thirty most recent issues and pull requests contain three Dependabot groups (one proposing 51 version bumps at once, two for a singleundicisecurity patch, one closed with the note that the dependencies were no longer updatable), CodeQL and image-scan workflows, a changeset bot on every pull request, and Cursor summaries and agent links embedded in the bodies. Two contributions in that window come from outside the company and are both worth reading. One reports thatparseSandboxArtifactsinSandboxArtifactDownload.tsxuses the regular expression([^)]*), so the path/tmp/report(final).csvis truncated to/tmp/report(finalbefore the download URL is built; the reporter proposes a parser that recognises balanced parentheses, explains that a broader expression would swallow neighbouring links, and asks a maintainer to confirm the scope and assign the issue to him. The other asks for NVidia’s OpenShell to be supported as a sandbox provider. A third item is internal: a security pull request closing an INFOSEC disclosure by adding state to a browser cookie so that an MCP OAuth callback binds to the user who started it. The project also publishes its own benchmark, comparing TrueForge with Claude Managed Agents and deepagents on the same tasks, tools and model and claiming the same accuracy at lower cost, with the scripts to reproduce it committed underbenchmark/.
Adjacent records
All records →No. 065
GSD Core
Git. Ship. Done. — a meta-prompting, context-engineering and spec-driven development framework that runs the same five-step loop on every milestone: discuss, plan, execute, verify and ship. The heavy work is pushed into fresh-context subagents so the main session stays lean, and every decision is written into Markdown and JSON under a planning directory instead of living in the conversation.
No. 129
Agents Universe
An open-source agent platform that keeps one project context shared by everyone working in it. Agents read the whole knowledge base when a project is opened and write what they learn back into the same files while they work; knowledge is Markdown on disk with a database index behind it, and there is no embedding model or vector search.
No. 102
DeepSeek Harness
DeepSeek’s agent harness, built so that the model adapter, the tool registry, the session log and the agent loop itself are plugins — swapped from a configuration file rather than a fork.