Skip to content

TrueForge

A self-hosted runtime that takes over the agent loop — model calls, MCP tool servers, git-backed skills, an optional sandbox, human approvals, context compaction and session state — and hands it back three ways: a chat UI, an HTTP API with a TypeScript SDK, and an embeddable UI SDK, on SQLite in a single process or on Postgres and Redis when it is hosted.

Screenshot of TrueForge
Editor screenshot, 1 Oct 2026TrueForge ↗

What it is

The open-source agent harness from TrueFoundry: the runtime layer that owns the agent loop — model calls, MCP tool servers, git-backed skills, an optional sandbox, human approvals, context compaction and session state — and exposes it as a chat UI, an HTTP API with a TypeScript SDK, and an embeddable UI SDK. The repository was created on 2026-07-23 with a first commit that extracts the harness into a pnpm workspace, and in the ten weeks since it has taken 778 commits from 38 accounts to produce 2,938 files, six packages published to npm or PyPI, a Helm chart and 6,033 stars. One codebase runs two ways: a single process on SQLite with no login for local use, or Postgres and Redis behind Docker Compose, Helm or Railway. Its vocabulary is agents, sessions, turns, events and deltas, and both the TypeScript and the Python SDK are generated from one OpenAPI specification.

Who built itA company repository rather than one person’s project: the 778 commits come from 38 linked accounts, 776 of them tied to a GitHub account. The largest single author is Chirag Jain (chiragjn) with 119, followed by debajyoti-truefoundry at 77, bhaveshpatel640 at 75, govindavashishtha at 67, sr07asthana at 59 and thesujai at 54; three bot accounts add 62 more and twelve accounts have exactly one commit each. Of the 386 co-author trailers, 155 say Cursor and 19 name Claude Opus 4.8. The README’s only contact route is the two founders’ addresses at truefoundry.com.

How it is put together

The parts · 6

One pnpm workspace, four layers, and a contract file that all of them have to agree on. packages/trueforge-core is the harness itself — the agent loop, capabilities, MCP clients, sandbox providers, the event schema, the session and turn handles — carrying no HTTP and no interface. packages/trueforge is the server: handlers, routes, two parallel database trees, catalogs loaded from YAML, authentication and one 54 KB configuration module. packages/trueforge-ui is the interface as a library — atoms, containers, layouts, themes and overridable slots — and packages/frontend is the thin shell that ships inside the server package. packages/assistant-ui-runtime is the adapter joining a React chat runtime to the server’s event stream, and it owns the port types everybody else is forbidden to redeclare. The generated artifacts hang off the same contract: an OpenAPI document of 333,161 bytes, regenerated in CI and kept in two byte-identical copies, produces a TypeScript SDK of 1,010 files and a Python SDK of 403, and no human is allowed to edit either. Two consequences follow the layering around. Because the store is an interface with three implementations, local and hosted deployments differ only in which one gets constructed; and because the sandbox, the model providers, the skill sources and the web-search providers are all catalogs loaded from shipped YAML, the shipped defaults are configuration rather than code.

packages/trueforge-core/ (188 files, 1,165 KB)
The harness without a transport: AgentThread.ts (53,368 bytes) and AgentThreadOrchestrator.ts (19,894) run the loop; core/capabilities/builtins/ holds the subagents, compaction, large tool responses, OpenUI, web search and ask-user pieces; core/llm/ is a 53,436-byte adapter, with five test files of its own; core/mcp/ the remote and local tool servers; core/sandbox/ the sandbox, its providers and two Python helper scripts of 11,943 and 29,098 bytes; agent-session/ the session and turn handles, the store interface and the 128,154-byte contract suite.
packages/trueforge/ (459 files, 2,315 KB)
The server. src/apis/ holds the handlers — turns.ts at 35,765 bytes, sessions.ts at 22,989, mcpServers.ts at 19,634, schedules.ts at 19,299 — over src/routes/; src/db/ mirrors every store for Postgres and SQLite; src/sandbox/local/ implements a sandbox provider on the host, including a 27,969-byte provider and a 21,197-byte runner; src/truefoundry/ bridges the same interfaces to the company’s own platform through a 26,581-byte client; five YAML catalogs supply the shipped presets.
packages/trueforge-ui/ (559 files, 2,776 KB) and packages/frontend/ (24 files)
The embeddable interface: 58 files directly under src/atoms/, four layout modes (dock, drawer, sidebar, widget), a slot and theme system with a 26,577-byte theming document, and a changelog of 68,869 bytes beside a 25,328-byte README and its own contributor and security files — the shape of a package that used to live somewhere else. The frontend shell next to it is a small Vite and React application with login and logout screens, and it ships inside the server package.
packages/assistant-ui-runtime/, packages/trueforge-sdk/ and python/trueforge_sdk/
The adapter and the two generated clients: 75 files of React runtime glue whose server/types.ts is 39,681 bytes and is the single definition of the server ports, plus a 42,394-byte hook file and a 100,655-byte test; then a TypeScript SDK of 1,010 files and a Python SDK of 403, emitted from .github/fern/openapi/openapi.json (333,161 bytes) with wire tests for every resource.
docs/ and benchmark/
Fourteen files at the top of docs/ — docs.json, a 333,161-byte openapi.json and twelve pages — with three API pages, 22 UI SDK pages, five capability pages, one page each for agent creation, authentication and harness setup, 54 screenshots totalling 21,702 KB and 11 more assets at 1,960 KB. benchmark/ is a Python harness of 11 files whose largest script is 16,205 bytes, and it is what the README’s comparison with Claude Managed Agents and deepagents is built from.
Root files, workflows and deployment
A root AGENTS.md of 5,822 bytes with nine nested counterparts, nine CLAUDE.md files of 11 bytes each, .cursor/BUGBOT.md plus three rule files and two package-level bugbot documents, eleven pending changesets, nine workflows led by a 20,583-byte release file and a 12,319-byte CI file, CONTRIBUTING.md at 16,775 bytes and RELEASING.md at 13,607, a two-stage Dockerfile with an npm variant, two compose files, a Helm chart whose values file is 15,704 bytes and whose helpers template is 33,412, and a .railway directory with a 3,927-byte infrastructure file.

Choices, and what they beat

  • Sessions, turns and a streamed event log as the public contract over a prompt-in, response-out call

    The concepts page defines the hierarchy as one agent to many sessions to many turns to many events to some deltas, and then draws the consequences: turns chain automatically because previous_turn_id defaults to auto, so the caller never resends history; only one turn runs in a session at a time; every event carries an id, a thread id and a sequence number, which is what makes resuming after a disconnect possible; and deltas exist only on the live stream, because a turn’s events are already merged when they are listed afterwards.

  • One codebase, SQLite locally and Postgres when hosted over picking a single deployment target

    The README’s mode table pairs local use with one process, SQLite and no extra infrastructure against teams and multi-replica hosting on Postgres plus Redis; the code keeps two migration trees and one store interface with contract tests run against both, so the difference is a constructor rather than a fork. The documentation is equally direct about the price: standalone mode ignores OIDC even when it is configured, the chart’s dev defaults include a well-known Postgres password and an unauthenticated Redis, and local mode is described as something to keep on localhost.

  • Generated SDKs and an OpenAPI document nobody hand-edits over hand-written clients for each language

    Both copies of the specification are regenerated in CI and required to stay identical, the TypeScript and Python SDK trees are marked as generated in the pull request checklist, and fork pull requests are told to change source only while maintainers regenerate the SDK after merging. Regeneration carries its own changeset, so a generated change is still a versioned change — the same discipline that makes the SDK the contract surface rather than a copy of it.

  • A sandbox that is off by default and provisioned on demand over running agent code inside the server process

    The README states it as a feature and a boundary at once: an isolated environment for code, files and shell commands, off by default, provisioned only when needed, with secrets staying in the harness. It is required for skills and Code Mode, and it is written against a provider interface with three implementations and a shared contract suite, which is the part that makes an outside request for a fourth provider a configuration question rather than a rewrite.

  • Approval by label rather than by gating every tool call over stopping on every write-capable tool

    The default policy names two labels, @write and @destructive, and the page then records the failure mode itself: those labels only match tools the MCP server has labelled, many servers skip labels, and a tool that changes data will therefore run without asking unless it is named or the policy is set to @all. Persistent policies are keyed by tool name, so a rename on the MCP side silently voids one, and the open pull request on paused turns shows the other edge — a stream that stays open, and resources that stay paid for, while it waits for a decision.

Read fromREADME.md (9,423 characters, fetched in full from the GitHub API), AGENTS.md (5,822 characters), docs/api/overview.mdx (13,586), docs/authentication/overview.mdx (12,546), docs/create-agent/overview.mdx (25,194), .github/fern/generators.yml, .changeset/, the nine workflows under .github/workflows/, the two migration trees under packages/trueforge/src/db/, and the complete 2,938-file tree with sizes.

Build log

6 stages
  1. 01

    The unit of work is a session, a turn and a stream of typed events

    What this project actually sells is a vocabulary. An agent is a saved definition — model, instructions, MCP servers, config — and is explicitly not a running process; a session holds one conversation’s context; a turn is one request inside it; a turn emits events onto a stream as JSON objects, and model output arrives as deltas that the client merges into the base event by id. The documented event set is small enough to enumerate: turn.created, mcp.initialize, model.message, tool.response, tool.approval_required and turn.done, with thread.created and thread.done announcing subagent threads; every event carries an id, a thread_id (main for the root agent, a generated id for a subagent, null for run-level events) and a sequence number the docs say exists to resume after a disconnect. Turns chain automatically, because previous_turn_id defaults to the string auto, so an application stores a session id and never resends history. One turn runs in a session at a time, and a turn can stop in three places — an approval, a clarifying question (tool.response_required) or MCP OAuth (mcp.auth_required) — each of them resumed by sending a new turn. The code behind that vocabulary is where the weight sits: SessionHandle.ts is 19,965 bytes, TurnHandle.ts 18,924, the store interface 17,814, and the shared store contract suite 128,154.

  2. 02

    Ten weeks, 778 commits, and a version line that keeps being restarted

    The first commit is dated 2026-07-23 and reads "Initial commit: extract agent harness into pnpm workspace." — which is the whole origin story, a harness lifted out of a larger product and given its own repository. The commit curve is 57 in July, 412 in August and 309 in September, 778 in total, with the newest commit on 2026-09-30. Around it sit 6,033 stars, 486 forks, 17 watchers, 99 open issues and a repository of 44,151 KB holding 2,938 files. Releases are less tidy. Six things are published at once — @truefoundry/trueforge, trueforge-core, trueforge-ui, trueforge-sdk, assistant-ui-runtime, a Python trueforge_sdk — plus the Helm chart, and their versions do not move together: at the end of September the server was at 0.3.1, the UI at 0.4.1, the SDK at 0.2.1 and the chart at 0.3.0, each with -rc lines behind it. On 2026-09-29 a pull request titled "Reset versions to 0.0.0" set every package and both copies of the OpenAPI document back to 0.0.0, and the Changesets bot then opened release pull requests moving them to 0.0.1 and 0.4.0-rc.1. The visible tags are a third line again — v0.1.1 through v0.1.10-rc.1, next to per-package tags such as @truefoundry/trueforge-ui@0.3.0-rc.11 — and a release-workflow test plan still asks for a re-run on the branch release-v0.176.0 for 0.176.0-rc.1, a version the current packages never reach.

  3. 03

    Three bugs in the release pipeline, and a body of written law

    The most useful pull requests here are about the pipeline rather than the agent. One found that the release workflow wrote skip=false and passed it between jobs, and that GitHub Actions drops a job output whose value is the string false — so both image builds and the Helm chart publish were silently skipped; the flag became build or skip. Two more describe a job skipped because an earlier one was skipped, since a dispatch skips select-mode and a push skips version-release, leaving downstream conditions false; image builds, Helm, npm and PyPI publish now use !cancelled() and name the jobs they depend on. The fourth is a production image that had quietly grown: pnpm fetch in the store stage filled /pnpm/store but also left a full node_modules/.pnpm tree that later stages carried into the runtime image together with build-only packages such as esbuild and TypeScript, so the step became pnpm fetch && rm -rf node_modules. Beside those fixes sits an unusual amount of written law. AGENTS.md is 5,822 bytes of mandatory rules: no assertion escapes (as T, as unknown as T, non-null !, as never); an error thrown from a catch must carry { cause: caught }; every type, schema and helper has exactly one canonical owner, with no forwarding shims; server-port types exist only in packages/assistant-ui-runtime/src/server/types.ts and the UI package may re-export aliases and nothing else; environment reads go through a 54,846-byte config.ts; UI lengths use rem, not px; comments explain intent and must not carry tracker ids. Every nested AGENTS.md also needs a sibling CLAUDE.md containing only the line @AGENTS.md, so that Cursor and Claude Code load the same scoped rules — nine of each, and every one of the nine CLAUDE.md files is 11 bytes.

  4. 04

    Two databases, one contract, and a login mode that is deliberately absent

    Every store exists twice. packages/trueforge/src/db/postgres/ holds 38 migrations running from 2026-07-27 to 2026-09-29; packages/trueforge/src/db/sqlite/ holds 36 from 2026-07-30 to the same week, and the two trees stay in step by subject — sessions, agents, MCP OAuth, skills, sandbox providers, schedules, session metrics, sandbox environments, with the one Postgres-only GIN index on session metadata explaining part of the difference. What keeps them honest is a set of contract tests written once and run against the in-memory store, then Postgres, then SQLite: storeContractSuite.ts alone is 128,154 bytes, with sibling suites for the agent, MCP server, model provider and OAuth token stores, six jest configurations in the server package, and two of those six reserved for a local sandbox. That split is the product’s main deployment decision, and the documentation states the cost. Local mode is one process on SQLite with no login at all; standalone mode ignores OIDC even when the variables are set; and the README says local mode is not a production or internet-facing setup, asks that it be kept on localhost, and disclaims responsibility for data loss or unauthorized access beyond that. The authentication document adds that the Helm chart defaults to no login, a well-known bundled Postgres password and Redis without authentication, then lists two holes instead of hiding them: session history is owner-only, so admins are not global superusers today, and agents created by anyone are visible to everyone on the instance.

  5. 05

    A sandbox that is off by default, and approvals that trust a label

    Sandboxing is opt-in and lazy: the docs say it is off by default, provisioned only when an agent needs it, and that secrets stay in the harness rather than the sandbox. Skills and Code Mode both require it, and it is written against a provider interface — DaytonaProvider.ts at 21,556 bytes, a second provider backed by TrueFoundry itself at 11,818, and a local provider at 27,969 with a Lima configuration, a loopback probe and a smoke script — which is why an outside request to support NVidia’s OpenShell could be filed as an alternative to Daytona and answered as an extension point rather than a rewrite. Skills are git-backed SKILL.md packs mounted into the sandbox by a 29,098-byte downloader, and Code Mode talks to the sandbox over NATS. The human checkpoints have the more interesting documentation, because it records its own hole: tool approval defaults to require_approval_for_tools of ["@write", "@destructive"], and the page warns in a note that those labels only match tools an MCP server has labelled — many servers skip labels, so a tool that changes data can run without asking, and the remedy is to name the tool or set ["@all"]. A recent pull request adds persistent approval policies keyed by server and tool, records that "Approve once" is never stored, and admits that renaming a tool on the MCP side leaves an existing policy matching nothing. An open one would change the shape of a paused turn: instead of finishing with a result, the stream would emit paused, hold the connection and its resources open until an abort, and stay paused rather than closing if the consumer simply walks away.

  6. 06

    A company repository, a bot-heavy queue, and two outsiders

    The work belongs to a company and is spread over 38 accounts. Chirag Jain (chiragjn) has the most commits at 119 of 778, ahead of debajyoti-truefoundry at 77, bhaveshpatel640 at 75, govindavashishtha at 67, sr07asthana at 59 and thesujai at 54; three of the accounts are bots and twelve more have exactly one commit each. 386 commits carry a co-author trailer — 155 reading only "Cursor", 49 from the release bot, 43 from a project bot named trueforge-dev-bot, 19 naming Claude Opus 4.8 — the paper trail of a history written with agents in the loop. The queue is heavily instrumented: the thirty most recent issues and pull requests contain three Dependabot groups (one proposing 51 version bumps at once, two for a single undici security patch, one closed with the note that the dependencies were no longer updatable), CodeQL and image-scan workflows, a changeset bot on every pull request, and Cursor summaries and agent links embedded in the bodies. Two contributions in that window come from outside the company and are both worth reading. One reports that parseSandboxArtifacts in SandboxArtifactDownload.tsx uses the regular expression ([^)]*), so the path /tmp/report(final).csv is truncated to /tmp/report(final before the download URL is built; the reporter proposes a parser that recognises balanced parentheses, explains that a broader expression would swallow neighbouring links, and asks a maintainer to confirm the scope and assign the issue to him. The other asks for NVidia’s OpenShell to be supported as a sandbox provider. A third item is internal: a security pull request closing an INFOSEC disclosure by adding state to a browser cookie so that an MCP OAuth callback binds to the user who started it. The project also publishes its own benchmark, comparing TrueForge with Claude Managed Agents and deepagents on the same tasks, tools and model and claiming the same accuracy at lower cost, with the scripts to reproduce it committed under benchmark/.

Adjacent records

All records →