agentacct
Reads the session files your coding agents already write, records what they claim they did over MCP and hooks, prices the tokens from a public price list, and shows the whole thing in a local dashboard that never leaves the machine — keeping an agent’s word for done and a passing check as two separate facts.

What it is
A local-first ledger for coding-agent work. It reads the session files each client already writes — Claude Code’s JSONL transcripts, Codex’s SQLite state and rollouts, Hermes’s state database, OpenCode’s rollups, OpenClaw JSONL, DeepSeek Harness’s compressed logs, Kimi Code’s per-request usage records — and records what the agent says it did over MCP and host hooks, then joins the two on real session ids and labels every join with a confidence. A task becomes a receipt: participating sessions, the recorded outcome, checks with their exit codes, an activity timeline, tokens, and a cost estimate drawn with ≈ because it is a price-list estimate and not an invoice. Nothing leaves the machine: no account, no telemetry, no provider API key, and no attaching to agents it did not start. It ships as a Python CLI and TUI plus a notarized macOS app, and eight clients have integration paths of openly different strength.
Who built itThe repository is credited to the GitHub account mikehasa, which the commit history ties to 210 of its 484 commits; the second account, FZ2000, landed 159, three others one to three each, and 110 commits carry no linked account at all. No personal name appears in the metadata, and the homepage field points at https://trytofu.ai, a page titled “Tofu puts your AI-built app online in 10 minutes”, which describes a different product.
How it is put together
The parts · 6One evidence kernel with several thin edges. Everything a client reports is normalised into an immutable envelope, appended to a durable spool, and indexed into a rebuildable SQLite projection; the derived views — receipts, timelines, usage cubes, discrepancies — are recomputed from that spool instead of being written back into it. The organising rule is that no adapter may promote its own claim: a hook can prove that a tool call happened but not what it meant, MCP can prove that the agent said something but not what a provider billed, and a client log can prove tokens but not the goal. So the two streams stay apart and are joined by real ids carrying a confidence label, and missing, partial and conflicting are product states rather than errors to be smoothed over at query time. Cost sits deliberately outside the kernel: a local, refreshable, pinnable snapshot of a public price list applied to client-reported tokens, with the uncovered case left unknown. A second and separate store holds operational control — attempts, approvals, budgets, schedules — under the rule that agentacct controls only work it started itself, which is why the capture layer fails open and the whole product can be read without granting it any authority.
- src/agentacct/
- 93 files, 4.1 MB. A 553 KB
cli.pyand a 454 KBclient_usage.pycarry the command surface and every per-client importer; beside them sitwork_ledger.py(233 KB),evidence_store.py(234 KB),api.py(216 KB),tui.py(178 KB),hooks.py(143 KB),mcp.py(126 KB) and the pricing paircost.py(40 KB) andpricing_catalog.py(26 KB). Two small subpackages hold the evidence-capture adapters, manifests and registry, and the local connectors. - apps/agentacct/
- 305 files, 43.6 MB: a SwiftUI macOS app in which
SetupModel.swiftalone is 161 KB, withWorkPane(135 KB),DashboardPane(100 KB),ReceiptsPane(81 KB),WorksetsPane(80 KB) andTheme(52 KB) behind it. Its tests are snapshot-based, with light and dark reference images checked in pairs under a platform-named directory. - tests/
- 167 files, 3.6 MB, dominated by
test_client_usage.pyat 302 KB,test_mcp.pyat 164 KB,test_evidence_store.pyat 104 KB and a 99 KB tokens-explorer test. The suite runs on Python 3.11, 3.12 and 3.13 in CI, and release pull requests quote it — 4,895 passed with 3 skipped on the 0.12.7 tree. - docs/ and INSTALL.md
- Twelve documents under
docs/(193 KB) plus a 33 KBINSTALL.md: a 33 KB reference, a 28 KB per-client integration guide, a 24 KB usage and cost truth table, the coverage matrix, the architecture note, safety boundaries, a privacy threat model and the evidence RFC. The coverage matrix is generated from the capability manifest in the code byscripts/gen_coverage_matrix.py, so the published promise and the implementation cannot drift quietly. - packaging/ and .github/workflows/
build-dmg.sh,freeze-cli.sh(PyInstaller),verify-dmg.sh,source-provenance.shand a CLI payload validator, with four workflows including a 5.8 KB publish job that refuses a tag disagreeing with the packaged version. The app derives every version string frompyproject.tomlrather than keeping a second one.- design-plans/ and animation-plans/
- The working papers: a data-quality plan with seven documents and four audit or fixture scripts, a native-macOS overhaul with three large fixtures, a session-steps readability study across 100 numbered scenarios in three files, a usage-limits redesign with 100 numbered review files plus critique and synthesis notes, and a 215 KB scenario matrix under the animation plans.
Choices, and what they beat
Two evidence streams joined by id and labelled, not one merged number over folding usage and recorded work into a single figure
The README states it as a rule — “Missing beats wrong” — and says every join between usage and recorded work carries a confidence of
exact,high,mediumorlow, with an unproven link shown as a gap rather than a zero. The usage truth table extends the same idea to sources: MCP proves what an agent said, a local import proves what was parsed, and neither becomes provider billing.Cost from a public price list, always labelled an estimate over claiming provider or subscription billing
The README says costs are pricing-table estimates marked
≈and that there is no invoice access; the truth table adds that a subscription or coding-plan user is not charged per token, that the table is a snapshot of a community-maintained list rather than a provider price sheet, and that a model id the table does not cover stays unknown instead of taking a nearby price.Control only the work agentacct starts itself over attaching to, or supervising, agents already running
The safety document lists what it will not do — no scanning the machine for existing agent processes, no attaching to existing sessions, no pausing, killing or inspecting processes it did not start, no editing global client configuration by default — and the capture layer is fail-open, so a broken or moved adapter can never block a tool call in the host agent.
Unavailable instead of zero where a client reports no tokens over showing a measured zero
The Cursor lane reads composer identity, timestamps, explicit model metadata and child links only; the coverage matrix says it never fabricates token usage, cache usage, cost, titles or projects, and the integration guide says missing usage stays unavailable rather than becoming a measured zero.
Keep the pre-rename names accepted after two renames over a clean break with the old identifiers
0.5.0renamed the package and the MCP tools and0.5.1changed the environment prefix; the changelog states that the older names stay accepted forever, and that when two aliases are set to different values agentacct refuses rather than silently picking one. The project store directory still carries the first name,.agent-sentinel/state.
Read fromREADME.md and README.zh-CN.md, docs/architecture.md, docs/reference.md, docs/usage-truth-table.md, docs/coverage-matrix.md, docs/coding-agent-integrations.md, docs/safety-boundaries.md, docs/multi-source-evidence-architecture.md, docs/agentacct-workflow-instructions.md, CHANGELOG.md, the release pull-request bodies, and the file tree with sizes.
Build log
6 stages- 01
Two months, 42 versions, and two renames in the first week
The repository was created on 2026-07-24 and released
0.1.0the next day: a local-first ledger that imported client-reported tokens from local session files, took work context over MCP, joined the two with a confidence label and drew it on a local dashboard. Sixty-eight days later the changelog carries 42 version sections, ending at0.12.9on 2026-09-30, and the commit history 484 commits — 210 credited to the owner’s account, 159 to a second contributor, 110 with no linked account, and five spread across three others. It is not a lone hand at the keyboard: 79 commits carry co-author trailers, and 74 of those name a Claude model, the largest group being the 59 that say “Claude Fable 5.1”. The project also renamed itself twice in its first week.0.5.0renamed the MCP tools fromsentinel_*toagentacct_*and the Python package fromagent_chronicletoagentacct;0.5.1madeAGENTACCT_*the primary environment prefix while keeping the older names accepted forever. The trace that survived is the project store directory, still called.agent-sentinel/state. - 02
A work step is something the agent writes down, and evidence outranks wording
agentacct does not watch an agent and guess what it did; the step record is written by the agent itself. Over MCP,
agentacct_record_sectionopens a section asstarted, addscheckpoints, and closes it ascompleted,blockedorhanded_off; machine checks such as a test run are recorded separately with their exit code, and the installed hook bridges add tool-activity ticks and workspace-relative file metadata that never include prompts, responses, thoughts or tool bodies. Whether a check counts is decided by how it was observed rather than by the agent’s wording — text such as “tests passed” is never parsed into a pass — and a finished task is rendered either Reported, meaning the agent said so, or Verified, which requires a passing check that post-dates the last recorded change. Two open pull requests tighten that contract further. Number 324 would make a 20–260 characterprogressnote mandatory when a section closes as completed or handed off, with its last clause forced to open with one of six set words, and would print a corrected example call with every refusal; number 335 would put the agent’s own account — goal, newest progress note, next step — above the counted metrics on the receipt, labelled Agent reported. Every join to usage carriesexact,high,mediumorlow, and an inherited session id can never beexact. - 03
Seven local formats, one ledger, and one client that reports no tokens
Token truth comes only from the files the clients write themselves, and no two write the same shape. Claude Code is read from the JSONL transcripts under its projects directory. Codex is read from
state_5.sqliteplus rollout JSONL, where raw input already includes cached input, so the importer subtracts reported cache reads and writes to normalise a non-cached figure and refuses to add reasoning twice, which is already inside output. Hermes comes from session rows instate.db, OpenCode from per-session rollups in the nativeopencode.db, with exported JSON events as fallback, OpenClaw from assistant rows in JSONL, DeepSeek Harness from Zstandard-compressed logs under~/.dshthat record tokens but no cost, and Kimi Code from asession_index.jsonlindex plus per-requestusage.recordevents in each session’s wire streams, which are deltas and never cumulative, so the importer sums them. Cursor is the deliberate exception: its primarystate.vscdbyields composer identities, timestamps and child links only, and an unavailable token figure stays unavailable rather than becoming a zero. Support is published per capability rather than per logo: sessions, usage, mechanical capture, MCP semantics, model attribution, cache reads, cache writes and install are rated on their own, and five further agents sit on the roadmap with nothing behind them. - 04
Where the money number comes from, and what it is not
The money figure is an estimate and is labelled as one. Tokens are
client_reported, read out of the client’s own session store. Cost isestimated_from_tokens: those tokens multiplied by the project’s local snapshot of LiteLLM’s public, community-maintained model price list, drawn as≈$, while~$marks a known-partial subtotal. The built-in table is deliberately small and stops atgpt-5.5; a model id no row covers stays cost-unknown instead of taking a nearby price, cache reads price at 0.1 times input and cache writes at the input rate, and three aliases map a client’s own route names onto the table —claude-codetoanthropic,codextoopenai, and DeepSeek Harness’sdeepseek-officialtodeepseek. The snapshot refreshes itself once it is older than seven days, best-effort, with a failed attempt throttled to one an hour so an unreachable network never blocks an import. The documentation is blunt about what the number is not: a subscription user is not charged per token, so the estimate is not what they paid, and there is no invoice access at all. A stored row also does not heal itself, because a scan reprices only what a catalog row now covers — which is why one patch adds a command that says why a row has no cost, and why a missing price for a single Codex model shipped as its own release. - 05
The maintainer’s own store: 22.8 GB of evidence, and a release process written after it failed
The maintainer runs it on the machine it is written on, and that store’s numbers are in the changelog. The append-only evidence spool had reached 22.8 GB while the projection built from it held 277 MB, and sampling showed about 94% of those bytes were dead shadow rows left behind by pruning. A verified cold compaction landed in number 328, but proving it was the hard part: equality against the live projection can never hold on a store that has been pruned, and a from-zero replay of the 3-million-row spool runs at roughly 137 rows a second — about six hours. The gate became a containment proof, and then had to exclude rows the live usage lane writes during the offline window. After the repair the maintainer’s store went from 21.24 GiB to 131.7 MiB. The same upkeep shows in the tests: number 340 recalibrated a display pin that had drifted from under 15% to 22.4% of 3,592 unique recorded titles, and number 342 chased a failure that appeared only because a two-day fixture sat across different calendar days and the weekly buckets are Monday-anchored. Release pull requests must state whether the macOS disk image will be rebuilt or reused, since the app embeds a frozen copy of the Python CLI — a rule written after four releases in a row shipped with no disk image at all.
- 06
What it says it cannot do, and the small round of attention from outside
agentacct calls itself early alpha and writes its limits where they can be checked. It does not scan for or attach to agents it did not start, does not copy prompts, responses, thoughts or transcripts, stores no provider API key, and offers no hosted service; Windows is supported only through WSL, and the signed macOS app exists for people who would rather not install Python. The coverage matrix admits that one client can hold a proven usage lane and an experimental capture lane at the same time, and it uses that: for Kimi Code the hook bridge has fired on live desktop sessions — 64 payloads across five — but no live session has been seen binding a recorded section to its session id, Cursor proves presence and nothing about tokens, and OpenClaw’s routing metadata is not integrated. Dates are whole session totals attributed to an activity date, which the API exposes as a limitation instead of pretending to a daily split. The repository keeps the paper trail: 100 numbered review files under the usage-limits redesign, a critique and a synthesis document beside them, a 215 KB scenario matrix in the animation plans, and a data-quality plan with its own audit scripts. Attention from outside has been small and specific — a contributor’s readability cleanups, credited by name in the changelog, a fix for a stale export file from another account, and one inbound integration request, issue 349 from the founder of MemCode, asking whether an importer for memory save-and-recall metadata would be accepted and promising never to ingest memory text, prompts, credentials or source files.
Adjacent records
All records →No. 077
Lody
A workspace where a team shares the coding agents it already runs: connect a machine, bring Claude Code, Codex, Kimi or any other agent that speaks the protocol, and dispatch work from desktop, phone, web or terminal while sessions delegate to each other and the code stays on the machine its owner connected.
No. 073
Zeron
A Rust desktop app that runs the coding agents you already use — Claude Code, Codex, Cursor, Devin, Grok, Hermes, Pi and Antigravity — on your own machine, with no account required, and syncs the sessions to your other devices only if you sign in.
No. 070
OpenChatCut
A local-first video editor whose editing surface is a conversation: the built-in agent and external Codex or Claude Code sessions call the same editing tools the interface itself uses, so every change lands on a real multi-track timeline as a clip, transition, caption, effect or audio item that can still be dragged, undone and exported. Projects and media stay on the machine, and preview and final render both come out of Remotion.