Skip to content

cocoindex-code

A command-line tool and MCP server that keeps a local semantic index of a codebase — tree-sitter chunks, embedding vectors in SQLite — so a coding agent can find code by description instead of reading files, with a background daemon holding the model and an index that only re-embeds what changed.

Screenshot of cocoindex-code
Editor screenshot, 1 Oct 2026cocoindex-code ↗

What it is

A Python command-line tool that builds a semantic index of a codebase so a coding agent can search it by description. The ccc command initializes, indexes, searches and diagnoses a project, and can run as an MCP server; a second command, ccc grep, does structural pattern matching with no index, daemon or embeddings involved. Indexing is an application written on CocoIndex v1, the Rust data-transformation engine from the same team: files are matched with include and exclude patterns, .gitignore rules and a size limit, split at tree-sitter boundaries around a target of roughly a thousand characters, embedded locally by sentence-transformers or remotely through LiteLLM, and written to SQLite with vector search from the sqlite-vec extension. A long-lived daemon keeps the model in memory and serves both the CLI and the MCP server. It is published as a plugin marketplace for Claude Code and Grok, an extension for Oh My Pi, and a skill; the README claims a seventy percent token saving and one-minute setup, and the package describes itself as Alpha under Apache-2.0.

Who built itThe repository belongs to the GitHub organization cocoindex-io, the team behind the CocoIndex data-transformation engine this tool is built on, and the README signs off with a maintainer address on the same domain. Its 249 commits come from 23 accounts and two people carry almost all of them: georgeh0 (Jiangzhou He, 133 commits, nearly all under a jiangzhou@cocoindex.io address) and badmonster0 (Linghua Jin, 35 commits under linghua@cocoindex.io, the address the README gives for enterprise help). Next come mareurs (Marius Ailinca, 23) and the dependency bot (22), so the outside share is small and specific. Of the 77 co-author trailers counted, 53 name a Claude model — Claude Sonnet 4.6 on 15, Claude Opus 4.6 (1M context) on 15, Claude Opus 4.6 on 9, Claude Fable 5 on 8 — alongside 22 from the dependency bot, one for a bot named Clawdbot and one naming a maintainer.

How it is put together

The parts · 6

The shape is a daemon and a thin client wrapped around one CocoIndex application. The expensive object — the embedding model — is loaded once into a long-lived background process listening on a Unix socket, a named pipe on Windows; the command-line tool and the MCP server connect for a single request, hand over a message-pack-encoded tagged union defined in one protocol module, and close. Indexing is expressed as a function whose results are memoized by content, so a second run re-embeds only what changed, and because the memo state lives in a database file rather than in the process, a restart does not by itself invalidate anything. Search is a query against a SQLite table of vectors read through an extension, with two plans: a nearest-neighbour plan when no path filter is given, and a full scan when one is. Everything the agent touches is deliberately thin — the command-line layer holds no logic beyond argument handling and delegation — while the parts that must be fast live in the Rust engine underneath, which is also why the project can offer a second search mode that needs no index, no daemon and no embeddings at all.

src/cocoindex_code/
Twenty files, 214 KB, in one flat package. The command-line layer is the largest at 39,570 bytes, then the daemon at 31,661 and the client at 27,130, with settings and path resolution at 26,226. The rest is one concern per file: the structural search command at 16,211 bytes, the MCP server at 13,760, the per-project wrapper around a CocoIndex environment at 12,603, file matching at 7,819, the message protocol at 7,305, shared context keys and the embedder factory at 6,614, the vector query at 6,021, the embedder defaults and parameters, daemon socket paths, the rate-limited cloud embedder, the indexing application at 3,615, the chunker registry at 1,133 and a small schema module.
tests/
Twenty-three files and 221 KB, weighted towards end-to-end tests on user-facing commands rather than units: the largest single file is 37,690 bytes, the settings and file-matching suite is 33,724, and there are dedicated suites for command-line helpers, structural search, the client, the daemon, daemon idle behaviour, the protocol, and two end-to-end paths. Docker end-to-end tests exist as a separately marked set that the default test command excludes, and the project’s own guidance says that when a test fails the underlying issue is fixed rather than the test skipped.
scripts/
Two files that exist to answer one question with evidence: a 15,274-byte script that queries the live MTEB results dataset and renders a ranked table of embedding models with parameter counts, architecture and an estimated CPU speed, and the 7,236-byte report it produced, committed so that the numbers in the README can be traced back to a command.
skills/ccc/ and the plugin manifests
The agent-facing surface: a 3,298-byte skill that tells the agent it owns initialization, indexing and searching for the project and should not ask the user to do any of it, plus two references — one on settings at 4,493 bytes and one on management and troubleshooting at 3,501. Around it sit two marketplace catalogs, for Claude Code and for Oh My Pi, each with a manifest, and a one-line MCP server definition that both of them can read.
hooks/ and extensions/
How the index stays current without being asked. A 947-byte hook file was registered on session start and after edits so that an incremental index runs when the project directory exists, and because one agent platform does not execute that file format, the same behaviour is reimplemented for it as a 2,518-byte TypeScript extension, with a small manifest declaring which events it listens to.
docker/, .github/workflows/ and the documents
A 4,719-byte image definition with two variants — one that installs only the cloud path and one that bundles the local model stack — plus a compose file, an entrypoint that aligns file ownership on mounted paths, and a 7,678-byte release workflow against a 1,616-byte check workflow. The documents are the README at 41,991 bytes, an embedding guide at 12,958 and a 7,677-byte file of guidance for coding agents working in the repository.

Choices, and what they beat

  • An embedded SQLite index instead of a database to install and run over requiring a vector database as a separate service

    The feature list states it as “Embedded: Portable and just works, no database setup required!”, and the whole index is two files inside the project directory. What the choice costs is written down rather than hidden: the incremental state sits in an LMDB database whose maximum size is fixed when the daemon starts and defaults to 4 GiB, so a large repository can hit an environment mapsize error, and the documented answer is an environment variable plus a daemon restart until an upstream issue lands that grows the map automatically.

  • A long-lived daemon that keeps the embedding model resident over loading the model once per command

    The README says plainly that the daemon holds the model in memory, which is why it exits after a configurable idle timeout, 180 minutes by default, and is restarted transparently by the next command. The trade-off is negotiated per session instead of being fixed: an MCP client sends heartbeats that keep the daemon warm by default, and a documented setting disables only that heartbeat so a long-lived editor session can let the model go between real requests.

  • A structural search command that uses no index at all over sending every lookup through the vector index

    The by-example pattern command matches the syntax tree through the engine’s structural matching feature and, in the README’s words, “runs entirely locally: no index, daemon, or embeddings required”. It is also where the project leans on something not yet released: the same README notes that until that feature ships upstream, the command needs a local build of the engine to run against.

  • Two install flavors, one of them without the local model stack over one package that always pulls in sentence-transformers

    The batteries-included extra brings in sentence-transformers so local embeddings work with no API key, which is the recommended default; the slim install is cloud-only and exists for people who do not want roughly a gigabyte of the local inference stack on the machine. The same split is mirrored in the two Docker image variants published from every release.

  • Teach the agent through a skill first, and offer MCP second over shipping only an MCP server

    The README marks the skill as the recommended integration and says that with it installed no initialization or indexing step is needed, because the agent handles that lifecycle itself; the skill file instructs the agent to initialize, index and refresh on its own rather than asking the user. MCP is documented as the alternative path, with the search tool refreshing the index by default.

Read fromREADME.md (41,991 characters), EMBEDDINGS.md (12,958 characters), CLAUDE.md (7,677 characters), pyproject.toml, scripts/find_best_models.py, scripts/MTEB-RANKINGS.md, .github/workflows/release.yml, .github/workflows/pre-commit.yml, skills/ccc/SKILL.md, and the complete 81-file tree with sizes.

Build log

6 stages
  1. 01

    Four months of building, and a release line that stops in August

    The repository was created on 2026-02-01 and its oldest commit is dated 2026-01-31, titled feat: initial version. It has taken 249 commits in eight months and the shape is front-loaded: 90 in February, 73 in March, then 20, 7, 15, 27, 13 and finally 3 in September, the last month with any activity at all. Twenty releases are listed, none of them a draft or a prerelease, running from v0.2.22 on 2026-04-09 to v0.2.41 on 2026-08-07. The gap at the end is the plainest fact about the project’s rhythm: the last push is dated 2026-09-22, six weeks after the last release, so work continued past the point where versions stopped being cut. Release mechanics are automated end to end — publishing a GitHub release builds the distribution, pushes it to PyPI through the standard publish action, attaches the same files back to the release, and builds Docker images for two architectures on both Docker Hub and GHCR with a registry-backed cache. One comment in that workflow records a race that actually bit: the image installs the tool from the checked-out source tree rather than from PyPI, because a just-published version had not propagated through PyPI’s CDN at the v0.2.24 release, and building from the tag also guarantees the image matches it. Around all of it sit 2,729 stars, 225 forks, 17 watchers and 44 open issues.

  2. 02

    How a file becomes a chunk, and where the index is kept

    Indexing is one CocoIndex function, process_file, decorated so its results are memoized by content; a second run re-embeds only what changed. It walks the project through a matcher that applies include patterns, exclude patterns, nested .gitignore rules and an optional max_file_size, detects the language from the extension with per-extension overrides available, and splits the text — through a custom chunker registered in settings.yml for that extension, otherwise through the default recursive splitter. The splitting is what the README argues for: chunks are cut at tree-sitter boundaries, so a function or a class tends to stay whole, and the target is about a thousand characters, roughly three hundred tokens, chosen so a chunk fits the 512-token window most local models have while cloud models offer eight to thirty-two thousand. The embedder is created once when the daemon starts, and two parameter dictionaries travel with each request so asymmetric models get different arguments for documents and for queries. Storage is two files in the project’s .cocoindex_code/ directory: an LMDB database holding the incremental state, and a SQLite file holding the vectors, searched through the sqlite-vec extension. Their location can be remapped by path prefix for containers, because LMDB does not behave well on a bind mount.

  3. 03

    The percentage in the headline, and the method behind the numbers beside it

    The README leads with a percentage: “Instant token saving by 70%.” Nothing in the README says how it was obtained — no benchmark, no dataset, no command to reproduce it — so it is an authors’ claim, and this record carries it as one. The repository does ship a reproducible entry point for the neighbouring question of model choice: scripts/find_best_models.py, a single Python file with inline dependencies, queries the live MTEB results dataset on Hugging Face and renders a report committed in the tree as scripts/MTEB-RANKINGS.md. It regenerates with one documented command, uv run scripts/find_best_models.py --clear-cache --output MTEB-RANKINGS.md, and stamps its freshness, here a dataset updated on 2026-06-23. What it is careful about matters more than its table: CPU speed is labelled “estimated from parameter count and architecture”, and the figure most often quoted — that a decoder model of the same parameter count can be three to ten times slower on CPU — is an estimate, not a measurement, hardcoded in the script’s architecture table and repeated in the README as an expectation. The one measured latency figure comes from an outside contributor: an issue filed on 2026-09-09 reports that on a synthetic dataset of 100,000 chunks a broad path filter costs about 450 ms at 384 dimensions, while queries without a path filter use the index’s nearest-neighbour plan.

  4. 04

    A crash report at 19:34, a release at 00:15, and a disagreement

    The clearest exchange in the material starts with issue #270. A user running v0.2.39, installed through pipx, reported that any ccc search with --path killed the daemon with TypeError: unsupported operand type(s) for *: "NoneType" and "NoneType", that --lang alone worked, and that the client only re-raised whatever the daemon sent back. He offered a root cause: path-filtered requests are dispatched to a full scan, and a NULL distance was reaching the score calculation. A fix landed the same evening — pull request #271, which carries daemon-side tracebacks to the client by populating an error field, consolidates five duplicated raise sites into a single _daemon_error() helper, and rejects rows whose distance is NULL, reasoning that the column is hidden and populated only under the nearest-neighbour plan, so a NULL means the plan that ran is not the plan that was asked for; the full-scan path needs no guard because it computes the distance itself. Version v0.2.41 was published at 00:15, under five hours after the report. Ten minutes later the maintainer replied that he could not reproduce it, that the diagnosis looked off because the distance function raises on bad input rather than returning NULL, and asked the user to upgrade and re-run for a real frame. The user did, confirmed the traceback fix worked, and reported the same TypeError still there.

  5. 05

    Two inputs the memo key had forgotten

    The most substantive outside contribution is a pair of reports about incremental indexing being wrong, not slow. The first observes that process_file is memoized but two of its inputs were never part of the memo key: the language overrides, read from the settings file inside the function, and the custom chunker, which arrives through an untracked context key. Editing either left unchanged files on their old chunks and language, and a daemon restart did not help: memo state persists in the database file. The report quotes the README’s own promise back at it: after editing those settings, no index deletion and no daemon restart are needed. The proposed fix adds a fingerprint over the effective language overrides and each registered chunker, passed into the memoized function only to enter the memo key. A second report explains why the slow path is reached more often than the flag suggests: a search from a subdirectory scopes itself to it, and any path filter sends the query to a full scan. A third open pull request is a memory bug: the chunker treats its size setting as a target instead of a bound, so a line with no recognised separator comes back whole however long it is — sixty thousand characters on one line is ordinary in a scene file — and the embedder pads a batch to its longest member, so one file can kill the daemon. All three were still open at the end of September.

  6. 06

    Who reviews, and how long a good pull request waits

    The contributor list has 23 names, but there is effectively one reviewer, and the queue is visible in the threads. On a fix making the diagnostic command respect the file size limit, the answer is a thank-you and a pointer: “thanks @shixi-li, @georgeh0 can help take a look!” That pull request, opened on 2026-08-10, came back more than six weeks later — the branch clean and mergeable, all four continuous-integration checks green, the suggested reviewer still not having looked — and is still open. A proposal for a dry-run preview of the next index met a more useful reception: the maintainer asked what the motivation was, warned that it would list new and deleted files but not updated ones, which “may be unexpected and misleading for many cases”, and pointed at CocoIndex’s own shadow-run and preview API work; the contributor said his case was discovering where unexpected files came from — another worktree created by an agent — and revised the pull request. Outside work does land: Elixir support, the Oh My Pi marketplace catalog, which drew a “Thanks a lot!”, a fix for daemon sockets whose path exceeded the system limit on a deep home directory, and a setting that lets the daemon exit while an MCP client stays connected. One issue shows the other side: a user asking the maintainers to stop reading .gitignore as a rule source for the application, closed with no reply recorded.

Adjacent records

All records →