
Moonshine Voice
A speech stack for voice agents and interfaces — transcription, intent recognition and text to speech — built to run at low latency on small devices, with 74 of its recent commits co-authored by Cursor.
GitHub avatar of moonshine-ai, not the project’s own logo — taken from github.com on 2026-10-02.

What it is
Moonshine Voice is a speech stack for building voice agents and interfaces: very low latency speech to text, intent recognition and text to speech. It ships as a C++ core with per-platform builds and a micro/ tree for small models, alongside documentation that takes positions rather than only explaining APIs — word-level timestamps, execution providers, diarization models, and a page comparing it with Whisper.
Who built it594 of the repository’s 666 commits are his, against Kyle Howells on 26, Evan King on 17, Giovannini Barbosa on 9 and a long tail of one- and two-commit contributions. 75 commits carry a co-author trailer: 74 of them name Cursor and one names a Claude model.
How it is put together
The parts · 4A C++ core with vendored dependencies and a separate tree for the small models, so the whole stack can be built and run without a package manager or a network round trip. The three capabilities — transcription, intent recognition and synthesis — are presented as one stack for voice interfaces rather than three products, and the documentation treats the surrounding questions (timestamps, execution providers, diarization, and how it compares with Whisper) as part of the product.
- core/
- The C++ engine: transcriptions, intent recognition and synthesis, with a third-party tree of 341 files and 63.4 MB vendored in so nothing has to be installed separately.
- micro/models
- Sixteen files and 30.6 MB of small models shipped in the repository, which is what lets a device run the stack locally rather than calling out to one.
- docs/moonshine-vs-whisper.md
- An 8 KB comparison page the project maintains against the model its readers are most likely already running — documentation as a position rather than a description.
- The per-topic documents
- Word-level timestamps at 12 KB, examples and quickstart at 8 KB and 7 KB, the release process at 6 KB, large-file-storage purge notes at 5 KB, and one page each for execution providers and diarization models.
Choices, and what they beat
Vendor the dependency tree into the core over relying on a package manager at build time
341 files and 63.4 MB of
core/third-partyis the cost of a C++ stack that builds without an install step — which is what a voice interface on a device needs.Ship the small models in the repository over downloading them on first use
Sixteen files and 30.6 MB under
micro/modelsare checked in, so latency at first run and a missing network are not failure modes.Maintain a comparison with Whisper as part of the documentation over leaving the comparison to readers
docs/moonshine-vs-whisper.mdexists at 8 KB, which names the incumbent explicitly rather than implying the question.
Read fromREADME, docs/ and the repository tree of moonshine-ai/moonshine, read 2026-10-02.
Build log
3 stages- 01
One author, 594 commits, and a co-author trailer that says Cursor
The repository was created on 2024-10-04 and has been pushed to continuously since; the commit counts by month run from 28 in December 2025 to 143 in July 2026. 666 commits, of which 594 belong to Pete Warden, against Kyle Howells on 26, Evan King on 17, Giovannini Barbosa on 9 and a tail of single commits. 75 commits carry a co-author trailer and 74 of them name Cursor — the largest single AI-attribution count in this batch after the agent-maintained projects, and a plain statement of how the recent work was written.
- 02
The documentation argues instead of describing
docs/is not an API reference. It opens withword-level-timestamps.mdat 12 KB, thenmoonshine-vs-whisper.mdat 8 KB — a comparison page the project maintains about the model most readers will already be using — followed byexamples.mdandquickstart.mdat 8 KB and 7 KB,release-process.mdat 6 KB,lfs-purge.mdat 5 KB, and separate pages forexecution-providers.mdanddiarization-models.mdat 4 KB each. A repository that documents its release process and how to purge large-file-storage history is one whose author expects other people to build on it. - 03
Where the weight is
The tree holds 2,384 files, and the two heavy directories are
core/third-partyat 341 files and 63.4 MB, andmicro/modelsat sixteen files and 30.6 MB. That distribution is the design in one line: the C++ core carries a vendored dependency tree so there is nothing to install, and the small models are shipped rather than downloaded at first use — which is what makes a low-latency speech stack something a device can hold rather than a service a device calls.
Adjacent records
All records →No. 137
Agents Universe
An open-source agent platform that keeps one project context shared by everyone working in it. Agents read the whole knowledge base when a project is opened and write what they learn back into the same files while they work; knowledge is Markdown on disk with a database index behind it, and there is no embedding model or vector search.
No. 075
Colibrì
An inference engine in pure C with no engine dependencies that treats storage, RAM and VRAM as one hierarchy, so 744B to 2.8T-parameter mixture-of-experts models run on hardware people already own.
No. 068
GSD Core
Git. Ship. Done. — a meta-prompting, context-engineering and spec-driven development framework that runs the same five-step loop on every milestone: discuss, plan, execute, verify and ship. The heavy work is pushed into fresh-context subagents so the main session stays lean, and every decision is written into Markdown and JSON under a planning directory instead of living in the conversation.