Skip to content

Moonshine Voice

A speech stack for voice agents and interfaces — transcription, intent recognition and text to speech — built to run at low latency on small devices, with 74 of its recent commits co-authored by Cursor.

GitHub avatar of moonshine-ai, not the project’s own logo — taken from github.com on 2026-10-02.

Screenshot of Moonshine Voice
Editor screenshot, 2 Oct 2026Moonshine Voice ↗

What it is

Moonshine Voice is a speech stack for building voice agents and interfaces: very low latency speech to text, intent recognition and text to speech. It ships as a C++ core with per-platform builds and a micro/ tree for small models, alongside documentation that takes positions rather than only explaining APIs — word-level timestamps, execution providers, diarization models, and a page comparing it with Whisper.

Who built it594 of the repository’s 666 commits are his, against Kyle Howells on 26, Evan King on 17, Giovannini Barbosa on 9 and a long tail of one- and two-commit contributions. 75 commits carry a co-author trailer: 74 of them name Cursor and one names a Claude model.

How it is put together

The parts · 4

A C++ core with vendored dependencies and a separate tree for the small models, so the whole stack can be built and run without a package manager or a network round trip. The three capabilities — transcription, intent recognition and synthesis — are presented as one stack for voice interfaces rather than three products, and the documentation treats the surrounding questions (timestamps, execution providers, diarization, and how it compares with Whisper) as part of the product.

core/
The C++ engine: transcriptions, intent recognition and synthesis, with a third-party tree of 341 files and 63.4 MB vendored in so nothing has to be installed separately.
micro/models
Sixteen files and 30.6 MB of small models shipped in the repository, which is what lets a device run the stack locally rather than calling out to one.
docs/moonshine-vs-whisper.md
An 8 KB comparison page the project maintains against the model its readers are most likely already running — documentation as a position rather than a description.
The per-topic documents
Word-level timestamps at 12 KB, examples and quickstart at 8 KB and 7 KB, the release process at 6 KB, large-file-storage purge notes at 5 KB, and one page each for execution providers and diarization models.

Choices, and what they beat

  • Vendor the dependency tree into the core over relying on a package manager at build time

    341 files and 63.4 MB of core/third-party is the cost of a C++ stack that builds without an install step — which is what a voice interface on a device needs.

  • Ship the small models in the repository over downloading them on first use

    Sixteen files and 30.6 MB under micro/models are checked in, so latency at first run and a missing network are not failure modes.

  • Maintain a comparison with Whisper as part of the documentation over leaving the comparison to readers

    docs/moonshine-vs-whisper.md exists at 8 KB, which names the incumbent explicitly rather than implying the question.

Read fromREADME, docs/ and the repository tree of moonshine-ai/moonshine, read 2026-10-02.

Build log

3 stages
  1. 01

    One author, 594 commits, and a co-author trailer that says Cursor

    The repository was created on 2024-10-04 and has been pushed to continuously since; the commit counts by month run from 28 in December 2025 to 143 in July 2026. 666 commits, of which 594 belong to Pete Warden, against Kyle Howells on 26, Evan King on 17, Giovannini Barbosa on 9 and a tail of single commits. 75 commits carry a co-author trailer and 74 of them name Cursor — the largest single AI-attribution count in this batch after the agent-maintained projects, and a plain statement of how the recent work was written.

  2. 02

    The documentation argues instead of describing

    docs/ is not an API reference. It opens with word-level-timestamps.md at 12 KB, then moonshine-vs-whisper.md at 8 KB — a comparison page the project maintains about the model most readers will already be using — followed by examples.md and quickstart.md at 8 KB and 7 KB, release-process.md at 6 KB, lfs-purge.md at 5 KB, and separate pages for execution-providers.md and diarization-models.md at 4 KB each. A repository that documents its release process and how to purge large-file-storage history is one whose author expects other people to build on it.

  3. 03

    Where the weight is

    The tree holds 2,384 files, and the two heavy directories are core/third-party at 341 files and 63.4 MB, and micro/models at sixteen files and 30.6 MB. That distribution is the design in one line: the C++ core carries a vendored dependency tree so there is nothing to install, and the small models are shipped rather than downloaded at first use — which is what makes a low-latency speech stack something a device can hold rather than a service a device calls.

Adjacent records

All records →