Skip to content

VibeRave

A fork of Strudel that lets you hold a key, say a genre out loud, and have the music change on the next bar without the beat ever stopping.

Screenshot of VibeRave
Editor screenshot, 28 Sep 2026VibeRave ↗

What it is

A fork of Strudel, the browser live-coding language, with a multimodal agent loop built on top. Hold a key and speak, type a line, or click a preset chip, and the same pipeline turns it into Strudel code and swaps it into the pattern that is currently playing; the audio scheduler keeps the cycle, so the edit arrives on the next bar instead of stopping the music. Speech recognition and code generation are both pluggable, and every backend can be run locally and free — whisper or vosk for speech, Ollama for the model — or pointed at any OpenAI-compatible endpoint.

Who built itHis GitHub display name is `d2VpanQ2MDY`, which is the base64 of his login — and it is also the second identity that appears in this repository’s commit log. His profile describes him as shipping “AI agents, tools, and weird ideas” with “less hype, more working systems”, and gives his location as Mars. Of twenty-one public repositories the most-starred is a writing tool that strips the tells out of AI-generated prose, and two others are records of AI hackathons he attended in Berlin and Paris.

Build log

7 stages
  1. 01

    The loop, and what hot-swapping actually means

    Three input modes — push-to-talk voice, a text box, and ten clickable preset chips — all feed one pipeline: transcribe if the input was audio, send the text with the current session to a model, get Strudel code back, validate it, and replace the pattern that is running. The claim the project is built around is that this never interrupts playback. Everything in the editor is scheduled by Strudel’s cycle clock, and a new pattern is swapped in at a cycle boundary rather than the scheduler being restarted, so an edit lands on the next bar. If the generated code fails validation the swap is reverted and the previous pattern keeps playing, which is the difference between a live tool and a demo: the failure mode is “nothing changed” rather than silence.

  2. 02

    The speech backend that is fast enough to use on a stage

    Speech recognition comes in three swappable backends, and the latency table in the README is the most interesting engineering in the project. Whisper, running locally, takes 700–900 ms and is described as medium accuracy. Pointing at an API takes one to two seconds and is accurate on free-form speech. Vosk, also local, runs in about ten milliseconds — because it is not doing open-vocabulary recognition at all: it is matched against a closed grammar of the commands the tool understands. Accepting a restricted vocabulary is what buys two orders of magnitude, and it is a reasonable trade here, since the things a performer says into a microphone mid-set (“harder”, “drop the bass”, “more reverb”) are a small, known set. The same file also notes that the browser shares Strudel’s AudioContext so that recording does not glitch playback.

  3. 03

    The commit log records the moment the AI started being credited

    Twenty-six of the fifty-nine commits carry a `Co-Authored-By: Claude Opus 4.7` trailer — nineteen naming the model plainly and seven naming the 1M-context variant. The interesting part is where they start. The repository’s first day, 26 April 2026, is sixteen commits with no trailer at all; the first co-authored commit is on 27 April, and from then on most carry one. Whatever the reason for the change — a new habit, a new tool, a decision to disclose — the commit log dates it precisely, which is rare. It is the same kind of evidence as a repository whose commits change author identity partway through: the record of how a project was made is often in the metadata rather than in anything anyone wrote down.

  4. 04

    The fork inherited somebody else’s database key

    The first pull request is titled “drop upstream Strudel Supabase key + hackathon residue”, and it describes a real defect inherited from the fork. A file in the Strudel codebase carried a hardcoded Supabase URL and anon key pointing at the strudel.cc project’s own database, so every short share link or public-pattern lookup issued from a VibeRave deployment was being written to somebody else’s backend. The fix moves both values into environment variables and makes the call sites short-circuit when they are unset. The same commit clears what it calls hackathon residue: an ignore rule for an editor directory, a comment in an audio metrics file reading “good enough for a hackathon demo”, and a vocabulary file whose author field said “voice pipeline / hackathon”. Between that, his two hackathon repositories, and the banner image at the top of the README — a rack diagram of the pipeline captioned “Made in Berlin — built for the rave” — the project’s origin is legible without anyone stating it.

  5. 05

    Every genre came out sounding the same

    The second pull request is a post-mortem on the prompt design, and it is unusually specific about the failure. The agent was producing “same-y, medium-busy output across every genre because three rules were over-applied”: a lushness rule that required atmosphere on every track, so even trap and minimal got an automatic ghost pad and everything converged on the same lush electronic character; a variation rule that hard-required at least four layers, which left sparse genres unable to breathe and gave dense ones no room to exceed the floor; and no complexity axis at all, so “complex deep house” received the same four-layer template as “deep house”. The fix splits the lushness rule into tiers — atmosphere required for ambient, dub, lo-fi and jazz, dry by default for trap, minimal, IDM, breakcore and chiptune — and adds the missing axis. It is a useful document because the symptom (music that all sounds alike) and the cause (three rules fighting each other) are usually much harder to connect than this.

  6. 06

    The rave tool grew a classroom mode

    Among the documentation is a 9,200-character guide in Chinese for using VibeRave in primary and secondary school music lessons, complete with ten-minute and twenty-minute lesson plans, a pre-class checklist, and prompt suggestions for nursery rhymes, marching bands and animal carnivals. Two details make it more than a translation exercise. The application has a family mode that suppresses sounds considered unsuitable for children — the guide names 808 sub-bass, FM distortion and dark filters as things it blocks when certain keywords appear. And it recommends running whisper and Ollama locally for the lesson, not for privacy but so a classroom demonstration does not depend on the school’s network. The default rave-styled background image is also configurable, with the guide suggesting a cartoon instrument or a stave instead.

  7. 07

    What the numbers look like

    Fifty-nine commits, forty-three of them in the first five days after the initial open-source release on 26 April 2026, and the newest on `main` is 16 July 2026. Five hundred and forty-eight files, twenty-four megabytes, mostly the inherited Strudel codebase; the README alone is fifty-five thousand characters, with a fifty-two-thousand-character Chinese translation beside it. No releases and no tags. Ten stars and one fork. Both pull requests were opened by the author and merged by the author, and neither drew a comment. There is no live deployment: it is a self-hosted application you install and point at your own model. One loose end is visible in the repository’s own event feed, which records a create on 4 September, three pushes through 7 September and a delete on 10 September, while the branch list still holds only `main` and there are no tags — so the most recent work on this repository is work that did not land.

Adjacent records

All records →