VibeGame
Describe a game in one sentence and a team of agents divides the work — an architect plans it, a programmer builds it, an auditor checks the code against the plan, and a player has to actually play it before the task is accepted.

What it is
An open-source framework that turns a sentence into a playable two-dimensional browser game. It is two things stacked. Underneath is an engine built on Phaser that loads a scene tree, assets and scripts from plain JSON and JavaScript, so a game is a directory of editable files rather than a binary. On top is a multi-agent harness that runs inside an ordinary coding agent — Claude Code or Codex — and divides the work among a team of named roles with a chain of documents between them. The published demos include a civilization prototype, a Hollow Knight boss slice, an arcade fighter, a dungeon roguelike, a parkour game and an AI-moderated werewolf party game.
Who built itTet lists Xi’an Jiaotong University as his affiliation; kiyotakali is the second maintainer and the address given for the commercial-deployment notification. The project ships a technical report, a Discord server and a WeChat group, and its acknowledgements place it alongside other research efforts in the same area rather than on its own — which is closer to a research release than to a personal tool.
Build log
7 stages- 01
Two layers, and the one that is the point
The repository is a game engine and an agent harness, and the README is honest that the second is the subject. The engine side is deliberately plain: Phaser 3 underneath, with a game described by a
project.jsonand a directory of.scene.jsonfiles,.node.jsontemplates, JavaScript scripts extending an engineNodeclass, and a manifest declaring which assets to preload. Anything dynamic — bullets, enemies, pickups — is spawned at runtime from a template rather than assembled from raw objects, and the instructions say so in those words. Six JSON schemas in the engine directory define nodes, scenes, tilemaps, manifests, project files and input maps, so the whole game state is something an agent can read and write without touching compiled code. The first layer exists mostly to make the second one possible: an agent needs a world made of text files with a schema, and a game engine is a good excuse to build one. - 02
A team with an adversarial step
The harness defines eight roles. An orchestrator owns product intent — deciding what gets built, splitting it into tasks, reviewing plans and accepting results. Below it sit a designer and an artist, both persistent; an architect, programmer, auditor and player, each scoped to one task; and a reviewer for the final quality gate. The task workflow is numbered and the order matters: align requirements, survey what already exists, create the task with a written feature spec covering trigger condition, initial state, flow, terminal states and edge cases, then architect, then review the plan, then implement, then audit, then play. Two of those stages exist to catch the others. The auditor compares the finished code against the plan and the requirements, and the player has to actually play the game before the task is accepted. Setting a non-authoring role as the gate is the same instinct as putting a reviewer in the loop, taken one step further: the last word belongs to something that behaves like a user.
- 03
Documents with owners, and consistency checks between them
The roles do not coordinate by chatting; they coordinate through files with declared owners, and the ownership rules are written down. A game design document holds the global gameplay truth, and a new mechanic has to land there before it can appear in a requirements document — the orchestrator is told explicitly never to invent mechanics in the requirements slice. The architect owns a plan containing the technical design, a runtime state contract and a verification plan. A log file is append-only and holds handoffs between the programmer, the auditor and the player in turn. The interesting part is who checks what: the architect verifies that the requirements are consistent with the design document before planning, and the auditor verifies that the code is consistent with the requirements and the plan after implementing. Each hand-off is a real document a person can read afterwards, so a failure can be located at a stage rather than blamed on a model.
- 04
A glossary enforced across every agent
One of the more quietly clever things in the repository is a glossary that all agents are required to use and prohibited from paraphrasing. It defines the engine’s vocabulary precisely and, in several entries, names the words that are forbidden: a unit in the scene tree is a node and must not be called an entity, an actor, a game object or a prefab, even though the last two are what a developer arriving from Unity or Godot would reach for. A reusable node definition is a template rather than a prefab. The rules add that terminology must not switch between turns, and that when a term is introduced to a non-developer the agent should give a plain-language analogy once and then keep using the canonical term. The pivot entry alone runs to a paragraph, specifying the cascade between clip, sprite and group settings and the separate cascade for colliders. Prompt drift is usually treated as a model problem; this treats it as a vocabulary problem, and fixes it by writing the dictionary down.
- 05
The harness, and what it costs to install
The orchestrator’s own instructions run to thirty-eight thousand characters, and it is supported by nine Python hooks that fire at defined points — session start, prompt submission, subagent stop, team setup, team mode enforcement, goal capture, status line, teammate messages, and a web hook. Between them they enforce the team structure rather than merely describing it. Installation is where the seams show, and the README is unusually candid about them. The framework launches agents with the flag that skips all permission prompts, and it warns that anyone who has not used that flag before must run it once in a terminal first, because the first-run consent screen cannot be answered from inside the managed session and agent startup fails without it. Codex users have to run a setup command and then separately open Codex, type a command and trust the installed hooks; until both are done, the README says, agents still run but never register, the dashboard stays empty, and startup waits forever for a session that never appears. That is a specific, unglamorous failure mode written down for the next person, which is what good setup documentation looks like.
- 06
Rules that read like scars
The engine instructions given to the agent contain a handful of rules that could only have been written after something went wrong. One forbids the pattern-matching kill command across the whole project, with the reason stated in the rule itself: it can terminate the entire agent team, the dashboard, the runtime and the terminal session at once. Another requires that files be deleted by a recoverable method rather than a permanent one. Another tells the agent to read the design document, the asset inventory and the engine specs before acting, and not to guess at game state or what assets exist. A fourth says that when showing a file path to a user it must be a markdown link to an absolute path, and specifically not the file-scheme form that an agent might reasonably emit. And a fifth is a small piece of prompt engineering worth copying: the agent is forbidden from using structured question tools and must instead print the question and its options as plain text, because a templated dialogue box interrupts the flow the rest of the harness is trying to maintain.
- 07
What the numbers look like, and one unusual licence clause
Two hundred and fifty-six stars and eighteen forks from twenty-one commits is an unusual ratio, and it is explained by what the project is: a research release with a technical report, eight recorded demonstrations and a launch post, rather than something that grew in public. Two contributors, one of whom has sixteen of the commits and the other five. Three hundred and ninety-four files, eleven megabytes, no releases, no tags, and a single issue ever opened — a reader asking whether the demos could be linked so they could play them, and whether the games were one-sentence generations or the result of several rounds of adjustment. The answer was both, in that order, and then a follow-up saying the demos did not work, which turned out to be a reader looking at a video section rather than the play area further down the page. The licence is Apache-2.0 with one addition that is worth quoting in substance: commercial deployments are asked to send a notification with the organisation, the product name and a short description. The text says plainly that this is for tracking and outreach only, requires no approval, carries no fee and modifies no rights under the licence. It is an unusual thing to attach to a permissive licence, and stating its own limits that clearly is what makes it acceptable. The roadmap wants Godot, then Unity, then three dimensions.
Adjacent records
All records →No. 061
DeepSeek Harness
DeepSeek’s agent harness, built so that the model adapter, the tool registry, the session log and the agent loop itself are plugins — swapped from a configuration file rather than a fork.
No. 048
gstack
Twenty-three specialist roles and eight power tools for Claude Code, all written as Markdown slash commands, plus the evaluation harness their author uses to decide whether any of it is working.
No. 035
Nexus Agents
A control plane that sits above coding agents rather than being one: it admits work through a single entry point, puts every real fork to a multi-agent vote, records every action in a hash-chained audit log, and requires the repository owner to ratify any change to the rules that govern it.