MiroFish
A prediction engine that builds a parallel world of hundreds of AI agents out of a document you upload, then runs it forward to see what happens next.

What it is
A simulation engine that turns documents into a world you can run forward. Upload seed material — a news dossier, a policy draft, the first eighty chapters of a novel — and it extracts entities and relationships into a graph, generates a persona for each of them, and starts hundreds of agents posting and replying to each other across two simulated social platforms. What comes back is a prediction report plus the world itself, which you can question afterwards. The agents carry long-term memory, and the whole thing runs on your own machine against whatever model you point it at.
Who built itA student, listing USTC and BUPT as his institutions and Shanghai as his location. His nine public repositories include two with more than forty thousand stars each, both in Chinese and both about the same subject: reading public opinion and projecting it forward. Press coverage reports that Chen Tianqiao, the founder of Shanda Group, invited him to join the company and that Shanda then invested thirty million renminbi in this project — and the repository carries Shanda’s logo in its README.
Build log
8 stages- 01
What it does, and what “simulation” means here
The pipeline has five stages. Seed material is turned into a temporal knowledge graph, with entities and the relationships between them extracted and stored in a way that also records when things happened. Each entity then gets a persona — background, memories, behaviour patterns, social position. Those personas become agents in a two-platform social simulation, one modelled on X and one on Reddit, and the simulation runs for a configurable number of rounds with agents posting, replying and quoting one another. A report agent then works over the finished world to produce a prediction report, and the same environment stays open for questions: you can ask any agent what it did and why, or ask the report agent about the run. Two details are doing more work than they look like. Every agent has long-term memory attached through Zep, so what it saw in round four affects what it says in round twenty. And the simulation is seeded by whatever document you supply, which is why the same engine handles public opinion about a university and the lost ending of an eighteenth-century novel.
- 02
The novel run, in numbers
The demonstration the project is best known for takes the first eighty chapters of “Dream of the Red Chamber” — about 150,000 characters — and asks what the missing ending would have been. The graph stage produced 905 entities and 3,822 relationships; the persona stage produced 580 of them, so 580 agents, each with a written biography and, in at least one case, a Myers-Briggs type. Thirty rounds of the two-platform simulation produced close to two thousand actions. The report agent’s conclusion was quoted in the press coverage and reads better than most machine-written prose: that the collapse of the garden was not an accidental tragedy but the necessary dissolution of a ritual order resonating with individual fates. Some of what it predicted matches the surviving ending — Daiyu burning her manuscripts, Xianglian taking vows — and some does not: where the accepted continuation has Baoyu sit the imperial examination, this simulation has him break down and leave with the mad Taoist. The author also published what the run cost: about fourteen renminbi. His stated limitation is worth recording too — when the input text is very large, the output can drift into mixing Chinese and English.
- 03
Ten days, and what the commit history shows
Press coverage of this project leads with a number: ten days of vibe coding to build it. The repository’s own history gives a slightly different shape without contradicting the claim about the build. The first commit is on 26 November 2025 and the repository is public from that day. Across the following nine months there are 320 commits on 64 distinct days, spanning 281 calendar days. Twenty-one of those commits fall inside the first ten calendar days; the tenth day on which anything was committed is 11 December 2025; and the first tagged release, v0.1.0, lands on 22 December, twenty-six days after the first commit. The ten-day figure describes how long the building took, not how long the repository has existed, and the two numbers are not in tension — but only one of them is in the press, and the other is in the commit log.
- 04
The star curve, from the repository’s own data
The repository ships a subsystem that tracks its own star count and commits the chart, which means the growth curve is published rather than estimated. It starts on 28 November 2025 at one star. On 22 December, the day v0.1.0 is tagged, it is 40. The next day it is 270 — a multiple of nearly seven in twenty-four hours. It reaches 683 by new year, 3,234 by the end of January, and 4,042 by the end of February. Then March happens: 10,278 by 9 March, and 46,118 by the end of the month, an elevenfold rise in thirty-one days. From there the curve flattens into something more ordinary — 57,873 in April, 63,036 in May, 67,430 in June, 68,993 in mid-July, 71,809 by the start of September. The API reports 75,157 now. What the shape says is that the project was public and nearly unread for a month, took off the day it was released properly, and had one extraordinary month five months later.
- 05
The investment, and the logo in the README
The README carries the logo of Shanda Group next to the project title, which is an unusual thing for an open-source project and is explained by the coverage. Chen Tianqiao, Shanda’s founder, invited the author to join the company after an earlier project of his went viral, and the author built the prediction feature he had wanted since that earlier project in ten days at Shanda. The account in the Chinese press is that the investment decision was made within twenty-four hours of the demonstration video being submitted: thirty million renminbi, roughly four million US dollars, for the project to be developed further inside the group. The repository also carries a Trendshift badge for topping GitHub’s trending list, a Discord server, and accounts on X and Instagram. This is an open-source project with a marketing operation behind it, and the README does not pretend otherwise. The split between the two halves is worth stating plainly, because it is easy to miss: the repository is AGPL-3.0 and self-hosted, while the site at mirofish.us sells a hosted version by subscription — the first screen offers it “in your browser, from $2.99/month”, alongside pages for use cases, research and pricing. Same project, two ways to use it, two very different price tags.
- 06
The predecessor, and where the idea came from
The same author’s earlier project, BettaFish, has 42,308 stars and 7,615 forks and is licensed GPL-2.0 rather than this project’s AGPL-3.0. It does public-opinion analysis: you give it a topic and it searches social platforms, has a team of agents summarise and argue over what they find, and returns a report. By the author’s account it began as his graduation project, and press coverage describes it gaining twenty thousand stars in a week after it went public. The repository’s own star data fills that in and corrects the obvious reading of it: the project was public from July 2024 and sat at 272 stars as recently as August 2025, then took on roughly nineteen thousand stars in the seven days from 1 to 8 November 2025 — the week in which v1.1.0 and v1.2.0 were released. The week is real; it was a release, sixteen months after the repository opened, that started it. MiroFish is the next step in the same direction, and the two connect through the part that matters: a BettaFish report was the seed document for the university public-opinion simulation that is one of this project’s two demonstration videos. Analysis feeds simulation, which is the closed loop both READMEs describe. Both repositories are written from scratch rather than on top of an agent framework — a phrase that recurs in the author’s project descriptions — and both are Chinese-first with English alongside.
- 07
How he says he built it
The most useful source on this project is not the repository but the author’s own account of the process, quoted at length in the Chinese coverage. Most of his time went into market research and choosing the stack rather than writing code: decide why you are building it, who it is for and how it will work, and only then direct the model. The flow he describes is a Figma sketch refined with a model, a frontend demo in Google AI Studio, the pages folded back into the project documentation, and the work then split into modules and handed to a coding agent in batches. He used two models by role — Gemini 3 Pro for the frontend, which he describes as having more instinct for page structure and interaction polish, and Claude for backend structure, interfaces and stability. And he ran several agents on the same task at once and kept the best result, often eight on one module. He is direct that this burns tokens heavily, and equally direct about why he kept doing it: it is the fastest way to learn where each model’s competence actually ends.
- 08
What a forty-thousand-star project attracts
Three hundred and twenty commits from sixteen contributors, and 148 open issues. The contributor list is worth reading because of what the outside work is about. One contributor has a series of open pull requests that read as a production-readiness audit: identifier validation before filesystem paths are built (the description notes that two of those paths feed `shutil.rmtree`, so a crafted identifier reaching a delete endpoint could remove an arbitrary directory); API key authentication, which the project had none of, leaving every endpoint including destructive ones reachable anonymously; resource-level ownership checks, absent even after authentication went in; atomic writes and a state machine for two JSON files that were being written with plain `open(..., "w")` and could be left truncated by a kill; crash recovery for an in-memory task manager; and a reproducibility manifest recording which model, documents and ontology produced a given result. The everyday issue list is more mundane and more telling: the deployed site was displaying its own raw HTML source, five languages are registered in the locale file with no translation behind them, and a Spanish translation has now been submitted three times by three different people. A project at this scale generates work that has nothing to do with its idea.
What they would tell you
- Most of the time goes into market research and choosing the stack, not into writing code. Work out why you are building it, who it is for, and how it will work — then tell the model what to do, not the other way round.
- Run several agents on the same task in parallel and keep the best result. He describes often having eight on one module at once. The token bill is real and so is the speed, and it is the quickest way to find where each model’s ability actually stops.
- The faster you go, the better the brakes need to be. Version control and written documentation are what stop a change in one place from breaking another.
- Read the generated code line by line and follow the execution, not just the diff. His claim is that most bugs are not a wrong line but the model drifting on one key assumption, and that correcting the assumption makes a group of symptoms vanish together.
- Do not try to cover everything. Cut the scope, keep re-testing the positioning, and do not wait for it to be perfect before showing it to anyone.
- Marketing can be minimal, but the material that lets other people promote you has to exist in advance — above all a demonstration video clear enough that someone else can post it.
- Code is cold; stories are warm. Being able to tell the story behind the code is a required skill for anyone working alone.
Adjacent records
All records →No. 017
BettaFish
A public-opinion analysis system in which several kinds of research agent are made to argue with each other on purpose, so the output is not one model’s opinion written up at length.
No. 023
VibeFlow-P5
A p5.js particle sketch, a 356-byte Go file server and a hardened Docker image, published alongside a manifesto for vibe coding — with the conversation’s German left in the comments.
No. 022
goose
A general-purpose AI agent that runs on your own machine, shipped as a desktop app, a CLI and an API, with extensions built on the Model Context Protocol.