BettaFish
A public-opinion analysis system in which several kinds of research agent are made to argue with each other on purpose, so the output is not one model’s opinion written up at length.

What it is
A from-scratch multi-agent system that reads public opinion and writes a report about it. You describe what you want to know and a crawler cluster goes out across thirty-odd social platforms — comment sections included — while several specialised agents work the material from different angles: a query engine, a media engine that handles video and images, an insight engine that mines a private database you supply, a forum host that moderates debate between the agents themselves, and a report engine that assembles the result. It is self-hosted, runs in Docker, accepts any OpenAI-compatible model, and has been public since July 2024.
Who built itA student who lists USTC and BUPT as his institutions, and whose later project is also in this archive. The two repositories are worth comparing: this one has been public since July 2024, carries a contributor list of forty-two names, has six tagged releases, and its README shows three sponsor logos with referral links beside the title. He also handles incoming contributions through an automated agent — every maintainer reply in the repository is prefixed “Agent automated reply on behalf of the maintainer”, and the agent reproduces pull requests, runs the test suite and reports findings in detail.
Build log
7 stages- 01
What it is, and why it is named after a fish
The project is called 微舆, which is a homophone of 微鱼, “tiny fish”, and the README explains the choice: a betta is a small, beautiful and notably aggressive fish, standing for “small but strong, unafraid of challenges”. What it does is take a question in ordinary language and return a research report built from public opinion. A crawler cluster works across more than thirty domestic and international platforms — Weibo, Xiaohongshu, Douyin, Kuaishou among them — and goes past the posts into the comment threads, which is where most opinion actually lives. The README claims six advantages over comparable tools; the ones that are architectural rather than promotional are a composite analysis layer that mixes fine-tuned and statistical models in with the language model rather than delegating everything to it, a multimodal path that parses short video and extracts structured cards such as weather, calendars and stock quotes, and a mechanism for joining public opinion data to a private business database so a company can read outside sentiment against inside numbers. The worked example shipped with the repository is a brand-reputation analysis of Wuhan University, and a video of the complete run is on Bilibili.
- 02
The forum, which is the part worth copying
The distinctive mechanism is what the README calls an agent forum. The project does not run one agent with one prompt and ask for a report; it gives different agents different toolsets and different habits of thought, and it introduces a separate moderator model whose job is to run a structured debate between them. The stated purpose is to prevent exactly the failure that plagues single-model output: agents that share a model and a context converge on the same reading, and a report assembled from four copies of one opinion is no better than a report assembled from one. Putting a moderator in the middle, whose role is to surface disagreement and force it to be resolved rather than smoothed over, is a cheap structural fix for a problem that is otherwise usually addressed by prompt wording. The successor project in this archive takes the same instinct much further — it stops trying to reach a conclusion at all and simulates a population instead — and the README here already gestures at it, describing a goal of becoming a general data-analysis engine and noting that changing the agents’ tools and prompts is enough to turn it into a financial market analyser.
- 03
A year at one star, then one week in November
The repository is unusual in this archive in that it publishes its own growth curve, and the curve is the story. The first recorded day, 5 July 2024, is one star. It reaches 272 stars by the end of August 2025, fourteen months later — a long, nearly flat line for a project that was public the whole time. Then October: 674 at the start of the month, 2,082 by the 31st. Then 1 November at 2,994, and the following week is vertical — 6,097 on the 3rd, 10,802 on the 4th, a single-day gain of 4,705, and 21,983 by the 8th. Roughly nineteen thousand stars in seven days. It passes 29,000 by the end of November, and from there the curve flattens into something ordinary: 32,941 at the end of 2025, 39,556 by the end of March, 41,501 by the end of June, 42,308 now. Press coverage of the author describes this as twenty thousand stars in a week after the project went public, which is right about the week and misleading about the timing — the release that started it came sixteen months after the repository opened.
- 04
Six releases, nine tags, and a version that is not the version
There are six published releases: v1.0.0 on 1 September 2025, v1.1.0 on 4 November — the day of the largest single-day star gain — v1.2.0 four days later, v2.0.0 on 28 November, v2.1.0 on 9 December and v3.0.0 on 23 December. Nine tags exist for the six releases. The repository also keeps a pinned issue, opened in July 2026 and written in Chinese, that is a frequently-asked-questions page, and its first answer is a warning about the word “latest”: the newest release is still v3.0.0, the main branch is forty-four commits ahead of it, and the two should not be treated as the same thing. Anyone reporting a problem is asked to include the output of `git rev-parse HEAD` rather than writing “latest version”, because there are now two plausible meanings of it. The same page records that a GraphRAG module was merged and later reverted, and that the current main branch does not contain it — with links to the merge, the revert and the follow-up pull request.
- 05
The maintainer answers pull requests with an agent
The most distinctive thing about this repository is not in its code. Nearly every reply from the maintainer on a pull request or issue begins “Agent automated reply on behalf of the maintainer”, and the replies are not acknowledgements — they are reviews. The agent checks out the exact commit a contributor proposed, compares it against the merge base, runs the repository test suite, and reports what it found, with findings graded and specific. It declines work as well as accepting it: on a pull request that added a routing gateway and claimed all seven of the project’s engines would go through it, the agent replied that the core claim was not met, because the report engine imports its own settings object rather than the one the patch modifies. On a pull request reporting a critical SQL injection, the agent replied that it could not establish the source-to-sink path, because the values in question come from hard-coded internal mappings rather than user input, and that the proposed fix would not run against the pinned version of the database library. On a third, a pickle-deserialisation fix, it accepted the direction and pointed out that the unsafe loader was still being called first, so the vulnerability was not actually removed — and the contributor agreed. A maintainer using a model to triage contributions is common enough; publishing the review text, unedited and attributed, under a 42,000-star project is not.
- 06
What a project with forty-two contributors attracts
A thousand and thirty-nine commits from forty-two accounts. The author has 402 of them; the next three have 301, 63 and 53, so there is a second substantial maintainer rather than a crowd of drive-by patches. A hundred and ninety-nine commits come from identities with no linked account, meaning contributors who set a git name without matching it to GitHub. What arrives through the pull request queue is largely not feature work: a PDF export path-traversal fix, where a report topic such as `../../tmp/pwned` was being interpolated into both a filename and a `Content-Disposition` header; an SBOM and provenance attestation for the published Docker image, whose author then corrected his own description in public after discovering that provenance attestations are already added by default for public repositories; container registry, star-chart maintenance and sponsor-link repairs. The security reports are not all accepted and one of them is not a real vulnerability, which the agent says plainly rather than merging to be agreeable. A project at this size spends most of its contribution bandwidth on the things that break once other people run it.
- 07
What it left behind
Two things connect this repository to the one that followed it. The first is concrete: the seed document for the public-opinion simulation demonstrated in the successor project is a report produced here, so the same institution — a university — is analysed by this system and then simulated by the next one. The second is a difference in how the two repositories record their own making. This one has a thousand and thirty-nine commits, of which nine carry a co-author trailer and three name Claude; the successor has fifty-nine commits, of which twenty-six carry a Claude co-author trailer. The same person, disclosing the same kind of help, went from doing it in under one per cent of commits to doing it in nearly half. Nothing in either repository explains the change, and it is only visible because commit trailers are metadata rather than prose: nobody had to decide to write it down.
Adjacent records
All records →No. 031
MiroFish
A prediction engine that builds a parallel world of hundreds of AI agents out of a document you upload, then runs it forward to see what happens next.
No. 035
autoresearch
Give an AI agent an LLM training script, a fixed five-minute budget and no supervision — then see what it found while you slept.
No. 034
N.O.R.A.Core
A single-author AI companion with two brains, an encrypted diary of its own, and 98 commits signed in its own name.