Skip to content

realworld-vibe-coded

A full-stack RealWorld implementation whose real subject is the harness around it: thirty-two custom Roslyn analyzers, one build entry point, and hooks that stop an agent from doing the wrong thing.

Screenshot of realworld-vibe-coded
Editor screenshot, 28 Sep 2026realworld-vibe-coded ↗

What it is

A complete implementation of the RealWorld specification — the standard Medium-clone exercise used to compare stacks — with a .NET 10 backend on FastEndpoints, MediatR and EF Core, and a React 19 frontend on Carbon Design. What makes it worth reading is not the application. It is the machinery built so that coding agents can work on it without supervision: thirty-two custom Roslyn analyzers that turn architectural rules into compiler errors, a four-layer test suite that gates every change, and a single Nuke build entry point that the agent instructions forbid it from going around.

Who built itBased in Auckland, New Zealand, with thirty-three public repositories since 2015 and almost no audience for any of them. This is his third implementation of the same specification — an earlier one is a .NET modular monolith — and the README says the reason for building it a third time was to learn an agent-first engineering workflow rather than to produce another app. One of his other repositories is a shell project named after the Ralph technique of running a fresh agent in a loop.

Build log

8 stages
  1. 01

    The application is the pretext

    RealWorld is a written specification for a Medium clone, published so that implementations in different stacks can be compared against one set of tests. Building one is a well-trodden exercise, and this repository is honest about that: it is the author’s third, and the README says the point was the workflow rather than the app. The stack is current and conventional — .NET 10 with FastEndpoints, MediatR for commands and queries, FluentValidation, EF Core against SQL Server, Serilog for structured logs and Audit.NET for change tracking; React 19 with Vite, TypeScript and IBM’s Carbon design system on the front — and the layering follows a published Clean Architecture template, with separate projects for domain models, shared abstractions, use cases, infrastructure, web endpoints and development-only endpoints. Endpoints stay thin: bind, authorise, delegate to a handler, map the response. There is a seventh project that is not part of the application at all, and it is the one that matters: an analyzers project holding thirty-two custom Roslyn rules.

  2. 02

    Even the commit log is mostly machine

    There are 1,935 commits across seven months, and the authorship is the most extreme in this archive. GitHub’s own Copilot coding agent, committing as `copilot-swe-agent[bot]`, authored 1,093 of them — fifty-six per cent. Another 402 carry the name and address of an identity called “Solo Yolo” with no GitHub account attached, which appears to be a second agent identity rather than a person. The repository owner accounts for 440. Nine hundred and sixty-seven commits carry a co-author trailer, and 839 of those name the owner, which is the same inversion seen elsewhere in the archive: the machine writes the commit and credits the human. The monthly distribution shows the shape of the work rather than a steady grind — 254 commits in the first nine days, then 498, then 156, then 711 in December 2025, then slowing to 119 in January and 149 across April before stopping. The README describes the sequence directly: Copilot for most of it, some hand engineering to make the architecture suit an agent, and Claude Code taking over towards the end.

  3. 03

    The invariants are numbered, and the first one is about context

    The project instructions begin with a list of rules rather than a description. The first is that `dotnet` must never be run directly; everything goes through a single build script target, and the agent is passed a flag whose stated purpose is to suppress verbose container output **for context efficiency**. The list continues in the same register: every feature must have its API and end-to-end tests passing before the next one starts; all compiler warnings and errors must be resolved and never suppressed; backend changes must be made first and the front-end API client regenerated from them, so the front end can never reference a field the generated types do not have. Two of the rules are about honesty rather than technique. One says every test failure must be investigated and that dismissing a failure as pre-existing, unrelated or flaky is not allowed — if it is genuinely out of scope it has to be filed with evidence. Another says the application URL must never be guessed from a port listing, because worktrees use port offsets and the dev server runs over TLS, and the build output is the only reliable source. A further rule requires front-end changes to be verified visually by driving a browser and taking screenshots before the work is considered done.

  4. 04

    Twenty rule files, loaded by path

    The instructions are deliberately thin and point at a rules directory where twenty markdown files sit, each with a stated scope. The scoping is the design: a rule file is loaded when the agent is working in the paths it covers, so a change to an end-to-end test brings in the Playwright conventions and a change to the server brings in the backend patterns. Six of the twenty are templates rather than rules — copy-paste skeletons for an endpoint with its request and response types, for a command handler, for a query handler with its validator, and for the persistence configuration — so that the shape of new code is decided by the repository rather than invented per session. Others cover feature flags and the analyzers that police them, the conventions for the generated API client, the rule-writing harness itself, and two research files: one describing a documentation server to consult before planning, and one recording the caveats of a Roslyn-based semantic-analysis server, including the blunt warning that source-generated symbols are invisible to it and its diagnostics are unreliable. A rule file whose content is “this tool lies to you about these things” is a good sign that someone actually used it.

  5. 05

    The hooks enforce what the prose asks for

    Written instructions are suggestions, so the repository also carries eight shell hooks that run at defined points in an agent session, and their names are a list of everything that had gone wrong. One enforces that the build entry point is used rather than invoking the .NET CLI. One enforces that the context-suppressing flag is present. One restricts which tools may be called. One protects files that must not be edited. One writes an action log. One runs when a test target fails. One installs the git hooks at session start, and one fires at the end of a session. Between the numbered invariants, the path-scoped rules and these hooks there are three layers saying the same things, which is the point: each layer catches what the one above it only asked for.

  6. 06

    Analyzers are how a lesson becomes permanent

    The thirty-two custom compiler analyzers are the most interesting part of the design, and the pull requests show how they get written. One example requires that every endpoint type has a matching validator, with a single documented exemption, so an endpoint cannot ship without input validation. A companion rule requires validators for paginated requests to inherit a shared base class, so the limit and offset rules live in one place instead of being rewritten per endpoint. A third exists specifically to stop a pattern reappearing: one pull request installed a global validation name resolver, deleted thirty-four manual override calls across seventeen validators, and then added an error-level analyzer so that anyone writing the override again gets a compile failure. Another request added two more analyzers alongside a pagination abstraction, and noted that changing paginated JSON from per-resource key names to generic ones was a breaking change requiring the generated client, the front end, the API test collections and the end-to-end tests to move together. Guardrails are added at the moment a rule is learned, which is why there are thirty-two.

  7. 07

    Unattended by configuration, and said so plainly

    The repository is set up to run coding agents without asking permission, and it does not hide it. A configuration file for one agent tool is commented as YOLO mode, explaining that it bypasses all approvals and disables the sandbox entirely, that the agent can read and write anywhere, run any command and use the network without asking, and that this is equivalent to the skip-permissions flag the other tool is run with. The instructions note that the approvals file is only loaded when the project is marked as trusted in the user’s own configuration, which is the one remaining gate. A separate file, added for a third tool, contains nothing but a pointer to the main instructions, with the reasoning recorded in the pull request: the main file is the single source of truth, duplication is kept at zero, and the pointer is a plain file rather than a symlink so it works on every platform. Three agent ecosystems, one set of rules, and an explicit decision that none of them should stop to ask.

  8. 08

    What every pull request gets, and what the numbers are

    Each pull request receives a series of automated comments reporting a four-layer test suite. Unit tests on both sides — the client suite runs 239 tests at seventy-four per cent coverage, the server suite 324 at eighty-five per cent, with eighty-five per cent a hard floor that fails the build. Five separate API collections run against the live application through Postman, split into auth, profiles, articles, feeds and an empty-state suite, totalling several hundred assertions and reported per suite with a pass rate. An end-to-end layer of 84 Playwright tests runs the real browser against the real app. The repository is 950 files, has no releases and no tags, is MIT licensed, and has twelve open issues. It has no stars and no forks, which for a project of this size and this much automation is the sharpest number in the record: a great deal of engineering, of a kind that is currently being discussed everywhere, and almost nobody has looked at it.

Adjacent records

All records →