autoresearch
Give an AI agent an LLM training script, a fixed five-minute budget and no supervision — then see what it found while you slept.

What it is
A small, complete LLM training setup designed to be handed to an autonomous agent overnight. The agent edits one file, trains for exactly five minutes, checks whether validation loss improved, keeps the change or reverts it, and repeats — roughly a hundred experiments while a person sleeps. The inversion is the interesting part: the human no longer writes the Python, the human writes the Markdown that instructs the agent.
Who built itA Slovak-Canadian AI researcher who co-founded OpenAI and later led artificial intelligence and Autopilot Vision at Tesla; he founded Eureka Labs, an AI education company, in 2024, and joined Anthropic’s pretraining team in 2026. The training code here is a single-GPU simplification of his own nanochat, and the three-file design is the point of the experiment rather than a shortcut.
Build log
6 stages- 01
The framing is part of the artefact
The README opens with a short piece of science fiction, signed and dated: that frontier research used to be done by “meat computers in between eating, sleeping, having other fun, and synchronizing once in a while using sound wave interconnect in the ritual of group meeting”, that this era is long gone, and that this repository is the story of how it all began. It is an unusual way to open a machine-learning repo, and it sets the register for everything after it.
- 02
The inversion: the human edits the Markdown
There are only three files that matter, and they are split by who is allowed to touch them. prepare.py holds the constants, one-time data preparation and runtime utilities, and is not modified. train.py holds the model, the optimiser and the training loop, and is the file the agent edits. program.md holds the instructions for the agent, and is the file the human edits. As the README puts it, you are not programming the Python — you are programming the Markdown that sets up your autonomous research org.
- 03
The loop, spelled out for the agent
program.md describes the experiment concretely enough to be mechanical. Work on a dedicated branch; hack train.py with an idea; commit; run the training with output redirected to a log; grep the validation bits-per-byte out of that log; record the result in a tab-separated file with a status of keep, discard or crash; then advance the branch if the number improved or reset it if it did not. Five minutes per experiment, roughly twelve an hour, about a hundred across a night’s sleep.
- 04
“NEVER STOP”
The instructions end with a section that is the most quotable thing in the repository, because it is where the design admits what it is asking for. Once the loop has begun, the agent must not pause to ask whether it should continue — no “should I keep going?”, no “is this a good stopping point?”. The human might be asleep. If it runs out of ideas it is told to think harder, read papers referenced in the code, re-read the files for new angles, try more radical changes. The loop runs until the human interrupts it, period.
- 05
A seed rather than a product
The repository is about ten files and 530 KB, and it was never going to be the thing people ran. Instead it became a template: 13,526 forks, with the README itself listing ports to macOS, Windows and AMD hardware contributed by other people, and a section of advice on shrinking the defaults for machines smaller than the H100 it was tested on. The discussion around it is larger than the code — Hacker News threads through 2026 apply the same loop to kernel optimisation, SAT solvers and reinforcement learning, and a scaling experiment drawing 237 points asks what happens when the agent is handed a whole GPU cluster.
- 06
Where it stands
Thirty-six commits between 6 and 26 March 2026, twenty-eight of them by Karpathy, the rest one each from eight outside contributors; three carry a co-author trailer for Claude. The last commit was on 26 March 2026, and it has been quiet since, which for a repository that is explicitly a starting point rather than a maintained tool is closer to the intended outcome than a failure. Ninety-six thousand stars, 13,526 forks and 195 open issues. The README states an MIT licence, though the repository contains no licence file.
Adjacent records
All records →No. 031
MiroFish
A prediction engine that builds a parallel world of hundreds of AI agents out of a document you upload, then runs it forward to see what happens next.
No. 029
realworld-vibe-coded
A full-stack RealWorld implementation whose real subject is the harness around it: thirty-two custom Roslyn analyzers, one build entry point, and hooks that stop an agent from doing the wrong thing.
No. 017
BettaFish
A public-opinion analysis system in which several kinds of research agent are made to argue with each other on purpose, so the output is not one model’s opinion written up at length.