Vibe Projects
One person’s monorepo of six AI-generated projects — an Android agent harness, a per-hunk diff reviewer for VS Code, and four small Android apps — kept deliberately as a running measure of what the models could build.

What it is
A single repository holding six AI-generated projects and a record of which model built each: a native Android agent harness that drives your own phone using your own logged-in accounts, a VS Code extension that reviews an agent’s edits one hunk at a time, and four small Android apps — a screen dimmer that goes below the hardware minimum, a scroll-distance tracker, a guitar tuner, and a Flutter rebuild of an offline strength-training logger. The README calls the collection a personal benchmark of what large language models can currently do, and its disclaimer is unusually direct: the code is heavily generated, only lightly reviewed, meant for prototyping rather than production, and used at your own risk.
Who built itCommits as steveimm, with 99 public repositories and a profile created in 2018 that lists a personal site at steveimm.id and leaves company, location and bio empty. All 103 commits here are his, under a personal Gmail and an address at the lunit.io domain. The repository has been running since January 2026 and is still moving.
Build log
9 stages- 01
A benchmark that is also a pile of apps
Vibe Projects opened on 2026-01-25 with an initial commit and, the same day, a release tagged 1.0.0 carrying APKs. The README says what it is in two sentences: somewhere to build random ideas with AI, and — the more interesting claim — “my personal benchmarks of the current capability of LLM”. Its project table carries a column this archive has no equivalent for anywhere else: the model that built each entry. extradim and scrollmuch are credited to Claude Opus 4.5 and 4.8; cherry-diff to Opus 4.8, Fable 5 and GPT 5.6 Sol; Fretmate to GPT-6. The table is behind the repository. It lists four projects and the tree holds six, and the two it omits are the largest and the newest: PocketPilot, which is 913 of the 1,195 files and 19.2 of the 37 megabytes, and fitnotes++, which has no row at all. The remaining numbers are small and unembarrassed: 0 stars, 0 forks, 0 watchers, 103 commits from one author across eight months, and no commits at all in February, June or July.
- 02
The models are named in the commit trailers
Git is a reasonable place to look for who did the work, and here it names the models. Nineteen commits carry co-author trailers and every one of them is a model: Claude Fable 5 eight times, Claude Opus 4.8 (1M context) eleven. The repository also keeps its agent configuration as files, per project rather than globally. PocketPilot has a 2,880-byte
CLAUDE.mdwithAGENTS.mdas a nine-byte link pointing at it — the reverse of the arrangement DeepSeek Harness uses — plus seven.claude/skills/covering action debugging, an autotune loop, build fixes, prompt tuning and visual UX debugging, and an architect subagent. cherry-diff and scrollmuch each carry anAGENTS.mdof their own, at 12.6 and 8.8 kilobytes. Three of the six projects therefore ship the instructions they were built under, which is the closest thing this repository has to a stated method. - 03
PocketPilot is 95% of the repository
The largest project is an open-source agent harness for Android, and the README argues for it against the alternatives: most phone-use agents either need a computer tethered over ADB to drive the handset, or run inside a cloud virtual phone that does not have your accounts logged in. PocketPilot runs on the phone you are holding, against the apps you are already signed into, with no laptop and no re-logging-in. The architecture is a ReAct loop with no external orchestrator. The primitive toolset is deliberately small —
mobile_actionfor tap, type and swipe,open_appandsystem_buttonfor navigation,scratchpadand a task-list tool as in-session working memory — withdelegate_taskfor subagents and a preliminary long-term memory held as markdown at user, device and per-app scope. Three tools are escapes from tap-and-swipe:shellfor Android toybox commands one at a time with no pipes or redirects,termux_shellfor a full Linux toolchain on the device, andbrowser_scriptfor JavaScript against real Chrome over the DevTools protocol, where loops and retries happen inside one tool call. Perception is the accessibility tree by default with optional screenshots, and a virtual display through Shizuku lets the agent work on a parallel screen while the foreground stays yours. - 04
Two kinds of skill, and the second is the interesting one
PocketPilot’s skills split in two. Agent skills follow the agentskills.io format and load progressively on demand; today they ship bundled with the app, and a discovery engine is described as in progress. App skills are the author’s own design: a
SKILL.mdper package teaching the agent how to operate one specific app, loaded automatically whenever that app comes to the foreground. Seventeen of them are in the tree — Chrome, Photos, Calendar, Files, Settings, VLC, OsmAnd, Markor, Tasks, OpenTracks, Retro Music, Broccoli, Audio Recorder, Expense, Simple Calendar and Simple Gallery among them. It is a plain idea with a real consequence: the general agent stays general, and what is known about one app lives in a file beside it rather than in the system prompt. - 05
An evaluation loop that edits the agent
The repository contains an
eval/directory, and the README says what it is for: run an AndroidWorld task suite against the agent, let an autotune harness analyse the failures, have it propose prompt, tool and skill fixes, then run again. The parts of that loop are all present as directories —eval/tests,eval/results,eval/analysisandeval/tools, with aninspection_tool/that has a replay mode and tests of its own — and.claude/skills/carries anautotuneskill beside anautotune-loop, both with reference material and scripts attached. The same README states the safety posture, and for a hobby project it is unusually absolute: banking, authenticator and crypto-wallet apps are hard-blocked and no setting can override that; unfamiliar apps prompt per app for always-allow, session-only or deny; screens markedFLAG_SECUREare invisible to perception by design; and there is no telemetry or third-party analytics, with traces kept on the device. - 06
The other five
cherry-diff is the one that is not an Android app: a VS Code extension that snapshots tracked files when you press start, then turns every later edit into a per-hunk review to accept or reject. It is agent-agnostic by construction, watching editor and filesystem events so that nothing has to integrate with it, and it does not need Git, because rejection restores from its own baseline snapshots. Two details read as having been used in anger: a ten-second poll catches creations and deletions the VS Code watcher misses, including paths suppressed by
files.watcherExclude, and rejection is byte-exact, so binary and non-UTF-8 files survive the round trip and a UTF-8 BOM is preserved. The Android apps are small and specific. extradim draws a software dimming overlay above the screen so brightness can go below the hardware minimum, run from a foreground service with a quick-settings tile and touch passthrough. fretmate is a guitar tuner and metronome for Android 15 and later, with 33 tests passing, a clean analyzer run and a passing format check — and an Android build the README states has never been verified, because no SDK was installed. fitnotes++ rebuilds FitNotes, an offline strength-training logger with no account, in Flutter, against a reverse-engineered reference document covering the schema, exercise types, graph metrics, the Brzycki one-rep-max formula and the CSV format; milestones M0 to M4 are done, and its two custom features are planned with their schema seams already in place so that they arrive without a migration. - 07
The scroll tracker, or what to do when the API lies
scrollmuch counts how far you have scrolled, across every app, in metres, with a per-app breakdown. The interesting part is the measurement. Android exposes
scrollDeltaXandscrollDeltaYon scroll events, and the README says plainly that most apps report either zero or a direction-only plus-or-minus one rather than a real pixel distance. So the distance is derived three ways in order of trust: use the reported delta when its magnitude exceeds one; otherwise difference the app’s absolute scroll position against its last value and accept the result only when the scroller reports pixel-scale positions and the change is not a teleport or a switch between scrollers; and when an app offers no usable pixel signal at all, fall back to a fixed estimate of roughly five centimetres per gesture. It is a small app whose central problem is that a platform API does not do what its name says, and the answer is written down in the README rather than hidden in the code. - 08
What the disclaimer says, and what is in the repository
The root README does not oversell anything. It says the projects are heavily generated by AI, that the author performed very limited review of the code, that they exist to realise ideas quickly as prototypes, that they are not for production, and that use is at your own risk. That paragraph and the repository are worth reading against each other without resolving them. The root licence is MIT; PocketPilot’s is Apache 2.0. PocketPilot — the project the table does not mention — carries a Play Store release folder with a privacy policy and store screenshots, a 21.8-kilobyte architecture document, a state-machine document, an evaluation harness with recorded results, seventeen per-app skill files and a replay inspection tool. None of that contradicts a disclaimer about code review; a repository can be lightly reviewed and carefully packaged at the same time. It is worth noticing only which of the two the disclaimer is about.
- 09
The unglamorous half
Nothing outside the repository has ever been written about it. Zero stars, zero forks and zero watchers, and the only two events in its history that did not come from its author are pull requests from ImgBot — one merged the day the repository opened, one still open since 2026-08-24, both reporting that images had been optimised. The single release, 1.0.0 with APKs, went out on the first day and has not been repeated. Several things are explicitly unfinished and labelled as such: long-term memory and the skill system are marked preliminary in PocketPilot, fitnotes++ is waiting on two planned milestones, fretmate has never been built for Android at all. And the model attribution that gives the repository its point is self-reported: the commit trailers are its only independent trace, and they name two of the models the README credits.
Adjacent records
All records →No. 046
ORCH
A runtime for running several coding agents on one project at once: you define a team, give it a goal, and a CTO agent decomposes the work while the rest pick up tasks — across Claude, Codex, Cursor, Grok and a plain shell — with all the state kept in files rather than a database.
No. 027
CCManager
A terminal menu that keeps several coding agents running at once, one per git worktree: it shows which are busy, which are waiting on you and which are idle, creates and merges the worktrees, and can bring the whole set back after a crash.
No. 022
goose
A general-purpose AI agent that runs on your own machine, shipped as a desktop app, a CLI and an API, with extensions built on the Model Context Protocol.