SWE-agent
A research project that argued the model was never the bottleneck — the interface was.

What it is
Takes a GitHub issue and tries to fix it. Its actual contribution is a concept rather than a product: the Agent-Computer Interface, the argument that models fail at software engineering partly because the tools they are given were designed for humans. The team redesigned the commands and feedback formats around the model, and reported resolving 12.29% of issues on SWE-bench — state of the art at the time.
Who built itSWE-agent comes out of a Princeton research group rather than a product team: the NeurIPS 2024 paper lists seven authors — John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik R. Narasimhan and Ofir Press — and the repository’s first commit is by Jimenez. Six of the seven also wrote SWE-bench, the benchmark the 12.29% figure is measured on, which the same group published in October 2023 (arXiv:2310.06770). The repository has since taken 101 contributors, and the group’s follow-up EnIGMA extends the same interface idea to capture-the-flag security tasks with researchers from Tel Aviv University and New York University.
Build log
3 stages- 01
Reframing the problem as an interface problem
The research question was not “which model is strongest” but “what should the model be looking at”. The team designed purpose-built commands and output formats so a language model could browse a repository, view, edit and execute code without wading through terminal output meant for a person. They named this the Agent-Computer Interface.
- 02
A number, published with a date on it
On the SWE-bench test set the agent resolved 12.29% of issues, state of the art in April 2024. That figure is two years old at the time of writing and should always be quoted with the date attached — a benchmark score without a date reads as a current claim.
- 03
The researchers now recommend their own successor
The documentation for SWE-agent carries a banner at the top recommending mini-swe-agent instead, describing it as the same performance with far more simplicity and flexibility. The original is still maintained, but the people who wrote it point new users elsewhere — worth stating plainly rather than filing the project as simply “active”.
What they would tell you
- When progress stalls, the interface between the model and the work is a legitimate place to look — not just the model.
- A benchmark result is a claim with an expiry date, and should be published with one.
- A research project that recommends its own replacement is doing its job, not failing.
Adjacent records
All records →No. 022
goose
A general-purpose AI agent that runs on your own machine, shipped as a desktop app, a CLI and an API, with extensions built on the Model Context Protocol.
No. 020
Browser Use
Gives a model a browser instead of a screenshot: the page is converted to text the model can act on.
No. 018
Cline
An open-source coding agent that reads files, runs terminal commands and edits across a repository.