AntOmniEvo
A Python framework that runs an evolution loop over a directory of your files: a coding agent rewrites them from the failure trajectories of the previous candidate, parent and child are scored on the same training batch, and only an improvement that clears a threshold is re-scored on the validation set and kept.

What it is
AntOmniEvo is a Python framework that treats optimization as file editing. You supply five things — a System that runs the thing under optimization, an Evaluator whose 0–1 scoring criteria are the objective, train and validation instances, a schema for the files that may change, and the starting files — then a coding agent reads the failure trajectories, rewrites those files as a new candidate, and the framework decides whether it survives. The genome is a real directory: a mutation can be a Markdown rule, a node script or a retrieval DAG’s pipeline.json. Each iteration rolls a parent and its child out on the same batch, keeps the child only if the summed improvement clears a threshold or is perfect, re-scores it on the validation set, then prunes the pool by Pareto dominance with a top-N fallback. A candidate that misses is not discarded: a reflection chain refines it for a bounded number of rounds. The run is written to a resumable workspace, and a React and Flask visualizer renders it. No release was ever tagged: version 0.1.1 of the two packages goes to an internal index.
Who built itAnt Research is an organisation account rather than a person: 45 public repositories, 368 followers, opened on 2022-12-17, with antresearch.com on its profile. Every commit in this repository comes from one member account, jacklv111 — opened on 2022-09-20, 13 public repositories and 5 followers — which is the author name on 9 of the 21 commits, the linked account on all of them, and the only entry in the contributor list. The paper the README cites lists six authors, led by Yubin Lyu; the repository’s internal release pipeline names two approvers, lyuyubin.lyb and hanpu.mwx. Which account belongs to which name is not stated in the material.
How it is put together
The parts · 6The organising idea is that optimization should be plain file editing: the thing being optimized is represented as a directory, and the machine that decides what survives is deterministic code the user does not write. That choice puts the intelligence in one place — a coding agent that reads a failed run and edits files — and keeps everything probabilistic out of the loop: deriving the batch, accepting or retiring, the validation re-score, the pruning and the budget are all counts, sums and comparisons. Two consequences shape the code. Because analysis and editing are the only steps an LLM is trusted with, they are split into two phases with a structured artifact between them, and the reflection path changes only that artifact, never the mutation machinery. And because a mutation loop built on an unreliable spender has to be auditable, memory is the filesystem: append-only changelogs that a child inherits from its parent, run records and analyses written before any error can short-circuit them, and a startup recovery that flips interrupted candidates back to selectable. The same documents are candid about the lineage: Pareto selection is borrowed from GEPA, and the parts presented as this project’s own are running several candidates at once and keeping rejected candidates alive.
- antomnievo/optimizer/
- 43,728 bytes in
optimizer.py: the main loop and its per-scenario subclasses for text2sql, AppWorld, TerminalBench, adconfig and the RAG pipeline. It owns the slot, the batch cursor, the accept-or-retire challenge, the reflection chain, the validation step, and the two counters a subclass cannot bypass. - antomnievo/proposer/
- The coding-agent side:
base_proposer.py(28,967 bytes) holding the two-phase analyze-then-mutate pipeline, with one thin subclass per backend —claude_code_proposer.py(6,054) drivingclaude,pi_coding_agent_proposer.py(5,776) drivingpi— three prompt templates led byreflection_analysis_prompt.py(26,192) withanalysis_prompt.py(9,508) andmutation_prompt.py(8,586), and the scripts the sandboxed agent is expected to run:validate_analysis.py,append_changelog.pyanddiff_dirs.py, the first two registered as console scripts so they are discoverable on PATH inside the sandbox. - antomnievo/interface/, store/, evolution_algorithm/, model/
- The contracts, the memory and the selection: six abstract interfaces (
System,Evaluator,Proposer,EvolutionAlgorithm,CandidateStore,DataInst), a 30,684-byte filesystem store beside its 26,020-byte interface, the 6,960-byte micro-pareto algorithm, and the pydantic models for candidates, budgets, statistics, trajectories and tunable-artifact schemas — including nine files undertunable_artifact_defs/, the largest being text2sql at 21,105 bytes and the RAG pipeline at 17,431. - antomnievo/system/, evaluator/, dataset/, example/
- Four reference targets and their scaffolding: a LangGraph ReAct agent, text2sql over BIRD, DP and Spider2Snow, AppWorld and TerminalBench systems; one evaluator each, the AppWorld and TerminalBench ones at 9,511 and 8,793 bytes; eight dataset loader directories; and eight example entry points that wire the components together. Every one of them expects an external codebase the framework does not install.
- visualizer/
- A second package,
ant-omnievo-visualizer: a Flask backend (server.py11,616 bytes,workspace_reader.py10,785,manage.py6,534) and a Vite and React front end whose 39 source files includeGymBackdrop.tsxat 30,427 bytes,MaraChains.tsxat 27,868 and a 58,913-byte stylesheet, plus 40 files and 32 MB underpublic/, nearly all of it sprite art. - docs/, skills/, .aci/
- Sixteen documents in eight English and eight Chinese pairs, led by
system-design.mdat 13,706 bytes with its counterpart at 12,517; a 74,597-byteSKILL.mdfor theeddy-antomnievo-experimentskill with nine recipe references, the longest 38,238 bytes, written for Ant’s internal users; and the two files of the internal release pipeline.
Choices, and what they beat
The tunable artifacts are a directory of files over a constrained parameter set or a domain-specific language
Stated as the first design point: a real directory lets the framework use a mature coding agent that edits files directly, and makes expressive power “whatever the filesystem can hold” — a markdown rule, a python implementation, a verification script or a runtime extension. Every edit is then forced into a structured record that has to cite the failed-trajectory observation it answers.
Keep a rejected candidate and refine it over retiring it as soon as it misses the threshold
The paper the README cites argues that discarded candidates contain information later proposals need, and that discarding them makes the next proposal revisit the same failure modes. The chain is bounded to a fixed depth and the pool is bounded by Pareto-filtered top-N selection, so the extra attempts cannot run away with the budget.
Micro-pareto: dominance pruning with a top-N fallback over pure Pareto weak dominance
The design document gives the reason — pure Pareto fails under probabilistic scoring — and adds that selection weights candidates by how many instances they lead on, to protect versatile candidates rather than specialists. The same document credits Pareto selection itself to GEPA.
A structured analysis layer between the runs and the edits over letting the mutating agent read the raw runs
Phase 1 turns a failed run into a
RunAnalysisof observations and prescriptions, each prescription naming the trajectory item it resolves, and self-checks it; Phase 2 then resolves contradictions across instances before editing, must record the diff through a changelog command, and fails when no file changed. Raw runs never reach the mutator.Enforce validation-set isolation in the runtime, not in the prompt over stating the rule in the prompt for both backends
The Pi backend is launched with two extensions, one hard-blocking reads of the validation runs and one blocking a changelog bypass, while the Claude Code backend relies on prompt-declared access rules. The design document draws the conclusion it wants the reader to draw: the Pi path is the more controlled one.
Stop opening slots, then let the running ones finish over killing work when a budget cap is reached
Hitting any cap stops new slots and drains the in-flight ones, which the features document pairs with checkpoint resume so that a long run neither overspends nor loses what it already paid for; it calls this a soft stop.
Read fromdocs/system-design.md (13,706 bytes), docs/features.md, docs/extensibility.md, docs/when-to-use.md, docs/quickstart.md, docs/workspace-artifacts.md, docs/checkpoint-resume.md, docs/visualizer.md, README.md, the pyproject.toml of both packages, .aci/python/.aci.yml with its README, and the bodies of the nine pull requests.
Build log
6 stages- 01
Fifteen days, 21 commits, nine pull requests and no releases
The repository was created on 2026-09-15, and its first commit,
init, landed two hours and eighteen minutes later; the newest is the merge of pull request 9 on 2026-09-30. That is 21 commits in fifteen days, all of them in September, and one account stands behind every one of them:jacklv111is the linked account on all 21 and the author name on 9, the other 12 carrying a second author name written in Chinese characters and an address at 163.com. Seven commits carry aCo-authored-by: Claudetrailer. Nine pull requests were opened in the same window and all nine were closed by the same account with no comments at all — their bodies are the author’s own change notes rather than conversations, which is where most of what follows comes from. No release and no tag was ever created; the version lives only as0.1.1in twopyproject.tomlfiles, and what publishes them is an internal pipeline,.aci/python/.aci.yml, whose README describes a manual “Verify & Confirm” gate, a requiredPIP_INDEX_URLso the isolated build cannot reach the public index, and two named approvers. Around it sit 866 stars, 86 forks, 35 watchers and 0 open issues in a repository of 45,613 KB. - 02
One slot, and what has to be true before a child is kept
The framework calls an iteration a slot, and the design document writes out every step of it. The optimizer derives a batch from the parent’s own
(epoch, dataset_index)cursor — tail-aligned, wrapping to the next epoch at the end, and advanced even when a rollout fails, so the data order is a snapshot rather than a reaction. Parent and mutated child are then rolled out on the same batch, and the comparison is a plain sum: improvement equals the sum of the child’s scores minus the sum of the parent’s. Two things clear the bar — an improvement at or abovemin_improvement_per_batch, or a perfect score. A child that clears it is re-scored on the validation set, gets asummary.jsonand flips back topending; one that misses goes to the reflection chain if reflection is enabled, and is retired if it is not. After every slot the parent advances its cursor,eliminateruns, and a line goes toiteration_records.jsonl. The quickstart expectsbatch_size=3,max_iterations=1000,num_proposals=1,max_reflection_iterations=2,min_improvement_per_batch=2.0andmax_candidate_num=3. Slots run concurrently inside one process, and any budget cap stops new slots opening while running ones drain — the features document counts four hard caps, the design document five, the extra one beingmax_system_runs, which also counts the validation runs the proposer may not read. - 03
A rejected candidate is kept and refined, not discarded
When a batch misses its threshold the framework does not stop at a single natural-language reflection. A mara chain starts: the reflection prompt is pre-loaded with the whole chain context — the chain id path, a per-instance score table with delta columns, the attributed changelog, every earlier link’s analysis for that instance, and the chain links’ run files — and the analysis agent is then made to run a five-method diagnosis, fill in an A–I diagnosis table, prefer REMOVE and REPLACE over adding, name in
artifact_issuewhich link of the chain it holds responsible, and hunt for “lost wins”. Up tomax_reflection_iterationsrounds run this way; the best chain member that beats the original parent is validated and kept and the rest are retired. Phase 2, the mutation step, is mechanically unchanged on the reflection path, which is where the design puts the whole of the chain’s learning: in the analysis JSON, not in new machinery. The features document is explicit about what is borrowed and what is not — Pareto selection is adopted from GEPA, and the deltas claimed here are concurrency and the mara chain — and the paper the README cites rests on the observation that a discarded candidate often holds information later proposals need, so throwing it away only makes the next proposal visit the same failure again. - 04
Where “optimizes anything” actually stops
The description promises a framework that optimizes anything, and the boundary is written down: the thing under optimization must be expressible as a directory of tunable artifacts, and its evaluation must be repeatable and reasonably cheap. Three shapes are claimed — an AI agent’s skill, harness or memory directory; a workflow of config plus node code; and a single-file algorithm plus its description — and it says twice that the target need not contain a model. The proposer, unlike the target, must be an LLM coding agent. The exclusions are just as plain: no tunable artifact at all (a few scalar parameters, best handled by black-box tuning), evaluation that is not repeatable or is extremely expensive, and outputs under strict compliance, where it warns that the proposer edits files in a sandbox and the reader must audit the sandbox permissions. Two more limits sit in the code rather than the pitch: the framework does not manage the runtime or scoring harness of the target, and distributed evolution is out of scope — a cloud store is a “local plus transparent sync” subclass you write, and the document says this is not several workers sharing one store across hosts. The example entry points confirm a third limit by not running: each depends on an external codebase, and one still carries an absolute path into the author’s home directory as its workspace.
- 05
Ten successes and an empty results directory
The most useful pull request in the repository is a bug fix, and its body states the root cause, the wrong behaviour and the fix. Under rate limiting at the model gateway the model returns empty assistant messages; the agent CLI retries internally and then exits 0, so nothing ever reaches
trajectory.errors. Phase 1 counted every such session as a success — the log said “10 succeeded” — while the analysis result directory stayed empty, and Phase 2 then proposed on “(no analysis results available)” and edited nothing. Two of the three changes in that pull request are spelled out in the material: afind_degenerate_endingcheck that treats two or more trailing empty model spans, or zero tool calls with zero text, as a failure and raisesPiEmptyResponseErrorinto the retry path the 429s already used, with the same backoff and the existing retry hook; and a Phase 1 assertion, after a clean exit, that the analysis result file really was written or rewritten since the run started, so a silent session counts as a failed instance rather than a successful one. The third change, scoped to Phase 2, is cut off in the excerpt available here. What makes it worth recording is the shape of the failure: a counter, a metric and a log line all agreed that work had been done, and the only thing that disagreed was the empty directory. - 06
What is recorded about a run, and what is not
Nothing in the repository publishes a result: no benchmark table appears in the README or the eight English documents beside it. The only evidence of real runs is visual — the visualizer’s documentation embeds six screenshots: a gym where candidates are drawn as lobsters in score-tiered rooms, a lineage tree, a best-score curve, a coverage grid of the questions no candidate solved, and a file viewer. What it publishes instead is a method for checking a run yourself: compare
baseline_avg_scorewithbest_avg_scoreinlogs/statistics.json; readiteration_records.jsonlfor each slot’s accepted flag, its old and new batch score sums and itsreflection_depth, a value above zero meaning reflection rescued it; then walkparent_idto the root and readchangelog.jsonl, the append-only lineage the workspace document calls the history of why the artifacts look as they do. The numbers it does have are in the paper the README began linking on 2026-09-30: its abstract claims up to 20.5% relative improvement over GEPA, ACE and SkillOpt-Lite on AppWorld with 65.5% fewer rollouts to reach the target score, pass-rate gains of 20.2 and 22.5 points over AHE and Meta-Harness on TerminalBench 2.1, and MuSiQue nDCG@10 and Recall@10 gains of 0.104 and 0.131 over a hand-written retrieval pipeline. Those are claims in an abstract by the same team; nothing in the material reproduces them.
Adjacent records
All records →No. 129
Agents Universe
An open-source agent platform that keeps one project context shared by everyone working in it. Agents read the whole knowledge base when a project is opened and write what they learn back into the same files while they work; knowledge is Markdown on disk with a database index behind it, and there is no embedding model or vector search.
No. 123
agent-memory
A long-term memory runtime for AI agents that keeps plain Markdown files as the single source of truth, ranks them locally without calling a model, answers recall with file paths the agent opens one level at a time, writes at conversation boundaries rather than on the agent’s initiative, and runs an independent sleep-time layer that may add and update on its own but can only ever file a deletion as a proposal — one store shared by Claude Code, Codex CLI and Hermes, with no API key.
No. 117
sepia
A portable de-AI writing skill: four operations over one canonical rules file, narrative architecture repaired before word choice on fiction, a thin rule file matched to the venue on professional prose, and every rule labelled as measured, consulted or the project’s own inference.