Skip to content

AntOmniEvo

A Python framework that runs an evolution loop over a directory of your files: a coding agent rewrites them from the failure trajectories of the previous candidate, parent and child are scored on the same training batch, and only an improvement that clears a threshold is re-scored on the validation set and kept.

Screenshot of AntOmniEvo
Editor screenshot, 1 Oct 2026AntOmniEvo ↗

What it is

AntOmniEvo is a Python framework that treats optimization as file editing. You supply five things — a System that runs the thing under optimization, an Evaluator whose 0–1 scoring criteria are the objective, train and validation instances, a schema for the files that may change, and the starting files — then a coding agent reads the failure trajectories, rewrites those files as a new candidate, and the framework decides whether it survives. The genome is a real directory: a mutation can be a Markdown rule, a node script or a retrieval DAG’s pipeline.json. Each iteration rolls a parent and its child out on the same batch, keeps the child only if the summed improvement clears a threshold or is perfect, re-scores it on the validation set, then prunes the pool by Pareto dominance with a top-N fallback. A candidate that misses is not discarded: a reflection chain refines it for a bounded number of rounds. The run is written to a resumable workspace, and a React and Flask visualizer renders it. No release was ever tagged: version 0.1.1 of the two packages goes to an internal index.

Who built itAnt Research is an organisation account rather than a person: 45 public repositories, 368 followers, opened on 2022-12-17, with antresearch.com on its profile. Every commit in this repository comes from one member account, jacklv111 — opened on 2022-09-20, 13 public repositories and 5 followers — which is the author name on 9 of the 21 commits, the linked account on all of them, and the only entry in the contributor list. The paper the README cites lists six authors, led by Yubin Lyu; the repository’s internal release pipeline names two approvers, lyuyubin.lyb and hanpu.mwx. Which account belongs to which name is not stated in the material.

How it is put together

The parts · 6

The organising idea is that optimization should be plain file editing: the thing being optimized is represented as a directory, and the machine that decides what survives is deterministic code the user does not write. That choice puts the intelligence in one place — a coding agent that reads a failed run and edits files — and keeps everything probabilistic out of the loop: deriving the batch, accepting or retiring, the validation re-score, the pruning and the budget are all counts, sums and comparisons. Two consequences shape the code. Because analysis and editing are the only steps an LLM is trusted with, they are split into two phases with a structured artifact between them, and the reflection path changes only that artifact, never the mutation machinery. And because a mutation loop built on an unreliable spender has to be auditable, memory is the filesystem: append-only changelogs that a child inherits from its parent, run records and analyses written before any error can short-circuit them, and a startup recovery that flips interrupted candidates back to selectable. The same documents are candid about the lineage: Pareto selection is borrowed from GEPA, and the parts presented as this project’s own are running several candidates at once and keeping rejected candidates alive.

antomnievo/optimizer/
43,728 bytes in optimizer.py: the main loop and its per-scenario subclasses for text2sql, AppWorld, TerminalBench, adconfig and the RAG pipeline. It owns the slot, the batch cursor, the accept-or-retire challenge, the reflection chain, the validation step, and the two counters a subclass cannot bypass.
antomnievo/proposer/
The coding-agent side: base_proposer.py (28,967 bytes) holding the two-phase analyze-then-mutate pipeline, with one thin subclass per backend — claude_code_proposer.py (6,054) driving claude, pi_coding_agent_proposer.py (5,776) driving pi — three prompt templates led by reflection_analysis_prompt.py (26,192) with analysis_prompt.py (9,508) and mutation_prompt.py (8,586), and the scripts the sandboxed agent is expected to run: validate_analysis.py, append_changelog.py and diff_dirs.py, the first two registered as console scripts so they are discoverable on PATH inside the sandbox.
antomnievo/interface/, store/, evolution_algorithm/, model/
The contracts, the memory and the selection: six abstract interfaces (System, Evaluator, Proposer, EvolutionAlgorithm, CandidateStore, DataInst), a 30,684-byte filesystem store beside its 26,020-byte interface, the 6,960-byte micro-pareto algorithm, and the pydantic models for candidates, budgets, statistics, trajectories and tunable-artifact schemas — including nine files under tunable_artifact_defs/, the largest being text2sql at 21,105 bytes and the RAG pipeline at 17,431.
antomnievo/system/, evaluator/, dataset/, example/
Four reference targets and their scaffolding: a LangGraph ReAct agent, text2sql over BIRD, DP and Spider2Snow, AppWorld and TerminalBench systems; one evaluator each, the AppWorld and TerminalBench ones at 9,511 and 8,793 bytes; eight dataset loader directories; and eight example entry points that wire the components together. Every one of them expects an external codebase the framework does not install.
visualizer/
A second package, ant-omnievo-visualizer: a Flask backend (server.py 11,616 bytes, workspace_reader.py 10,785, manage.py 6,534) and a Vite and React front end whose 39 source files include GymBackdrop.tsx at 30,427 bytes, MaraChains.tsx at 27,868 and a 58,913-byte stylesheet, plus 40 files and 32 MB under public/, nearly all of it sprite art.
docs/, skills/, .aci/
Sixteen documents in eight English and eight Chinese pairs, led by system-design.md at 13,706 bytes with its counterpart at 12,517; a 74,597-byte SKILL.md for the eddy-antomnievo-experiment skill with nine recipe references, the longest 38,238 bytes, written for Ant’s internal users; and the two files of the internal release pipeline.

Choices, and what they beat

  • The tunable artifacts are a directory of files over a constrained parameter set or a domain-specific language

    Stated as the first design point: a real directory lets the framework use a mature coding agent that edits files directly, and makes expressive power “whatever the filesystem can hold” — a markdown rule, a python implementation, a verification script or a runtime extension. Every edit is then forced into a structured record that has to cite the failed-trajectory observation it answers.

  • Keep a rejected candidate and refine it over retiring it as soon as it misses the threshold

    The paper the README cites argues that discarded candidates contain information later proposals need, and that discarding them makes the next proposal revisit the same failure modes. The chain is bounded to a fixed depth and the pool is bounded by Pareto-filtered top-N selection, so the extra attempts cannot run away with the budget.

  • Micro-pareto: dominance pruning with a top-N fallback over pure Pareto weak dominance

    The design document gives the reason — pure Pareto fails under probabilistic scoring — and adds that selection weights candidates by how many instances they lead on, to protect versatile candidates rather than specialists. The same document credits Pareto selection itself to GEPA.

  • A structured analysis layer between the runs and the edits over letting the mutating agent read the raw runs

    Phase 1 turns a failed run into a RunAnalysis of observations and prescriptions, each prescription naming the trajectory item it resolves, and self-checks it; Phase 2 then resolves contradictions across instances before editing, must record the diff through a changelog command, and fails when no file changed. Raw runs never reach the mutator.

  • Enforce validation-set isolation in the runtime, not in the prompt over stating the rule in the prompt for both backends

    The Pi backend is launched with two extensions, one hard-blocking reads of the validation runs and one blocking a changelog bypass, while the Claude Code backend relies on prompt-declared access rules. The design document draws the conclusion it wants the reader to draw: the Pi path is the more controlled one.

  • Stop opening slots, then let the running ones finish over killing work when a budget cap is reached

    Hitting any cap stops new slots and drains the in-flight ones, which the features document pairs with checkpoint resume so that a long run neither overspends nor loses what it already paid for; it calls this a soft stop.

Read fromdocs/system-design.md (13,706 bytes), docs/features.md, docs/extensibility.md, docs/when-to-use.md, docs/quickstart.md, docs/workspace-artifacts.md, docs/checkpoint-resume.md, docs/visualizer.md, README.md, the pyproject.toml of both packages, .aci/python/.aci.yml with its README, and the bodies of the nine pull requests.

Build log

6 stages
  1. 01

    Fifteen days, 21 commits, nine pull requests and no releases

    The repository was created on 2026-09-15, and its first commit, init, landed two hours and eighteen minutes later; the newest is the merge of pull request 9 on 2026-09-30. That is 21 commits in fifteen days, all of them in September, and one account stands behind every one of them: jacklv111 is the linked account on all 21 and the author name on 9, the other 12 carrying a second author name written in Chinese characters and an address at 163.com. Seven commits carry a Co-authored-by: Claude trailer. Nine pull requests were opened in the same window and all nine were closed by the same account with no comments at all — their bodies are the author’s own change notes rather than conversations, which is where most of what follows comes from. No release and no tag was ever created; the version lives only as 0.1.1 in two pyproject.toml files, and what publishes them is an internal pipeline, .aci/python/.aci.yml, whose README describes a manual “Verify & Confirm” gate, a required PIP_INDEX_URL so the isolated build cannot reach the public index, and two named approvers. Around it sit 866 stars, 86 forks, 35 watchers and 0 open issues in a repository of 45,613 KB.

  2. 02

    One slot, and what has to be true before a child is kept

    The framework calls an iteration a slot, and the design document writes out every step of it. The optimizer derives a batch from the parent’s own (epoch, dataset_index) cursor — tail-aligned, wrapping to the next epoch at the end, and advanced even when a rollout fails, so the data order is a snapshot rather than a reaction. Parent and mutated child are then rolled out on the same batch, and the comparison is a plain sum: improvement equals the sum of the child’s scores minus the sum of the parent’s. Two things clear the bar — an improvement at or above min_improvement_per_batch, or a perfect score. A child that clears it is re-scored on the validation set, gets a summary.json and flips back to pending; one that misses goes to the reflection chain if reflection is enabled, and is retired if it is not. After every slot the parent advances its cursor, eliminate runs, and a line goes to iteration_records.jsonl. The quickstart expects batch_size=3, max_iterations=1000, num_proposals=1, max_reflection_iterations=2, min_improvement_per_batch=2.0 and max_candidate_num=3. Slots run concurrently inside one process, and any budget cap stops new slots opening while running ones drain — the features document counts four hard caps, the design document five, the extra one being max_system_runs, which also counts the validation runs the proposer may not read.

  3. 03

    A rejected candidate is kept and refined, not discarded

    When a batch misses its threshold the framework does not stop at a single natural-language reflection. A mara chain starts: the reflection prompt is pre-loaded with the whole chain context — the chain id path, a per-instance score table with delta columns, the attributed changelog, every earlier link’s analysis for that instance, and the chain links’ run files — and the analysis agent is then made to run a five-method diagnosis, fill in an A–I diagnosis table, prefer REMOVE and REPLACE over adding, name in artifact_issue which link of the chain it holds responsible, and hunt for “lost wins”. Up to max_reflection_iterations rounds run this way; the best chain member that beats the original parent is validated and kept and the rest are retired. Phase 2, the mutation step, is mechanically unchanged on the reflection path, which is where the design puts the whole of the chain’s learning: in the analysis JSON, not in new machinery. The features document is explicit about what is borrowed and what is not — Pareto selection is adopted from GEPA, and the deltas claimed here are concurrency and the mara chain — and the paper the README cites rests on the observation that a discarded candidate often holds information later proposals need, so throwing it away only makes the next proposal visit the same failure again.

  4. 04

    Where “optimizes anything” actually stops

    The description promises a framework that optimizes anything, and the boundary is written down: the thing under optimization must be expressible as a directory of tunable artifacts, and its evaluation must be repeatable and reasonably cheap. Three shapes are claimed — an AI agent’s skill, harness or memory directory; a workflow of config plus node code; and a single-file algorithm plus its description — and it says twice that the target need not contain a model. The proposer, unlike the target, must be an LLM coding agent. The exclusions are just as plain: no tunable artifact at all (a few scalar parameters, best handled by black-box tuning), evaluation that is not repeatable or is extremely expensive, and outputs under strict compliance, where it warns that the proposer edits files in a sandbox and the reader must audit the sandbox permissions. Two more limits sit in the code rather than the pitch: the framework does not manage the runtime or scoring harness of the target, and distributed evolution is out of scope — a cloud store is a “local plus transparent sync” subclass you write, and the document says this is not several workers sharing one store across hosts. The example entry points confirm a third limit by not running: each depends on an external codebase, and one still carries an absolute path into the author’s home directory as its workspace.

  5. 05

    Ten successes and an empty results directory

    The most useful pull request in the repository is a bug fix, and its body states the root cause, the wrong behaviour and the fix. Under rate limiting at the model gateway the model returns empty assistant messages; the agent CLI retries internally and then exits 0, so nothing ever reaches trajectory.errors. Phase 1 counted every such session as a success — the log said “10 succeeded” — while the analysis result directory stayed empty, and Phase 2 then proposed on “(no analysis results available)” and edited nothing. Two of the three changes in that pull request are spelled out in the material: a find_degenerate_ending check that treats two or more trailing empty model spans, or zero tool calls with zero text, as a failure and raises PiEmptyResponseError into the retry path the 429s already used, with the same backoff and the existing retry hook; and a Phase 1 assertion, after a clean exit, that the analysis result file really was written or rewritten since the run started, so a silent session counts as a failed instance rather than a successful one. The third change, scoped to Phase 2, is cut off in the excerpt available here. What makes it worth recording is the shape of the failure: a counter, a metric and a log line all agreed that work had been done, and the only thing that disagreed was the empty directory.

  6. 06

    What is recorded about a run, and what is not

    Nothing in the repository publishes a result: no benchmark table appears in the README or the eight English documents beside it. The only evidence of real runs is visual — the visualizer’s documentation embeds six screenshots: a gym where candidates are drawn as lobsters in score-tiered rooms, a lineage tree, a best-score curve, a coverage grid of the questions no candidate solved, and a file viewer. What it publishes instead is a method for checking a run yourself: compare baseline_avg_score with best_avg_score in logs/statistics.json; read iteration_records.jsonl for each slot’s accepted flag, its old and new batch score sums and its reflection_depth, a value above zero meaning reflection rescued it; then walk parent_id to the root and read changelog.jsonl, the append-only lineage the workspace document calls the history of why the artifacts look as they do. The numbers it does have are in the paper the README began linking on 2026-09-30: its abstract claims up to 20.5% relative improvement over GEPA, ACE and SkillOpt-Lite on AppWorld with 65.5% fewer rollouts to reach the target score, pass-rate gains of 20.2 and 22.5 points over AHE and Meta-Harness on TerminalBench 2.1, and MuSiQue nDCG@10 and Recall@10 gains of 0.104 and 0.131 over a hand-written retrieval pipeline. Those are claims in an abstract by the same team; nothing in the material reproduces them.

Adjacent records

All records →