Skip to content

Open Mercato Skills

A pack of forty-one agent skills that installs into any repository and runs the whole pull-request pipeline — plan, implement in an isolated worktree, self-review, browser-checked QA, merge — with three entry paths (a task brief, a specification, or a tracker issue) converging on one review loop, every project-specific value read from a single committed configuration file, tracker commands confined to committed descriptor documents shipped for GitHub, GitLab, Linear and Jira, and each skill handing the next one a machine-parsed PR reference line.

Screenshot of Open Mercato Skills
Editor screenshot, 30 Sep 2026Open Mercato Skills ↗

What it is

Forty-one agent skills that run an entire pull-request pipeline, extracted from the ERP product their authors built at Open Mercato so that any repository can run the same process. They install with one command through skills.sh and work with a couple of dozen coding agents. Three entry paths — a free-form task brief, a specification, or a tracker issue — converge on one review loop and one QA gate, and each skill ends with a machine-parsed reference line that the next skill consumes. Everything project-specific comes from one committed configuration file: base branch, validation commands, label taxonomy and working paths; tracker and browser commands live in committed descriptor documents, shipped for GitHub, GitLab, Linear and Jira, so no skill invokes a tracker command of its own. The team’s own claim, unverified here: roughly 800,000 lines of code with no hand-written lines, 1,700+ merged pull requests, 4,000 unit tests, 730 integration tests and weekly releases; the repository description instead says 1.2M+ lines, a figure the README does not repeat. MIT licensed.

Who built itThe repository belongs to the Open Mercato organization, and Karwatka is its dominant committer: 274 of the 369 commits are attributed to the account pkarw, which matches 271 commits across his two email addresses plus three from a GitHub no-reply address. Ten other accounts have contributed, the largest being matgren with 41, and 154 commits carry a co-author trailer of which 153 name a Claude model — 102 of those reading only “Claude Fable 5”. The README closes by saying the team teaches this way of working in a Polish AI engineering course, and two HTML comments in it still ask him to rewrite the copy in his own voice.

How it is put together

The parts · 5

A document-first framework rather than a program: what runs is an agent reading numbered instructions, so the architecture is a set of contracts between documents. One committed configuration file holds every project-specific value; tracker and browser operations are named abstractly and defined in committed descriptor documents, which makes a new tracker a new file rather than a code change. Each skill is laid out as a router plus a map — the numbered algorithm in SKILL.md, the repeatable procedures in per-skill references/ files under standard names — because the body is paid for on every invocation and a reference only when it is read. The standard step files are deliberately duplicated inside every skill that uses them instead of being shared through pointers, so each skill installs and runs standalone; the stated cost is drift, and the stated management for it is a contributor rule to diff the other copies and ask before propagating a change.

skills/
Forty-two skill directories, each a SKILL.md plus its own references/, laid out under the same standard file names. The largest package is om-setup-agent-pipeline at 17 files and 197 KB, followed by om-auto-create-pr-loop (20 files, 90 KB) and om-auto-review-pr (17 files, 90 KB). The biggest single document in the repository is the GitLab tracker descriptor at 50,743 bytes, against 26,218 for the GitHub one it mirrors operation for operation.
scripts/
The enforcement layer: lint.sh (8,375 bytes) as the gate every pull request must pass, plus five Node contract suites — tracker providers 20,524 bytes, discovery contracts 19,632, the run classifier 8,438, close keywords 7,528, browser providers 6,124 and the spec-writing data model 2,151 — with a Codex browser test shell script, a skill installer, an audit script and a script that bumps a pinned browser release.
.ai/
The pipeline’s own workspace inside the repository that ships it: agentic.config.json at 1,210 bytes, seven specification documents totalling 84 KB, twenty run records totalling 77 KB, the GitHub tracker descriptor at 26 KB, and the installed browser descriptor at 9,302 bytes — the exact file whose drift turned the default branch red.
Root documents
UPGRADE_NOTES.md (66,220 bytes) and DECISIONS.md (66,206) are the two largest files in the repository, ahead of the README at 43,608, SDLC.md at 21,546, AGENTS.md at 15,637, BACKWARD_COMPATIBILITY.md at 8,577, CODE_REVIEW.md at 6,256 and CHANGELOG.md at 6,009. Five role pages sit under docs/roles/ and 41 skill cards under docs/skills/, an index with no card for om-qa-buddy.
.github/workflows/
Three workflows: lint.yml (670 bytes) as the frontmatter and product-agnosticism gate, skills-audit.yml (640 bytes) as an informational third-party audit, and agent-browser-pin.yml (2,052 bytes) whose weekly run opened PR #97 by itself.

Choices, and what they beat

  • Each skill carries its own copy of the shared step files over sharing them through cross-skill pointers

    Standalone installability over DRY, stated in the README and repeated in the cross-skill contract: a skill installs by itself, so it may not point into another skill’s references directory. The accepted cost is duplication, managed by a contributor rule — when one of those standard files changes, diff the same file in the other skills and ask whether to sync it, listing the skills that would change.

  • A descriptor document per tracker instead of calling a tracker command over each skill invoking the tracker command line itself

    No skill calls a tracker command line directly; skills name operations and one committed descriptor defines how each is executed. That is what let a stand-alone GitLab provider cover all 42 operations the GitHub descriptor defines, over REST through glab, with no change to any consuming skill.

  • A repo-local skill wins, but cannot weaken the installed one over repo-local skills replacing installed skills outright

    Stated in the README: local rules win, but a local skill can never relax the installed skill’s safety rules — no skipping tests, no bypassing commit hooks, no force-pushes. The same mechanism is what makes the collection a drop-in for repositories that already keep their own skills under the local path.

  • Replace the upstream-first merge gate with a verified import from a pinned commit over waiting for the upstream pull request to merge before the imported files could be re-diffed

    Recorded in the pull-request thread and then in the specification: after the skill was renamed and its workflow rewritten, a re-diff against a merged upstream would no longer verify anything, and this collection is the source going forward. The replacement gate is a pinned source commit, a mapping of all 13 imported files, accounted differences, retained evidence and a fresh review.

  • Leave four imported defects for upstream instead of patching them here over fixing known polish items locally

    The findings all live in files imported from the upstream monorepo, and the rollout gate required those files to be re-diffed against the merged upstream commit; patching them locally would have created exactly the drift the gate existed to catch, so they were filed as a follow-up issue with the verification steps recorded.

Read fromREADME.md (43,344 characters, read in full from the main branch), AGENTS.md (15,414 characters, including the task-routing table, the binding cross-skill contract and the two skill-authoring standards), the complete 459-file tree with per-file sizes, the two-level directory totals, and the pull-request and issue bodies cited above — PRs #96 to #124 and issues #99 to #121 with their bot comment threads.

Build log

6 stages
  1. 01

    A claim, an extraction, and three months of commits

    This repository earns its place because of the sentence it puts in its own description: Enterprise AI Engineering skills, coined at Open Mercato, over 1.2M+ lines of ERP code built with AI. That is the project’s own figure and the largest number it uses — the README credits the same workflow with roughly 800,000 lines, 1,700+ merged pull requests, 4,000 unit tests, 730 integration tests and weekly releases, with 100+ contributors. None of it is verified in this record; what is verifiable is the extraction around it. The organization registered the repository on 2026-06-26, the first commit arrived on 2026-07-02 with a message saying only that it is a repository stub plus a README and a licence, and 369 commits followed in under three months: 244 of them in July, 68 in August, 57 in September. Two releases came out of that — v1.0.0 on 2026-07-21 and v1.1.0 on 2026-08-13 — and around it sit 210 stars, 30 forks, 22 open issues and eleven contributor accounts, one of which holds 274 of the 369 commits.

  2. 02

    What the extraction left behind

    A pipeline lifted out of a product does not arrive intact, and the tracker keeps the receipts. Two shipped skills went on delegating work to skill names that release 1.1.0 had already removed — the first had been renamed, the second absorbed into another skill. Both references were written as optional dependencies, so at run time the agent read them as a component the user had simply not installed, skipped the hand-off and reported success; nothing failed loudly, and the regression survived a release. The fix pointed them at the current names and added a lint gate so the next stale name cannot ship. The second loss was quieter: the anti-bundling check in the specification skill had been shipped upstream in two halves, the rule in the skill body as a mandatory open question and the evidence that made the rule fire in a separate references file — and only the body travelled. The check came through the migration with nothing to fire on.

  3. 03

    The rules are documents, and the gate is a script

    What keeps forty-odd instruction documents from drifting apart is a task-routing table in AGENTS.md with a binding contract behind it: a skill’s frontmatter name must equal its directory name and carry a description of at most 500 characters, the lint gate greps the skills tree for product references and for a hard-coded base branch or package manager, tracker state may only be changed through named operations, and no configuration value may be written into a skill. On top of that sit scripts/lint.sh (8,375 bytes) and five Node contract suites covering the tracker providers, the browser providers, the discovery skills, the run classifier and the close keywords. A weekly workflow refreshes a pinned browser release and opened its own pull request, with a note that the descriptor it changes is what pins the binaries setup downloads and verifies by SHA-256. PR #114 shows how a rule gets installed here: a mandatory, exact-path local-override preflight added to all 37 skills that existed on 2026-09-11, a lint invariant that rejects a skill whose preflight is missing or points at the wrong path, and a negative probe as the proof — remove one preflight, confirm the exact error, put the file back.

  4. 04

    The repository that had to fix itself first

    This project configures the agent pipeline for other people’s repositories, and for two months it ran an out-of-date copy of that configuration on itself. Its own config was written on 2026-07-10, four days before the browser-provider contract landed; with no browser key in it, every browser-capable skill silently treated this repository as a Playwright repository, and the directory of provider descriptors the documentation describes did not exist here at all. The same file listed one of the five commands the integration workflow actually runs. A second self-inflicted break is the sharper one: the lint step requires the installed browser descriptor to be identical to the shipped one, so when a setup run used an older installed copy and wrote v0.34.0 content over the v0.35.2 descriptor, the default branch went red — and stayed red, because every pull request that merged the default branch inherited the failure. The repair was a verbatim copy of the shipped file, chosen precisely so that no reviewed content changed.

  5. 05

    Retiring a gate that had stopped verifying anything

    The prototype skill was imported into this collection from a pull request in the upstream product monorepo, and the rollout gate written for that import said the imported files could only be merged after the upstream pull request landed and the files were re-diffed against the merged commit. It stayed open, so the import sat in the queue for weeks — and when a reviewer finally took over the stale claim, the honest label was not changes-requested but blocked, because there was no code change left for the author to make. Four non-blocking findings were left unfixed on purpose rather than patched locally, since patching them would create exactly the drift the gate existed to catch. On 2026-09-15 the gate itself was replaced by agreement: a verified import from a pinned source commit, a verification record mapping all 13 imported files and explaining each intentional difference, and the new condition written into the specification instead of left in the pull-request thread. The reason given for retiring it is the interesting part — after the rename and the workflow rewrite, a re-diff against a merged upstream would no longer verify anything, and this collection is the source going forward.

  6. 06

    The pipeline running on the repository that ships it

    Read the pull requests and you are reading the pipeline’s own output: bot comments sealed with a skill marker and a purpose line, one consolidated label rationale per pull request with a full sentence per label, claim and release locks that make concurrent agents back off, and autofix iterations counted in the summary — eight review findings resolved in one iteration in one case, seven in another. The automation also keeps admitting a limit it cannot engineer around: GitHub forbids self-approval, so when the reviewing account is also the author, the verdict is posted as a comment and a second account has to add the native approval, which the reports say outright rather than claiming an approval. The same discipline produced a carry-forward pattern for fork contributions — a fork pull request with a base conflict was closed in favour of a replacement opened by a maintainer, the original author was credited in the closing comment, and a supersede credit rule exists so the changelog credits the author rather than the person who merged. What is still unwritten is the proof: the README carries two HTML comments asking the founder to rewrite the copy in his own voice, and its section on products built with the workflow says case studies are still being added — the claim that this pipeline shipped a real product has no case study in the repository yet.

Adjacent records

All records →