book-to-skill
A converter that reads a document — one file, a folder, a glob or a list of paths, in PDF, EPUB, DOCX, HTML, RTF, MOBI or plain text — and writes an agent skill: a core file of mental models with a chapter index, one file per chapter that loads only when a question touches it, and a glossary, a patterns file and a cheatsheet beside them.

What it is
book-to-skill turns documents into Agent Skills, the SKILL.md folder format that GitHub Copilot CLI, Amp, Claude Code, Hermes Agent, OpenCode and OpenClaw all read. It is two halves that the repository refuses to blur. A deterministic Python extractor takes a file, a folder, a glob or a list of paths — PDF, EPUB, DOCX, HTML, RTF, MOBI or plain text — and writes a merged full_text.txt and a metadata.json holding pages, words, tokens, chapters and table of contents; every format has a standard-library fallback, and one unreadable source is skipped rather than killing the batch. Then the reader’s own coding agent follows a 55,009-byte specification to distil that text into a folder: a core of mental models and a topic index at about 4,000 tokens, one file per chapter at about 1,000 tokens each, plus a glossary, a patterns file and a decision cheatsheet. Around that sit analyze-only, fold-in and publish modes. The claim underneath is a token bill: an agent reading a PDF re-processes the table of contents on every turn, and this pays that structuring cost once, at conversion.
Who built itA solo maintainer whose account wrote 89 of the repository’s 202 commits. GitHub lists 41 contributors in all, and the largest after the owner is dex0shubham with 23, then Hotragn with 11, Stamina9 with 10 and a dependency bot with nine. The owner’s commits are signed under three names — Virgilio Junior, Virgilio Borges and the account handle — which is why the per-name counts split. Eighty-six commits carry a co-author trailer, and most of the names in them are tools rather than people: Claude Opus 5 on 21, Claude Opus 4.8 on 16 with a further seven on a million-token context, Copilot on nine, Amp on three, Cursor on one.
How it is put together
The parts · 6Two programs with a contract between them is the whole design, and the line is drawn deliberately: everything that can be made deterministic and tested is a Python program, and everything that requires actually reading a book is left to the agent the reader already has, driven by a written specification. That keeps the risky half — extraction, chapter counting, sanitisation — out of model judgement, and keeps the generator half out of code, so conversion behaviour can be changed by editing a step in SKILL.md without releasing anything. The output shape comes from one economic claim: the expensive part is not reading the book but re-navigating it, so the navigation is done once, at conversion, and stored as an index. That is why a generated skill is a small front page of mental models, a topic index and a directory of chapter files, and why the project measures itself in tokens saved rather than books converted. Two habits shape the code around it: every format has a standard-library fallback, so a missing optional dependency degrades instead of failing, and every claim has to be reproducible, which is why there is more test and evaluation scaffolding here than product code.
- SKILL.md
- The generator half: 55,009 bytes and the largest file in the tree that is read as prose. Numbered Steps 0–10 plus a fold-in workflow that merges new material into an existing skill. It loads into the agent on every conversion, which is why the repository treats any net growth in it as something that has to earn its context.
- book_to_skill/
- The extractor package: nine modules and 97 KB, of which
utils.pyis 65,957 bytes — CLI parsing, multi-source resolution, chapter and table-of-contents detection, the runner — with configuration, an optional-dependency prober that answers--check, an exceptions module so a per-source failure stays batch-safe, and 6,357 bytes of text sanitisation. - book_to_skill/parsers/
- One module per format, eight files and 38 KB: PDF with its Docling path and its fallbacks, EPUB, DOCX, HTML, RTF, Calibre and plain text. Each prefers a library and falls back to the standard library, and most of the bugs in the report live in the fallback branches.
- tools/ and evals/
- Three programs and a sub-package:
scan_generated_skill.pyat 14,278 bytes,validate_skill.pyat 11,707 bytes with a per-host lens,discovery_tax.pyat 9,177 bytes to measure a context dump against an indexed lookup, and anevals/package with a manifest, a scorer, a replay tool and a flat-paper baseline. - tests/
- Forty-two files and 314 KB, one of which is a 111,646-byte single test module — larger than any source file in the repository. The fixtures are synthetic books and recorded trajectories, and five tests are gated behind optional imports, which is itself an open issue: continuous integration installs only
pytest, so those five skip on every run. - docs/ and the repository root
- Ten files under
docs/— led by a 25,559-byte evaluation ledger underdocs/research/— plus a MkDocs configuration and a deploy workflow; 1,106 KB of that area is seven images of the project’s wizard mascot. At the root sitAGENTS.mdat 5,198 bytes of execution contract, three READMEs, a 15,915-byte changelog that contributors are told not to edit by hand, a security policy and notice, and three workflows.
Choices, and what they beat
Ask what kind of book it is, then pay for the right extractor over one extractor for every document
The repository’s own measurement is the argument: Docling costs 164 seconds against 0.1 seconds and returns the same token count, so it is not a quality default — it is bought for books with tables and code, where it recovers 48 tables and 36 code blocks that the fast path loses entirely. The choice is made at Step 1.5, before extraction, by asking the reader.
A structure index instead of a summary or a copy over storing passages and retrieving them at question time
The architecture document and the specification say the same thing from opposite ends — extract named frameworks, decision rules and anti-patterns, never raw passages — and the copyright section is what that buys: a generated skill is described as the reader’s own notes, and the tool can say it ships no book content. The quality rule that forbids copying is cited by name in the licence argument, so the legal position rests on the generation rule rather than beside it.
Chapter files that load on demand over one large document per book
The cost is only paid where a question lands: the core file is about 4,000 tokens and each chapter about 1,000, and a chapter is read only when the topic index sends the agent there. The stated reason for the layout is compaction — the core is written front-loaded because truncation takes from the end, which is a property of the host, not of the book.
Keep the fallback, but make it say so over either dropping the optional dependency or leaving the silent fallback in place
Every format degrades to the standard library, and one report showed what that costs: in technical mode, when Docling was installed but failed at runtime it was treated as absent, the run exited zero on flattened text with no tables and no code blocks, and the metadata still said technical. The fix kept the fallback and the exit code and added a machine-readable reason per source — unavailable, empty or failed — so the generator can ask before producing a technical skill out of text.
A hypothesis stays out of the product until its gate passes over adding an idea because it sounds right
The repository’s agent contract forbids turning a paper-derived hypothesis into production behaviour before its gate passes, names the features that may not be added on plausibility — metadata keys of a particular shape, a library mode, deeper routing, more
SKILL.mdcontent — and requires measurement rather than assertion for any quality, token, routing or cost claim. The evaluation discipline is written down too: unit tests before live runs, a small discriminating sample before a sweep, cached generation keyed by source and configuration, and a cost ceiling registered before the run.
Read fromdocs/architecture.md and docs/how-it-works.md, both printed in full in the recon report; AGENTS.md and README.md, also in full; the issue and pull-request bodies quoted in the log; and the complete 119-file tree with sizes.
Build log
6 stages- 01
Two halves: a deterministic extractor and a 55,009-byte specification
The architecture document states the split as a rule, not a description: a deterministic Python extractor, and a spec-driven generator that is whatever coding agent the reader already runs.
SKILL.mdis the second half, and it is the largest thing in the tree that is read as prose — 55,009 bytes, smaller only than the extractor’s 65,957-byteutils.pyand a 111,646-byte test module. It is written as numbered steps: Step 1.5 asks whether the book is technical or text-heavy, Step 2.6 is a read-eval-print loop for large books that greps and slices instead of re-reading them, Step 7 writes the per-chapter summaries on a budget that is the product of book type and depth, and Step 9.5 scans the files that were just written. The extractor half is more conventional:parsers/is eight modules and 38 KB,sanitize.pyis 6,357 bytes, andscripts/extract.pysurvives as a 1,333-byte shim so that older invocations keep working. Its output is two files in a per-run work directory — merged text with source markers, and metadata with the counts — deleted when the conversion ends. The command surface is wider than one command: analyze-only, generate-from-analysis and an update or fold-in mode that merges new material into a skill that already exists, plus an optional publish step that puts the result into a private GitHub repository so any host can install it withnpx skills add. - 02
What a book costs to read, and what the conversion costs
The extractor picks its tool by book type instead of trying everything. Prose goes through
pdftotext, falling back topypdfand thenpdfminer.six, all three effectively instant. Books with tables and formulas go through Docling at roughly 1.5 seconds a page, and the repository publishes the trade in one table: on a 103-page technical book run on CPU only,pdftotextfinished in 0.1 seconds and produced 27K tokens with no tables and no code blocks, while Docling took 164 seconds and produced the same tokens, 1.2% more, with 48 tables and 36 code blocks kept as markdown. Four real conversions follow, with pages, extracted tokens, auto-detected chapters and a one-pass estimate priced on Claude Sonnet 4.5 at three dollars in and fifteen dollars out per million tokens: Think Python 2 at 244 pages, 119K tokens, nineteen chapters and about eighty-eight cents; Working Backwards at 371 pages, 175K tokens, ten chapters and about ninety-six cents; Pro Git at 501 pages and 229K tokens; and Moby-Dick as an EPUB at 301K tokens. The last two carry the honest footnote: automatic chapter detection needs an explicit “Chapter N” or “Capítulo N” heading, Pro Git uses section titles and Moby-Dick uses chapter titles and roman numerals, so neither segments on its own — extraction and conversion still work, and the reader points at sections by hand. - 03
Chapter detection is where the project actually lives
A chapter count decides what the generated skill looks like, and thirteen of the thirty issues and pull requests in the report concern what a chapter is or how it is counted. The multilingual half arrived as the same patch over and over: Malayalam, Gujarati and Odia headings, each adding a pattern, a map from native digits to Latin ones, a dispatch branch placed next to the existing Brahmic blocks and tests beside their neighbours, with the contributor behind three of them at 23 commits, second only to the owner. The other half is false positives, and they are the more interesting ones. A sentence wrapped at the column limit whose continuation line began with a cross-reference was counted as a chapter; so was a fenced unified diff; and on the structural path the rule took the shallowest heading depth with at least two distinct titles, which counted a preface, a table of contents, an appendix, a glossary, a bibliography and an index along with twenty real chapters. Two headings called “Part” collapsed the count in the other direction, reporting two chapters for a book of ten. The maintainer reproduced each with synthetic input and split them apart: numeric false positives in one change, structural selection and the sample of headings that produced the count in another, and a blanket prefix exclusion refused because it would have deleted a genuine “Introduction” chapter.
- 04
Eight reports in one evening, each reproduced against the real library
Between 2026-09-30 and 2026-10-01 one account filed eight of the report’s last items, each reproduced against the real dependency, not a mock.
ebooklib0.20 extracts an EPUB in manifest order, ignoring the declared spine and reordering chapters the standard-library fallback keeps in order.python-docx1.2.0 emits every top-level paragraph before every table, so a table between chapters one and two lands at the end. On Windows, a Markdown source with ordinary CRLF endings is persisted with doubled carriage returns, so a plain read then sees a blank line between every line, though the count printed before the write is right. BeautifulSoup inserts a newline between every text node, inline ones included, so a synthetic two-chapter source containingChapter <span>1</span>yields zero chapters and exits cleanly. Three more aim at the guard that lets an interrupted run reuse its corpus:reuse_is_safe()says the work directory is intact and every source matches even when the extracted text has been deleted, emptied or replaced, and when two inputs in different directories share a basename and a hash, a substituted third file is never examined. The reporter is careful about the claim — what fails is the helper’s verdict, not a known bad skill — and the gap is the seam an earlier change left: every source’s name and hash are recorded and checked, nothing checks the output. - 05
What it refuses to do with the file, and what it refuses to ship
Scanned pages are handled by stopping, not by trying harder: a PDF with no text layer is caught on the first pages and exits with an explanation rather than grinding through the book to produce an empty skill; the README says to run
ocrmypdffirst. Copyright is the same kind of decision, and the design rests on it: the tool ships no book content, extraction runs on the reader’s machine, and a generated skill is a structured derivative — framework names, definitions, takeaways — because the specification forbids copying raw passages at all; its own tests use synthetic or licensed fixtures for the same reason. The security work follows the document’s path into an agent: invisible and zero-width characters are stripped from every parser’s output before anything is measured; a DOCX part declaring a DTD is rejected before parsing; paths are absolutised before reachingpdftotextorebook-convert, so a name starting with a dash cannot be read as a flag; and the generator ends with an advisory scan that names a rule and a location, never the matched text. When a reader showed that nineDefault_Ignorablecode points survived the scrub, the maintainer confirmed the gap, refused a blanket format-character filter because it would strip semantic controls, called it a coverage gap and not a successful prompt injection, and asked for the rest through private vulnerability reporting. - 06
Five releases, a quiet September, and 33,182 stars the material does not explain
The first commit is stamped six seconds before the repository itself, so the project arrived from a working copy rather than the platform. From there: 202 commits in 151 days, in a monthly shape that reads like a response rather than a schedule — twenty in May, sixty-two in June, nineteen in July, seventy-four in August, twenty-seven in September. Five releases came out of the first three months, none since:
v1.0.0on 2026-06-08,v1.1.0four days later,v1.2.0on 2026-06-17, an installable package plus multilingual chapter detection,v1.3.0on 2026-07-30 andv1.4.0on 2026-08-10. September brought twenty-seven commits, a long stretch of triage and no release — the last commit is a contributor’s EPUB spine fix merged on 2026-09-29 — so the record says active, not finished. A reviewer also showed the translated READMEs had drifted past the English destination policy, each pinning the revision it was made from with nothing enforcing the pin, and the fix was accepted on two conditions: the pin may only move to a source actually reviewed, and no check may depend on git history, which shallow checkouts would fail. The material contains no account of the 33,182 stars; the only traces of how it travelled are two Trendshift badges for repository 27038, a companion use-cases repository whose index takes a one-line pull request, and 3,466 forks against 41 contributors.
Adjacent records
All records →No. 055
LeanCTX
A local layer that sits beside a coding agent and decides what reaches the model: file reads are compressed and cached, command output is compressed by per-command rules, session findings persist across chats, and a local proxy rewrites each request without breaking the provider’s prompt cache — with a savings ledger, a budget and a dashboard for what it measured.
No. 125
OKF Agent Memory
Keeps what a coding agent learns as plain Markdown inside the repository — an OKF v0.2 knowledge bundle searched in-process by BM25 — so the memory can be diffed and reviewed instead of living in a database.
No. 123
agent-memory
A long-term memory runtime for AI agents that keeps plain Markdown files as the single source of truth, ranks them locally without calling a model, answers recall with file paths the agent opens one level at a time, writes at conversation boundaries rather than on the agent’s initiative, and runs an independent sleep-time layer that may add and update on its own but can only ever file a deletion as a proposal — one store shared by Claude Code, Codex CLI and Hermes, with no API key.