Skip to content

feat(library): Read filenames the way download sites and library managers write them - #31

Merged
OffCrazyFreak merged 3 commits into
mainfrom
feat/library-wild-filenames
Sep 16, 2026
Merged

OffCrazyFreak merged 3 commits into
mainfrom
feat/library-wild-filenames

Conversation

@OffCrazyFreak

@OffCrazyFreak OffCrazyFreak commented Sep 16, 2026 •

Copy link
Copy Markdown
Owner

Changes:

  • Recognise the naming schemes measured in the wild (OceanofPDF, Anna's Archive, Z-Library, Library Genesis, PDFDrive, scene folders, sharing channels, ebook-tools, Standard Ebooks and dokumen.pub slugs, Springer, Calibre and LazyLibrarian's title-first form) and undo each one's encoding before splitting author from title
  • Read a plain A - B author first as before, unless B reads more like a person, and carry the other reading along as FilenameFacts.alternate when both halves could be a person
  • Try the other reading only when no source identifies the book as first read, and keep it only if a source then does
  • Report a name that carries no title (a Gutenberg number, a bare ISBN, an ASIN) as such in the CLI and on the page instead of "no source answered"; print how a non-plain name was read
  • Have the page ask the worker how every filename reads, so pending rows use the same parser as the verdicts, and show the reading and scheme in the detail view
  • Record each catalogue once per row however many rounds ask it, and pace the page between books by the rounds the last book took
  • Document the audit and the decisions in docs/filenames.md

A file named the OceanofPDF way had no - in it, so the whole stem was read as an author with no title, nobody was asked, and the page said "no source answered", which was false. Renamed by hand the same three web sources reached HIGH. Nobody should have to rename anything: half the tools out there write the title first and every site adds its own marks, so the parser now knows the shapes rather than one convention. The audit behind the rules (Calibre, Calibre-Web, LazyLibrarian and Readarr defaults read from their source, Anna's Archive's filename builder, Z-Library, libgen, OceanofPDF and PDFDrive names quoted from public issues, Gutenberg, Standard Ebooks and Internet Archive checked live) is in the new docs page.

Safety model: matching.py, tags.py and the gain rules are untouched. What changed is how the filename is read and how the catalogues are asked; the bar a source must clear is the same. A wrong reading cannot score, since a source would have to name a book whose title is the author's name and whose author is the title, twice over.

Checked against: the two OceanofPDF pairs that started this, live with the web sources (one MED on Apple alone since the site left no trace of the colon, one LOW, the LOW being a publisher's summary edition the model rightly refuses), plus the newer OceanofPDF pair whose double underscore keeps the colon (HIGH with all three sources); the same book named the Anna's Archive, Z-Library and Calibre ways (HIGH each); the pristine sample replayed against fixtures-v2 (Kobo, Google, Open Library) and fixtures-wide-raw (web sources) with every verdict, source list, gain and figure identical to main and no extra query asked. A retry with the head of a glued title was built, then removed after review reproduced it proposing a different volume of a series (a strict prefix the pre-existing prefix rule scores 0.95); only the whole title is ever asked for now, and docs/filenames.md records why.

Checks: ruff check, ruff format --check, pytest (395 passed); pnpm typecheck, pnpm format:check, pnpm test (44 passed), pnpm build, node scripts/pyodide-smoke.mjs (401 passed under Pyodide). The page was run end to end in headless Chromium with the real files through the file input at desktop width; the phone width was covered by the sweep in #30.

Notes:

@coderabbitai

coderabbitai Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 43a8091b-0b01-4f28-bb0c-8b03e0183c3c

📥 Commits

Reviewing files that changed from the base of the PR and between cd4f7ed and d20c3b3.

📒 Files selected for processing (22)
  • README.md
  • docs/README.md
  • docs/filenames.md
  • src/ebook_metamend/cli.py
  • src/ebook_metamend/enrich.py
  • src/ebook_metamend/library.py
  • src/ebook_metamend/web.py
  • tests/test_library_and_opf.py
  • tests/test_pipeline.py
  • tests/test_web.py
  • web/src/App.tsx
  • web/src/__tests__/client.test.ts
  • web/src/__tests__/run.test.ts
  • web/src/__tests__/use-run.test.ts
  • web/src/mock/books.ts
  • web/src/mock/simulate.ts
  • web/src/run/client.ts
  • web/src/run/use-run.ts
  • web/src/styles/blueprint.css
  • web/src/types.ts
  • web/src/worker/metamend.worker.ts
  • web/src/worker/protocol.ts
 __________________________________________
< The GIL can't stop these rabbit threads. >
 ------------------------------------------
  \
   \   \
        \ /\
        ( )
      .( o ).
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…gers write them

Changes:
- Recognise the naming schemes measured in the wild (OceanofPDF, Anna's Archive, Z-Library, Library Genesis, PDFDrive, scene folders, sharing channels, ebook-tools, Standard Ebooks and dokumen.pub slugs, Springer, Calibre and LazyLibrarian's title-first form) and undo each one's encoding before splitting author from title
- Read a plain "A - B" author first as before, unless B reads more like a person, and carry the other reading along as FilenameFacts.alternate when both halves could be a person
- Try the other reading only when no source identifies the book as first read, and keep it only if a source then does; retry the head of a long title once when a subtitle's colon was dropped and the verdict is below HIGH
- Report a name that carries no title (a Gutenberg number, a bare ISBN, an ASIN) as such in the CLI and on the page instead of "no source answered"; print how a non-plain name was read
- Have the page ask the worker how every filename reads, so pending rows use the same parser as the verdicts, and show the reading and scheme in the detail view
- Record each catalogue once per row however many rounds ask it
- Document the audit and the decisions in docs/filenames.md

A file named the OceanofPDF way had no " - " in it, so the whole stem was read as an author with no title, nobody was asked, and the page said "no source answered", which was false. Renamed by hand the same three web sources reached HIGH. The user should not have to rename anything: half the tools out there write the title first and every site adds its own marks, and the parser now knows the shapes rather than one convention.

Notes:
- The safety model is untouched: matching.py, tags.py and the gain rules are unchanged. What changed is how the filename is read and how the catalogues are asked; the bar a source must clear is the same. A wrong reading cannot score, since a source would have to name a book whose title is the author's name and whose author is the title, twice over.
- Checked live against the two OceanofPDF pairs that started this (one HIGH, one LOW, the LOW being a summary edition the model rightly refuses), the same book named the Anna's Archive, Z-Library and Calibre ways (HIGH each), and replayed against the pristine sample with fixtures-v2 (Kobo, Google, Open Library) and fixtures-wide-raw (the web sources): every verdict, source list, gain and figure identical to before, no extra query asked.
- The head retry was narrowed after the replay showed it drawing a single volume out of Open Library for an omnibus, a strict prefix that the pre-existing prefix rule scores 0.95; it no longer cuts mid-phrase. That prefix rule itself is not new and is noted in docs/filenames.md.
Changes:
- Turn "Title__Subtitle" into "Title: Subtitle" in an OceanofPDF name, so the query stops at the main title
- Pin it with a parser test and note the shape in docs/filenames.md

A newly downloaded pair showed that the site keeps the colon as a double underscore and cuts the title at about forty characters. Read as a colon, the main title is queried directly and all three web sources answer at once instead of one source needing the head retry.
Changes:
- Remove the retry with the head of a long title; only the whole title is ever asked for, pinned by a test
- Multiply the page's pause between books by the rounds the last book took, so Apple's rate holds when a name is read both ways
- Read a PDFDrive name as a title alone; a dash inside it is a subtitle, never an author
- Turn "Last, First" round only when a bare surname stands before the comma, so a comma-separated pair of authors is left as it is
- Drop the title-only second reading of "Title by Author": it names nobody and could never be strong
- Skip the second reading in replay when the recording lacks its key, reporting it, instead of aborting the run

Review of the pull request reproduced the head retry proposing a different volume: a series name glued to a title drew the volume of that name out of two sources, a strict prefix of the filename title that the prefix rule scores 0.95, and its ISBN and series index were proposed at HIGH where the code before stopped at MED. Anything that makes the tool more willing to write is a change to the safety model, so the retry is gone; a glued subtitle now stays at what the full query earns.

Notes:
- Replayed the pristine sample against fixtures-v2 and fixtures-wide-raw: every verdict, source list, gain and figure identical to main
- Live on the four wild names: the Z-Library and double-underscore OceanofPDF forms HIGH with all three sources, the summary edition LOW, the OceanofPDF name with no colon trace MED on Apple alone, which is the honest result
@OffCrazyFreak
OffCrazyFreak force-pushed the feat/library-wild-filenames branch from d20c3b3 to 7fb32ea Compare September 16, 2026 19:11
@OffCrazyFreak
OffCrazyFreak merged commit 7bef0df into main Sep 16, 2026
3 checks passed
@OffCrazyFreak
OffCrazyFreak deleted the feat/library-wild-filenames branch September 16, 2026 19:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant