feat(library): Read filenames the way download sites and library managers write them - #31
Merged
Merged
Conversation
|
Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (22)
✨ Finishing Touches📝 Generate docstrings
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…gers write them Changes: - Recognise the naming schemes measured in the wild (OceanofPDF, Anna's Archive, Z-Library, Library Genesis, PDFDrive, scene folders, sharing channels, ebook-tools, Standard Ebooks and dokumen.pub slugs, Springer, Calibre and LazyLibrarian's title-first form) and undo each one's encoding before splitting author from title - Read a plain "A - B" author first as before, unless B reads more like a person, and carry the other reading along as FilenameFacts.alternate when both halves could be a person - Try the other reading only when no source identifies the book as first read, and keep it only if a source then does; retry the head of a long title once when a subtitle's colon was dropped and the verdict is below HIGH - Report a name that carries no title (a Gutenberg number, a bare ISBN, an ASIN) as such in the CLI and on the page instead of "no source answered"; print how a non-plain name was read - Have the page ask the worker how every filename reads, so pending rows use the same parser as the verdicts, and show the reading and scheme in the detail view - Record each catalogue once per row however many rounds ask it - Document the audit and the decisions in docs/filenames.md A file named the OceanofPDF way had no " - " in it, so the whole stem was read as an author with no title, nobody was asked, and the page said "no source answered", which was false. Renamed by hand the same three web sources reached HIGH. The user should not have to rename anything: half the tools out there write the title first and every site adds its own marks, and the parser now knows the shapes rather than one convention. Notes: - The safety model is untouched: matching.py, tags.py and the gain rules are unchanged. What changed is how the filename is read and how the catalogues are asked; the bar a source must clear is the same. A wrong reading cannot score, since a source would have to name a book whose title is the author's name and whose author is the title, twice over. - Checked live against the two OceanofPDF pairs that started this (one HIGH, one LOW, the LOW being a summary edition the model rightly refuses), the same book named the Anna's Archive, Z-Library and Calibre ways (HIGH each), and replayed against the pristine sample with fixtures-v2 (Kobo, Google, Open Library) and fixtures-wide-raw (the web sources): every verdict, source list, gain and figure identical to before, no extra query asked. - The head retry was narrowed after the replay showed it drawing a single volume out of Open Library for an omnibus, a strict prefix that the pre-existing prefix rule scores 0.95; it no longer cuts mid-phrase. That prefix rule itself is not new and is noted in docs/filenames.md.
Changes: - Turn "Title__Subtitle" into "Title: Subtitle" in an OceanofPDF name, so the query stops at the main title - Pin it with a parser test and note the shape in docs/filenames.md A newly downloaded pair showed that the site keeps the colon as a double underscore and cuts the title at about forty characters. Read as a colon, the main title is queried directly and all three web sources answer at once instead of one source needing the head retry.
Changes: - Remove the retry with the head of a long title; only the whole title is ever asked for, pinned by a test - Multiply the page's pause between books by the rounds the last book took, so Apple's rate holds when a name is read both ways - Read a PDFDrive name as a title alone; a dash inside it is a subtitle, never an author - Turn "Last, First" round only when a bare surname stands before the comma, so a comma-separated pair of authors is left as it is - Drop the title-only second reading of "Title by Author": it names nobody and could never be strong - Skip the second reading in replay when the recording lacks its key, reporting it, instead of aborting the run Review of the pull request reproduced the head retry proposing a different volume: a series name glued to a title drew the volume of that name out of two sources, a strict prefix of the filename title that the prefix rule scores 0.95, and its ISBN and series index were proposed at HIGH where the code before stopped at MED. Anything that makes the tool more willing to write is a change to the safety model, so the retry is gone; a glued subtitle now stays at what the full query earns. Notes: - Replayed the pristine sample against fixtures-v2 and fixtures-wide-raw: every verdict, source list, gain and figure identical to main - Live on the four wild names: the Z-Library and double-underscore OceanofPDF forms HIGH with all three sources, the summary edition LOW, the OceanofPDF name with no colon trace MED on Apple alone, which is the honest result
OffCrazyFreak
force-pushed
the
feat/library-wild-filenames
branch
from
September 16, 2026 19:11
d20c3b3 to
7fb32ea
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Changes:
A - Bauthor first as before, unless B reads more like a person, and carry the other reading along asFilenameFacts.alternatewhen both halves could be a persondocs/filenames.mdA file named the OceanofPDF way had no
-in it, so the whole stem was read as an author with no title, nobody was asked, and the page said "no source answered", which was false. Renamed by hand the same three web sources reached HIGH. Nobody should have to rename anything: half the tools out there write the title first and every site adds its own marks, so the parser now knows the shapes rather than one convention. The audit behind the rules (Calibre, Calibre-Web, LazyLibrarian and Readarr defaults read from their source, Anna's Archive's filename builder, Z-Library, libgen, OceanofPDF and PDFDrive names quoted from public issues, Gutenberg, Standard Ebooks and Internet Archive checked live) is in the new docs page.Safety model:
matching.py,tags.pyand the gain rules are untouched. What changed is how the filename is read and how the catalogues are asked; the bar a source must clear is the same. A wrong reading cannot score, since a source would have to name a book whose title is the author's name and whose author is the title, twice over.Checked against: the two OceanofPDF pairs that started this, live with the web sources (one MED on Apple alone since the site left no trace of the colon, one LOW, the LOW being a publisher's summary edition the model rightly refuses), plus the newer OceanofPDF pair whose double underscore keeps the colon (HIGH with all three sources); the same book named the Anna's Archive, Z-Library and Calibre ways (HIGH each); the pristine sample replayed against
fixtures-v2(Kobo, Google, Open Library) andfixtures-wide-raw(web sources) with every verdict, source list, gain and figure identical tomainand no extra query asked. A retry with the head of a glued title was built, then removed after review reproduced it proposing a different volume of a series (a strict prefix the pre-existing prefix rule scores 0.95); only the whole title is ever asked for now, and docs/filenames.md records why.Checks:
ruff check,ruff format --check,pytest(395 passed);pnpm typecheck,pnpm format:check,pnpm test(44 passed),pnpm build,node scripts/pyodide-smoke.mjs(401 passed under Pyodide). The page was run end to end in headless Chromium with the real files through the file input at desktop width; the phone width was covered by the sweep in #30.Notes:
web/src/App.tsxon the same row-title line as fix(web): Keep the results list readable below 1024px #30, so whichever merges second needs a one-line rebase.