Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,6 +111,8 @@ Books are expected to be named `Author - Title - Subtitle.epub`, optionally with
└── Paulo Coelho - The Alchemist.epub
```

Files that arrived from elsewhere are read the way their source wrote them, so nothing has to be renamed first: Calibre's `Title - Author`, `_OceanofPDF.com_Title_-_Author`, Anna's Archive's `Title -- Author -- ... -- Anna’s Archive`, Z-Library's `Title (Author) (z-lib.org)`, Library Genesis, PDFDrive, Standard Ebooks slugs and the rest of what [docs/filenames.md](docs/filenames.md) lists. A plain `A - B` that could be read either way is read author first, and if no catalogue identifies the book that way the other reading is tried, so the catalogues settle the order rather than a guess. A name that carries no title at all (`pg1342.epub`, a bare ISBN) is reported as such instead of "no source answered".

Naive string similarity fails on real book titles, so two cases are handled specially:

- **Prefix containment is legitimate.** "Digital Minimalism" vs "Digital Minimalism: Choosing a Focused Life in a Noisy World" is the same book, main title plus subtitle. Scored 0.95.
Expand Down
1 change: 1 addition & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,4 +3,5 @@
Decisions and traps that the code cannot show. Nothing here is live state; anything with a date is a snapshot from that day.

- [sources.md](sources.md): which metadata sources were measured, what each is like, and why the web build uses a different set from the desktop tool.
- [filenames.md](filenames.md): how stores, library managers and download sites actually name files, which shapes the parser recognises, and why the pipeline may read a name the other way round.
- [web.md](web.md): how the browser version works, why it is hosted the way it is, what each browser can do, and how to check a change.
71 changes: 71 additions & 0 deletions docs/filenames.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
# Filenames in the wild

What ebook files are actually called when they arrive from a store, a library manager or a download site, measured on 2026-09-16, and how `library.parse_filename` reads each shape. The filename stays the ground truth; this page is about reading it correctly when another tool wrote it. Nothing here changes what a source must prove before a field is written.

## Why

A file named `_OceanofPDF.com_The_Quiet_Orchard_-_Mara_Voss.epub` has no ` - ` in it, so the old parser read the whole stem as an author with no title, asked nobody, and the page reported "no source answered". That report was false: no source was asked. Renamed to `Author - Title` the same three web sources reached HIGH. The user should not have to rename anything, so the parser now recognises the shapes below.

## The shapes, with their sources

Author first:

- This tool's own convention, `Author - Title` and `Author - Series - 02 - Title` (README).
- Readarr, default `{Author Name} - {Book Title}{ (PartNumber)}` (`NamingConfig.cs` in the Readarr repository).
- Library Genesis mirrors, `Author - Title (2020, Publisher) - libgen.li`, also dotted `Author.-.Title.2020.Publisher.-.libgen.lc.1`, and `_ ` standing for `: ` inside a title (filenames quoted in GitHub issues; a third-party parser documents the `_ ` convention).
- Scene release folders, `Author.Name.-.Title.Of.Book.2021.RETAIL.EPUB.eBook-GROUP` (shelfmark pull request 19).
- IRC sharing channels, `Author - Title [RSC] (retail).epub` (the Shadow Libraries IRC guide).
- ebook-tools output, `Author - [Series #1] - Title (2008) [ISBN].ext` (its README).
- Standard Ebooks, `author-name_title-of-book.epub`, `_advanced.epub`, `.kepub.epub` (checked live on standardebooks.org).

Title first:

- Calibre "Save to disk" and "Send to device", defaults `{author_sort}/{title}/{title} - {authors}` and `{author_sort}/{title} - {authors}`, authors joined by ` & ` (`save_to_disk.py`). Calibre's own guess-from-filename regex assumes the same, `(?P<title>.+) - (?P<author>[^_]+)` (`meta.py`, the manual).
- Calibre-Web downloads, `Title - FirstAuthor.ext` (`cps/helper.py`).
- LazyLibrarian, default `$Title - $Author` (`configdefs.py`).
- Anna's Archive, `Title -- Author -- Edition, Year -- Publisher -- ISBN -- md5 -- Anna’s Archive.ext`, every field at most 60 characters, the whole at most 150, and every `.` in the name turned into `_` so `Mara T. Voss` arrives as `Mara T_ Voss` (`allthethings/page/views.py`). Empty fields are dropped, so the second field is not always the author.
- Z-Library over the years, `Title (Author) (z-lib.org)`, `Title by Author (z-lib.org)`, `Title (Author)` followed by an em dash and `_Publisher_Language_ISBN (Z-Library)`, `Title (Last, First etc.) (z-library.sk, 1lib.sk, z-lib.sk)`. A colon in the title is dropped and leaves two spaces behind, which is how the subtitle boundary is recovered (filenames quoted in GitHub issues).
- OceanofPDF, `_OceanofPDF.com_Title_-_Author.ext`, underscores for spaces, a colon left as a double underscore (`Title__Subtitle`), the title cut at about 40 characters, dots dropped from initials, hyphens inside words kept (`Domain-Driven`) (three independent renaming scripts on GitHub and the files that started this).
- Renaming tools, `Title by Author.ext` (ebook-rename's README).

Title only:

- PDFDrive, `Title ( PDFDrive ).pdf` and `Title ( PDFDrive.com ).pdf`, spaces inside the brackets, `_ ` for `: `, a dash inside is a subtitle and never an author (filenames quoted on GitHub).
- Kindle "Download & transfer via USB", the title alone; Kindle for PC, `ASIN_EBOK.azw` (DeDRM issues).
- FanFicFare, default `${title}-${siteabbrev}_${storyId}` (`defaults.ini`).
- dokumen.pub, vdoc.pub, epdf.pub slugs, `the-title-of-the-book-9780465050659-9780465003945-2013024417`, `-1nbsped-` for "1st ed.", `-4u9bqm2ndpq0` record ids, `epdf-pub-...-pdf` (their page URLs).
- Springer, `2020_Book_IntroductionToScientificProgra.pdf`, CamelCase and cut at 30 characters (a Springer link in free-programming-books).
- Publishers and Humble Bundle, `Title_With_Underscores.pdf` (widely seen, not separately sourced).
- Scribd, `Document Title | PDF | Topic`.

No title at all:

- Project Gutenberg, `pg1342.epub`, `pg1342-images.epub`, `pg1342-images-3.epub` (checked live).
- Internet Archive, `atomichabitseasy0000clea_lcp.epub` (checked live).
- A bare ISBN, `978-1-4842-8853-5.pdf` (Springer's DOI links) or `9780465050659.epub`.

Other observations that shaped the rules: a spaced en dash, a spaced em dash, ` -- ` and ` _ ` all appear as the separator between the same two halves; browsers append ` (1)` to a second download; Kobo files carry `.kepub.epub`; Kavita, by contrast, reads the OPF first and uses the filename only as a fallback, the opposite of this tool's premise.

## What the parser does

1. Strips the site's own marks (`_OceanofPDF.com_`, `(z-lib.org)`, `( PDFDrive )`, `- libgen.li`, `-- Anna’s Archive`, `(retail)`, `(v5.0)`, `(epub)`), a trailing `(Year)` or `(Year, Publisher)`, a bracketed ISBN, a duplicate-download counter and `.kepub`.
2. Undoes the site's encoding: underscores or dots for spaces, `_ ` for `: `, `.-.` for ` - `, Anna's Archive's `_` for `.`, Z-Library's double space and OceanofPDF's double underscore for `: `, slugs back into words, Springer's CamelCase into words, `Last, First` into `First Last` (never `Smith, Jr.`, never `Mara Voss, Ann Person`).
3. Recognises the order when the scheme fixes it. A plain `A - B` is read author first, as the README asks, unless B reads more like a person than A (two or three capitalised words, an initial, no digits, no colon). When the name could be read either way and the other half could be a person at all, the other reading travels along as `FilenameFacts.alternate`.
4. Series shapes from other tools are read too: `[Series #2]` as its own segment and `Title (Series Book 2)`.
5. A name that carries no title (`pg1342`, an ISBN) is reported as such, in the CLI line and on the page, instead of "no source answered".

`FilenameFacts.scheme` names what was recognised, so the CLI prints `read as Author / Title (scheme name)` and the page can say the same.

## What the pipeline does with the other reading

Only when no source identifies the book as first read does `enrich.propose` ask the catalogues about the alternate reading, and it keeps that reading only if a source then identifies the book. A wrong reading cannot score: a source would have to name a book whose title is the author's name and whose author is the title, and two of them would have to agree. The bar for writing is unchanged; the cost is one extra round of queries for a book that was going to be LOW anyway.

A retry with the head of a long title (asking for `Quiet Orchard` when the name says `Quiet Orchard The Year Of Pruning`) was built, measured and removed. It did rescue a name whose colon had been dropped, because Open Library answers nothing for the glued form. It also manufactured a HIGH for the wrong book: a series name glued to a title (`The Dark Tower The Waste Lands`) drew the volume called `The Dark Tower` out of two sources, and a source title that is a strict prefix of the filename title scores 0.95 under the prefix rule, so volume VII's ISBN and series index were proposed for volume III. Anything that makes the tool more willing to write is a change to the safety model, so the retry is gone; a glued subtitle now stays at whatever the full query earns, usually MED, which is honest. The prefix rule itself predates this page and still applies to any source that answers with a strict prefix of the filename title.

The page paces between books, not inside one, so `web.pause_after` multiplies the wait by the rounds the last book took (`enrich.last_rounds`); Apple's twenty calls a minute hold even when every book is read both ways. In replay mode a fixture set recorded before names had two readings has no key for the second one; that round reports the missing recording and is skipped, while a missing key in the first round stays loud as before.

## Checked against

- The two OceanofPDF pairs that started this, live with the three web sources: one HIGH (Apple and Open Library agreeing), one LOW. The LOW is correct: that file is a publisher's summary edition of a well-known book, its own metadata names the summary publisher as the first author, and the author score of 0.26 against the original is the safety model refusing to dress a summary up as the book it summarises.
- The same book renamed the Anna's Archive way, the Z-Library way and the Calibre way (`Title - Author`): HIGH each time, with the reading printed.
- The pristine ten-book sample replayed against `fixtures-v2` (Kobo, Google, Open Library) and `fixtures-wide-raw` (the web sources): every verdict, source list, gain and figure identical to the code before this change, and no extra query asked.
14 changes: 13 additions & 1 deletion src/ebook_metamend/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,13 @@ def _write_json(path: str, payload) -> None:

def _report(index: int, total: int, book: Book, proposal: Proposal | None) -> None:
if proposal is None:
print(f'[{index}/{total}] {book.stem[:60]:<60} no source answered')
facts = book.facts()
why = 'no source answered'
if not facts.title:
why = 'filename names no title, expected "Author - Title"'
if facts.scheme:
why += f' ({facts.scheme} name)'
print(f'[{index}/{total}] {book.stem[:60]:<60} {why}')
return

sources = ','.join(proposal.sources)
Expand All @@ -36,6 +42,12 @@ def _report(index: int, total: int, book: Book, proposal: Proposal | None) -> No
f'fn={proposal.fn_score:.2f} au={proposal.au_score:.2f} '
f'src:{sources:<20} gains:{gains}'
)
if proposal.facts and (proposal.facts.scheme or proposal.facts.author != book.facts().author):
# The name was not the plain "Author - Title": say how it was read.
how = f' ({proposal.facts.scheme} name)' if proposal.facts.scheme else ''
print(
f' read as {proposal.facts.author or "(no author)"} / {proposal.facts.title}{how}'
)
if proposal.gains.get('title'):
was = (proposal.current.get('title') or '')[:40]
print(f" title {was} -> {proposal.gains['title'][:60]}")
Expand Down
47 changes: 38 additions & 9 deletions src/ebook_metamend/enrich.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@

from __future__ import annotations

import functools
import re
import sys
import time
Expand All @@ -14,7 +15,7 @@
from typing import Any

from . import calibre, matching, tags
from .library import Book, books
from .library import Book, FilenameFacts, books
from .sources import SOURCES, Pacer, Source, cache
from .sources.errors import SourceError, SourceUnavailable
from .writers import epub, pdf
Expand Down Expand Up @@ -56,6 +57,8 @@ class Proposal:
writes: list[tuple[str, bool, str]] = field(default_factory=list, repr=False)
#: Every source's score, including the ones that earned no say. Reporting only.
scores: list[matching.SourceScore] = field(default_factory=list, repr=False)
#: The reading of the filename the verdict was scored against. Reporting only.
facts: FilenameFacts | None = field(default=None, repr=False)

def to_dict(self) -> dict[str, Any]:
"""The serialised form. Explicit rather than ``asdict`` so that adding a
Expand All @@ -81,6 +84,10 @@ def to_dict(self) -> dict[str, Any]:
_pacer = Pacer()
#: Consecutive transport failures per source, for the shelving rule above.
_failures: dict[str, int] = {}
#: How many rounds of queries the last propose() made: two when a name was
#: read both ways. A caller pacing itself between books waits that many times
#: longer, so a source's rate holds even when a book cost two rounds.
last_rounds = 1


def reset_run_state() -> None:
Expand All @@ -90,10 +97,11 @@ def reset_run_state() -> None:
once is never retried, and back-off from a previous run still applies. Fine
for a single CLI invocation, wrong for anything longer lived.
"""
global _pacer
global _pacer, last_rounds
unavailable_sources.clear()
_failures.clear()
_pacer = Pacer()
last_rounds = 1


def query_sources(
Expand Down Expand Up @@ -401,18 +409,38 @@ def propose(
sleep inside Python and waits between books in JavaScript instead.
"""
facts = book.facts()
# A stem with no " - " parses as all author and no title, so every title
# score would be 0.0 against an empty string and any answer at all would
# look equally (un)related. There is nothing to score against, so do not ask.
# A stem that names no title (a Gutenberg number, a bare ISBN) leaves nothing
# to score against: every title score would be 0.0 against an empty string
# and any answer at all would look equally (un)related. So do not ask.
if not facts.query:
return None
answers = query_sources(
facts.query, facts.author, sources=sources, pause=pause, on_answer=on_answer
)
global last_rounds
last_rounds = 1
ask = functools.partial(query_sources, sources=sources, pause=pause, on_answer=on_answer)
answers = ask(facts.query, facts.author)
scores, conf = score(answers, facts) if answers else ([], 'LOW')
if facts.alternate is not None and not any(s.strong for s in scores):
# Nothing identified the book as first read, and the name could be read
# the other way round: half the tools out there write the title first.
# The catalogues settle it. A wrong reading cannot score: a source would
# have to name a book whose title is the author's name and whose author
# is the title, twice over, so the bar for writing is unchanged.
other = facts.alternate
try:
other_answers = ask(other.query, other.author)
except cache.MissingFixture as exc:
# A set recorded before names had two readings has no key for the
# second one. The first round stays loud about a missing key; this
# round is opportunistic, so it is reported and skipped.
print(f'replay has no recording for the other reading: {exc}', file=sys.stderr)
other_answers = {}
last_rounds = 2
other_scores, other_conf = score(other_answers, other) if other_answers else ([], 'LOW')
if any(s.strong for s in other_scores):
facts, answers, scores, conf = other, other_answers, other_scores, other_conf
if not answers:
return None

scores, conf = score(answers, facts)
trusted = trusted_names(scores)
title_score, author_score = reported_scores(scores, trusted)
surviving = {name: answers[name] for name in trusted}
Expand All @@ -437,6 +465,7 @@ def propose(
current=current or {},
unreadable=unreadable,
scores=scores,
facts=facts,
)


Expand Down
Loading
Loading