Skip to content

Latest commit

 

History

History
1474 lines (1102 loc) · 174 KB

File metadata and controls

1474 lines (1102 loc) · 174 KB

CalibreQuarry Patch Notes

3.54.0 (2026-09-26)

The multi-author seed fix and the stamp_pdf parse-crash recovery

  • _stamps_from_embedded cuts the author line at the [ bracket BEFORE splitting on " & " (the 2026-09-26 HFT 2nd-ed EPUB): calibre renders N authors as A & B & C & D [SortA & SortB & SortC & SortD], the bracket wrapping the WHOLE list, and the seeder split the display line before cutting it, so a correct 4-author stamp seeded 7 authors: the four display names plus sort-form fragments riding along as phantom authors in the manifest seeds. The OPF and the stamp_epub verifier were both correct in the live case, and the hand-corrected signed manifest stands. The fix is the display-segment shape stamp_epub's _display_author already compares for single authors: the single-author Name [Sort, Form] form seeds exactly 1, and the Unknown-placeholder drop still works in both its bare and bracketed forms.
  • stamp_pdf recovers the exiftool parse-crash class mechanically, exactly once (the 2026-09-26 C++23 STL Cookbook PDF): when the exiftool write exits nonzero WITH a Perl-space parse error in its output (today's shape: Can't find Root object off a malformed catalog Names array, on a file qpdf --check passes), the script attempts one qpdf --replace-input rebuild whose page count must be preserved (qpdf missing, an unreadable side, or a changed count restores the backup copy already made and no retry runs), then retries the stamp ONCE, re-reading the raw calibre author_sort property after any rebuild because the rebuild rewrote the XMP packet. A retry that still fails lands in the existing STAMP_FAILED path, and the stubborn-XMP class (write reports success, read-back disagrees) is untouched and keeps failing immediately. The live file was rebuilt by hand first, 496 pages preserved, the retry verified.
  • Docs and tests: both roadmap boxes are ticked with the fix notes; CLAUDE.md gains the 3.54.0 contract-notes section; the phase-1 library skill names the one-recovery exception in its stamp_pdf prose. New tests (9): the seeder's 4-author bracket cut, single-author suffix cut, and Unknown drop; and the parse-crash battery (rebuild verified and retried once, retry-still-failing STAMP_FAILED with no second rebuild, a page-count change restoring the backup byte-for-byte and failing without retry, missing qpdf and non-Perl failures never recovering, and the _is_parse_crash shape table). 681 tests, one skip on ebook-meta-less environments.

3.53.0 (2026-09-22)

stamp_epub: the pre-stamp ritual gets its EPUB tool (mandatory for ALL filetypes)

  • scripts/stamp_epub.py ships the EPUB sibling of stamp_pdf.py (the open roadmap box; the 2026-09-22 fiction wave ran the EPUB half on a one-off batch driver, 46 EPUBs stamped and three verifier iterations to learn the identifier spellings): one ebook-meta invocation writes the fixed field set (--title/--authors/--publisher/--isbn); Calibre's own writer rewrites the OPF in place and REPLACES any existing isbn identifier, which IS the correction mechanism for wrong embedded ISBNs, so dc: metadata is never hand-edited. Dry-run by default; --apply requires a --backup-dir refused when it resolves inside any target file's directory; repeated --author flags join with " & " (ebook-meta's own separator; cquarry --set-authors splits on ; instead, a CLI-layer difference only). The tool-boundary checksum gate rides stamp_pdf 3.52.0's shape: an --isbn failing cquarry's isbn_check_digit_is_valid is refused (exit 2) in dry-run and apply alike, before anything writes.
  • The verify contract is the EPUB-specific one the live exercise cost three iterations to learn. Title/authors/publisher verify from the ebook-meta read-back, the author compared as the display segment before the [ bracket (live files render Name [Sort, Form]). ebook-meta's read-back never displays ISBN, so the ISBN verifies by reading the OPF directly (zipfile -> the container's rootfile entry -> dc:identifier elements) under VALUE EQUALITY: each identifier's value AND id attribute is normalized by dropping an optional leading urn:/isbn: prefix and then every non-alphanumeric character, and the stamp verifies when any candidate equals it. Bare isbn:978... values, opf:scheme="ISBN", namespaced ns2:scheme="ISBN", dashed values, and the ISBN carried in the element id all verify; a file that ALREADY carried the right ISBN makes --isbn an honest no-op (the writer never destroys a matching identifier, it may add its own canonical one alongside) and verifies, never fails; when no --isbn was requested the ISBN check is skipped entirely.
  • A failed stamp STOPS the list (STAMP_FAILED, exit 1, never re-fought): the per-file report names the mismatched fields, every not-yet-stamped remainder is named, and the --json FILE machine report records every target exactly once (dry_run / stamped_verified / stamp_failed / left_unstamped), the bijection record the batch driver used to assert by hand. A target named twice is refused at the boundary (exit 2): one file, one stamp.
  • Docs and tests: the roadmap box is ticked with the live notes; spec § 5 gains the stamp_epub row and the write-capable closing list names it; CLAUDE.md gains the 3.53.0 contract-notes section. New tests (29): the verify contract (bracket display segment, value equality over every live shape, wrong-ISBN failure, no-ISBN skip), the normalizer table, the OPF reader (value + id attribute as candidates, the container rootfile honored over the entry name), the guard set (checksum refusals dry and apply, backup-dir refusals, dry-run invokes nothing, missing/non-EPUB/duplicate targets), the stop-and-report flow, and a real-ebook-meta seam class (skips on ebook-meta-less machines) pinning the wrong-ISBN replace, the four already-right no-op shapes, and the clean-file gain (672 tests, one skip). The phase-1 library skill's EPUB half now drives this tool instead of raw ebook-meta; the skill update ships in the same release. Erratum on the 3.52.0 entry below: its closing NOTE said the phase-1 skill's duplicate-screen prose was left for Brandon; it was in fact rewritten by ZCode that same evening, so nothing is owed there.

3.52.0 (2026-09-21)

The duplicate-screen semantics fix: the 2026-09-21 waves' false refusals and the silent masked duplicate, plus the ISBN and pubdate guards

  • screen_duplicate classifies title pairs instead of only comparing scrubbed equality (roadmap 2026-09-16 and 2026-09-19 boxes, driven by the 2026-09-21 mixed wave): classify_titles returns a verdict per pair, and the four failing shapes now each get the right one. Differing REAL subtitles over one shared base are distinct ("Introduction to Computer Organization: ARM" vs "An Under-the-Hood Look at Hardware and x86-64 Assembly", two books the scrub collapsed onto each other and refused twice in one day). Differing DECLARED volume annotations (a token word plus an arabic or valid roman numeral on both sides) are volume_sibling: the Moral Letters vols, the roman "Vol I" vs "Vol II" pair that was a false duplicate until now, and "Monster Vault" vs "Monster Vault 2" (separated by a trailing-ordinal check over the post-colon subtitle). Containment at a colon boundary (one full title equal to the other's pre-colon base or post-colon subtitle) is related: the Arcana base-vs-companion shape within the batch, and the masked DUPLICATE the old equality silently passed against the library ("Mothership: Wages of Sin" vs the bare "Wages of Sin" download, caught on 09-21 only by a manual collision sweep). Related candidates surface in the report (batch_related, related_hits), count toward exit 1, and are recorded in the phase-1 manifest under checks.related_works; volume sets land as batch_volumes / checks.volume_siblings, informational only. Neither class ever sets a verdict: the screen advises, the reviewer judges. The regression floor is untouched: normalize_title keeps the Capital scrub exactly, the roman-only annotation still matches the full-subtitle edition, and the 3.25.0 arabic-signature gate stays byte-for-byte (a "Book 8" next to the bare series title remains a silent gap-fill, so the 19-candidate Wandering Inn flood stays impossible).
  • stamp_pdf refuses an --isbn failing its check digit at the tool boundary (roadmap 2026-09-19 box): _check_isbn runs cquarry's isbn_check_digit_is_valid (10- or 13-form, separators tolerated) and exits 2 before anything stamps, in dry-run and apply alike, so the guard lives where the write happens instead of in every caller's batch script. Correction found while pinning the test: the Hewitt "valid form" 1-59863-503-5 named in the roadmap box ALSO fails its checksum and is refused like every other bad number; that ISBN needs re-sourcing, not re-typing.
  • The phase-2 metadata download writes pubdate at date-only precision (roadmap 2026-09-16 box, recurred 2026-09-21): _apply_opf truncates every OPF date to YYYY-MM-DD before set_pubdate, because downloaded metadata carries no trustworthy time-of-day and the downloader was landing 2010-05-25 01:49:00.233821+00:00-shaped values (42 of 47 books in one batch shared the same minute-band, all normalized by hand in phase 3). Bare years and exotic forms ride through unchanged, failing set_pubdate exactly as before; reconcile is untouched.
  • Docs and tests: the four roadmap boxes are ticked with fix notes; spec § 5's screen_duplicate and stamp_pdf rows name the new behavior; CLAUDE.md gains the 3.52.0 contract-notes section. New tests: the classifier case table (all three 2026-09-21 cases plus the Moral Letters, Arcana, and roman-volume shapes), the masked-duplicate and Monster Vault library fixtures, within-batch cross-reference classes, the seam's advisory notes, phase-1's advisory recording without refusal, the ISBN guard's refusals and passes, and the pubdate truncation end to end through _apply_opf (643 tests, one skip on ebook-meta-less machines). NOTE FOR THE LIBRARY SKILLS: the phase-1 skill's duplicate-screen prose describes the old normalizer behavior and needs one line updated (related candidates and volume sets now surface in the report and manifest instead of nothing) — left for Brandon, outside this repo.

3.51.0 (2026-09-21)

The bindery structured-fix records adoption: lossy consent reads data, with a version gate on the PATH binary

  • _mirror_lossy classes repairs from bindery's structured fix records, not summary substrings (the 2026-09-18 box on bindery-cli's roadmap, fulfilled by its v0.45.0): every per-book record in bindery's phase-1 report now carries the fixes dict plus ncx_uid_synced/watermark_refusals as data, so the lossy/structural split is judged by fixes.get(marker) over the same five keys bindery's own gate treats as lossy strips. The rendered summary stays in the sealed record as the display line the lossy_consent detail quotes. The practical gain: the class no longer depends on bindery's summary vocabulary staying stable; a summary that names no marker string still classes correctly from its data.
  • A PATH bindery below 0.45.0 is a hard error, never a silent downgrade: _bindery_phase1 probes bindery --version and refuses below _BINDERY_MIN_VERSION (0.45.0) with the upgrade command named, because the data-driven classing against an older report would class every strip as structural (the vacuous-consent hole again). The probe result is also why this is a minor release: the consent gate's contract now includes a versioned report shape.
  • Landed from the preservation branch: the work was written in a working tree a parallel release lane had to sweep aside mid-run on 2026-09-21; it survived on bindery-version-gate-wip (snapshot 9135493) and was cherry-picked onto main after 3.50.0 shipped. That branch's WIP message warned test_dry_phase1_emits_lossy_consent_decisions was broken; the warning was stale (the pre-fix state from the collision window), the snapshot carries all five migrated fixtures, and the full suite runs 622 green against the branch tree, verified in an isolated checkout before landing. The branch is deleted; the content lives on main.

3.50.0 (2026-09-21)

The calibre-touched PDF fix: stamp_pdf erases the stale author_sort that failed verify and would have poisoned imports

  • stamp_pdf clears a pre-existing XMP calibre:author_sort (roadmap 2026-09-21, found by the Effective C wave): any PDF Calibre has ever produced or touched carries that extension, ebook-meta renders the author as Display [Sort] from it, and the read-back disagreed on author even when every written field took (De Anima, How Software Works in that batch). The failure was never merely cosmetic: calibre's create_book_entry honors an embedded mi.author_sort verbatim instead of recomputing from the stamped authors, so the stale sort would have landed in the library on import; the verify-side tolerance idea was discarded for exactly that reason. exiftool cannot write the calibre namespace (not in its tables, and a user-defined -config table never associates with the packet parse), so the erase rides calibre's own writer: when the stamp sets authors and exiftool -s3 -Author_sort shows the property, a follow-up value-preserving ebook-meta pass with --author-sort "" (empty is null, so calibre writes no sort of its own) regenerates the XMP packet without ANY calibre-namespaced element, stale sorts, timestamps, and ratings included; the docinfo Keywords that carry the ISBN survive the rewrite, verified end to end on a calibre-touched repro (red before the fix, green after, clean XMP read). The erase re-carries exactly the stamped values, runs after the exiftool write so a failed write still leaves the file untouched, and its own failure is a STAMP_FAILED, not a blind continue. _verify stays strict: the file is made clean, the comparison is not loosened.
  • Docs and tests: tests pin the erase argv (stamped values re-carried, the null sort, the file path last), the skip when no sort is present, and the erase-failure path; the two exiftool call-count assertions now filter for the write invocation since detection shells out too (619 tests). README gains a "How this compares" section, and the mixed-wave roadmap findings were recorded without code changes (screen_duplicate title-prefix containment, the stamp_pdf ISBN checksum idea, reviewer verdict-flip propagation, an EVERY_BOOK_COVER allowlist) and stay open. A stopped agent's bindery version-gate work was preserved unreleased on the bindery-version-gate-wip branch.

3.49.0 (2026-09-18)

The roadmap findings batch: honest DJVU coverage, embedded-metadata stamp seeds, an in-batch cc6 check, and a lossy consent that only gates real strips

  • audit_drm's N/A verdicts reach the runner (the DJVU box, observed 2026-09-17): the directory- and library-scan loops continued past N/A verdicts (DJVU has no DRM scheme in practice) BEFORE the CSV append, so the manifest recorded checks.drm = "unscanned" for a format that was judged. N/A rows now land in the CSV (still out of the counters and the summary), the manifest says N/A, and N/A remains no quarantine reason. The instrument test is re-pinned from the old "unscanned" contract to the N/A one, and the phase-1 skill's coverage note matches the tool again.
  • Phase-1 stamps seed from the file's embedded metadata (the seeding box): _stamps_from_embedded reads ebook-meta's display output and merges title/authors/publisher/pubdate/language per-field over the filename parse; ISBN is deliberately never seeded (embedded identifiers are misidentification-prone). Guards that keep a bad seed from failing phase 2 or beating the decided conventions: a pubdate that does not parse as ISO is dropped (cquarry's add_book raises on a bare year), a read whose stderr carries calibre's traceback returns nothing (calibre exits 0 on unparseable files and prints its own reversed-order filename guess), Unknown placeholders are dropped, a missing ebook-meta binary degrades to filename-only with a warning, and a per-file read failure is recorded in the entry's repairs. The skill's step 5c is rewritten around the new seeding; the correction pass is now usually a verify.
  • The phase-2 #source stamp is verified in-batch (the cc6 boxes): inside the one batch(), a fresh book's first set_custom_column must report changed, or the import rolls back and the verb exits 1 with the library unwritten. The accompanying audit is an erratum on the observations: both signed manifests carry ZERO imported_ids (phase 2 saves the resume record immediately after the import), so neither batch went through the verb; the DB actually shows the reviewed provenance landed correctly on 77 of 79 books, with the two mismatches being the manifest-None files (one picked up Z-Lib, one Anna's Archive) and no all-Anna's-Archive state present. The out-of-verb door is outside the verb's reach; the effect-flag check is the in-verb half (a true read-back would need a cquarry custom-column reader and is recorded as a future option).
  • lossy_consent fires only on real lossy strips (the consent box, observed 2026-09-17 and 09-19): _mirror_lossy still records every gate-accepted bindery repair in the sealed record, but each repair now names its class ("lossy": true/false, judged by _LOSSY_REPAIR_MARKERS, the five fix keys bindery's own gate treats as lossy: stripped_pagination, stripped_broken_tags, stripped_watermarks, dropped_marker, stub_docs_dropped), and only lossy-marker repairs set the file's flagged. Structural-only files stop firing vacuous consent decisions, and the decision's detail carries the lossy summaries. Phase 2's consent re-drive flips and refuses over every recorded repair, lossy or structural, because bindery's re-drive applies all of them. The cleaner long-term home (a structured fixes dict in bindery's phase-1 report) is recorded on bindery-cli's roadmap.
  • Docs and tests: CLAUDE.md gains the 3.49.0 contract-notes section; the roadmap's five findings boxes are ticked with the audit results; the phase-1/phase-3 library skills and bindery-cli's roadmap are updated in the same release (the sweep rule). New tests: the ebook-meta seam parse and phase-1 seeding merge/fallback paths, the in-batch stamp rollback, the structural-record flip and refusal on consent re-drive, the lossy class mirrors, and the real audit_drm CSV round-trip (615 tests).

3.48.0 (2026-09-17)

The wave-2 refactor and consent batch: one dest-list source, one backup helper, and the lossy consent moves into the manifest

  • The parallel write-dest lists have one source (roadmap "new work noticed", L2.14): src/cquarry_cli/dests.py now holds SINGLE_BOOK_DESTS (writeops), set mode's target sources and --batch-* verbs (setwrite), and the restrict-refusal aggregate WRITE_FLAG_DESTS (restrict). The three consumers import their slices, and tests/test_dests.py pins every member against build_parser(), so a dest that exists only in a list (or only in the parser) cannot rot.
  • One backup helper: backups.make_backup replaces the triplicated _backup_db/_make_backup (run phase 2, the integrate verbs, set mode's --apply). The recorded error-mapping decision: a --backup-dir inside the library is a usage problem (exit 2) at every door, so the shared helper raises the stdlib-neutral ValueError and each dispatcher keeps its own usage path; unwritable destinations and sqlite failures raise ValueError too (set mode already wrapped them into its usage path; run/integrate previously propagated a traceback). Same timestamped sqlite-API backup, same on-disk shape.
  • bindery's manual_watermark_repair decisions mirror into the manifest's decisions_needed as manual_repair entries (_mirror_bindery_decisions, the sibling of 3.43.0's _mirror_lossy): books bindery refuses to auto-strip used to vanish from the durable record entirely.
  • The lossy-pending double-manifest wrinkle is closed: a dry phase 1 now emits a lossy_consent decision per lossy-flagged file, and the reviewer resolves it in the manifest (set the decision's "resolution" to "apply", re-sign). Phase 2 then drives bindery run phase1 --apply-lossy itself before the import batch, flips the lossy records to applied, and consumes the decisions, so consent no longer requires the phase-1 re-run that minted a second manifest and orphaned the first. Consent is all-or-nothing (a partial resolution refuses before anything runs); a failed strip fails the verb with the library unwritten; an unresolved lossy_consent still blocks like any open decision.
  • Docs truth: CLAUDE.md's contract-note sections are back in chronological order (newest first; 3.38/3.39 came in from the basement) with a 3.48.0 section on top, and the spec's script table gains check_pdf.py, comments_census.py, and db_util.py. Note: the 3.46.0 entry claimed this docs item; the work actually lands here.
  • Flag-coverage backfill: --show-tags, --show-id, --primary-only, --plugin-data, --show-author-details, --set-comments, and --clear-comments had zero tests; tests/test_flag_coverage.py pins all seven against the real render and write paths (including the comments table's id column and the metadata_dirtied queue).
  • The tag tree renders real libraries again: 3.47.0's rolled-up counts re-derived the arithmetic inline and crashed on depth-3 subtrees (a dict child recursed as an addend; the real library's Fic → Classic → African tripped it in run_tests.sh, the suite's fixtures stopped at depth two) and dropped a parent's own direct books wherever it also had children. The renderer now consumes cquarry's tag_rollup directly (the engine derives, the frontend renders), so every node shows its true subtree total, and a regression test pins the real shape.
  • Suite: 575 → 601 tests.

3.47.0 (2026-09-16)

The engine predicates adopted: identifierless listing, tag-tree rollups, pace year buckets, engine-driven duplicates

  • --identifierless: books carrying no identifiers at all, listed as [id] title through cquarry's find_identifierless (L4 rank 11's curation queue for Calibre-Companion-style lookups); a fully identified library reports clean.
  • Tag-tree rolled-up counts: every node of --tags --tree's taxonomy now shows its rolled-up book count -- a node's total is the sum of the leaf tags beneath it (cquarry's tag_rollup arithmetic, rendered in the tree).
  • --analytics pace --pace-granularity year (L4 rank 10): the addition timeline buckets by year instead of month.
  • Duplicate detection is the engine's predicate now: the audit's hand-rolled (title, primary author) grouping is replaced by cquarry 1.8's find_duplicate_books -- same key shape, but the ids inside a duplicate row are sorted numerically where book-iteration order used to decide (the one visible CSV difference).
  • The stale series-clear pin updated to the cquarry 1.23.1 fixed semantics: a cleared series resets series_index to 1.0, never NULL on real schemas.

3.46.0 (2026-09-16)

The L4 headliners land -- Markdown search results, a machine-readable health digest, a TUI that finally covers the whole read surface, and honest format refusals

  • --search QUERY --format md (L4 rank 1): the Markdown catalog emitter over any query's match set -- headings per author, bulleted bold titles, the library UUID provenance header, and the query recorded in a scope note. Composes with --restrict free (the view is already the database), and --output names the file as anywhere else; the default output is search_results.md. Delegates to write_catalog, so the emitter cannot drift from the catalogs.
  • --health --format json (L4 rank 6, first half): the digest's counts as a machine-readable payload -- issue counts, problem tallies, metadata-quality rows, tree/FTS findings, pending OPF sync, and the annotations-dirtied count riding beside the OPF line (cquarry's get_annotations_dirtied_books, unconsumed since 1.23). --fail-on-findings is the second half: opt-in, flips the digest's exit from the standing 0 to 1 when anything was found, turning the dashboard into a gate. Quiet mode still gates on the exit, so scripts get both.
  • Per-mode --format corners refuse instead of ignoring (roadmap 1351): --book --format md used to render plain text silently; --fts --format csv/ai/md did the same; --health --format <not-json> ignored the flag entirely. All three now exit 2 naming the one format the mode supports.
  • The TUI menu covers Format Stats and the Trash Listing (roadmap 1312's residue plus the 3.45.0 gap): the README's "menu covers every read mode" claim is true again, and the 3.45.0 trash surface has a menu door.
  • The stale series-clear pin updated: test_set_series_with_index_then_clear still expected series_index = NULL after a clear -- the exact IntegrityError the cquarry 1.23.1 fix closed on real schemas (REAL NOT NULL DEFAULT 1.0). The test now pins the fixed semantics (clear resets to 1.0, the link row gone).
  • Docs truth: the spec's script table gains check_pdf.py, comments_census.py, and db_util.py; CLAUDE.md's contract-note headers return to chronological order.

3.45.0 (2026-09-15)

The curation verbs land over the 1.19.0 riders, the trash gets a surface, and the live prose layer cleans up

  • --rename-entity KIND OLD NEW: the fix-the-misspelled-name verb over cquarry 1.19's rename_entity, which had ridden the floor for two releases with zero consumers. Renames a tag, author, series, or publisher everywhere; a rename into an existing name MERGES the rows (links move to the survivor, duplicate links drop); a no-match old name is a clean exit 1 and an unknown kind a usage exit 2. Author renames recompute author_sort and re-lay paths; series merges renumber incoming books.
  • --set-author-sort / --set-title-sort BOOK SORT: the passthrough sort setters, storing hand-tuned corrections verbatim; a later --set-authors / --set-title deliberately recomputes over them.
  • The trash surface: --trash lists the library's .caltrash entries (category, book id, age, files) through a pure read-only filesystem inventory, and run trash --empty / --expire DAYS owns the lifecycle through cquarry 1.20's verbs, dry-run listing by default with --apply executing. The apply half opens metadata.db writable to reach the upstream verbs, so the closed-Calibre guard applies; no backup is owed because the database itself never changes. --format json carries the plan/results shape.
  • restrict's refusal list, SINGLE_BOOK_DESTS, and the setwrite combination guard keep step with the three new write dests, with a test pinning that they must; the README's full-help dump is regenerated byte-identical from the live parser.
  • The live em-dash layer is gone: all eleven rendered CLI separators (including the catalog header frozen into README's sample output), the README/spec titles, and every live README/spec line are recast; the patchnotes 3.40.0 entry's ASCII-dash sites are recast; the "not crying wolf" echo now appears once; the CalibreQuarry (cquarry-cli) rename injections are down to one first-use mention (fixing the possessive grammar break). Comments keep their dashes by convention.
  • Comment truths: _fetch_metadata's docstring names the failed return; the download-segment comment no longer names calibredb; integrate's docstring says which verbs actually drive external programs; setwrite's JSON report shape names the ids key; treeaudit's tolerance wording matches its extra_cover_file class; audit_isbns' exit contract names the advisory rider; spot_check's 99 means "99 or more failures"; compress_pdf's workflow names both synced size records; the three "Stdlib only" headers that import vir_tui say so; cli.py and tui.py carry module docstrings.
  • Polish: the five dead [DRY RUN] conditionals are gone; audit's dead nested quiet-guard is unwound; run_cover counts already-so rows like run_convert; run_flush honors --format json; the FTS extraction-errors preview gains its tail; the run.py/setwrite.py/fetch_library_codes.py pgrep guards fail closed on OSError like integrate and reconcile already did.
  • Housekeeping: scripts/ exec bits normalized (every shebang'd script executable; the shared db_util.py helper and the JSON template not), and pyproject takes the PEP 639 shape (SPDX license = "MIT" + license-files, the License classifier dropped, Python :: 3 :: Only added; wheel metadata verified). GitHub topics drop the duplicate python3 and near-zero python-314.
  • New-work boxes recorded on the roadmap: the dest-list constants module, the _backup_db triplication, the per-mode --format corners, bindery's unmirrored manual_watermark_repair decisions, the lossy double-manifest wrinkle, and the UI-only pypi tag policy.
  • Suite: 542 → 559 tests.

3.44.0 (2026-09-15)

The truth-and-hardening batch: the raw-SQL reads retire, every lying header tells the truth, and the publish path locks down

  • The frontend tier's last unrecorded raw-SQL reads are gone. _precedent_tags (the phase-3 prompt's four-table JOIN) is promoted to cquarry 1.22's CalibreDB.precedent_tags -- released upstream under the cross-repo grant and consumed here with the floor bump (cquarry >= 1.22.0) -- and _remove_book_dry_run reads through cquarry's CalibreDB instead of raw sqlite3. modes/fts.py's sidecar read remains the one recorded exception. Both reads are pinned end to end by tests.
  • The publish path hardens (the workspace batch, byte-identical across cquarry and CalibreQuarry; vir-tui verified as already shipped): SHA-pinned actions (the pypa publish action was a moving branch holding id-token: write), a top-level contents: read permissions block, a no-cancel concurrency group, the CI's pinned ruff gates before the suite, a strict twine check plus a wheel smoke-install, and a create-release job that mints the GitHub Release from the tag's verbatim message. Tag protection ships as a release-tags-protected ruleset (deletion blocked for refs/tags/v*), and an actions-only dependabot keeps the pins current. One piece is recorded as reverted: REST-created environment deployment policies are branch-type only and reject tag deployments outright (cquarry's v1.22.0 publish proved it live); a UI-applied tag policy is the reopen item.
  • The comment/header truth batch: the --fields help no longer advertises comments (backfill refuses it by design); validate_metadata's header lists all fourteen checks (three were missing); librarything's module header documents the real cquarry --exportlt surface instead of a standalone CLI that never existed; integrate's convert skip comment names both shapes it swallows; SINGLE_BOOK_DESTS actually carries every single-book dest (add_tag/remove_tag included, and setwrite drops its duplicate inline pair); the -> None annotations hiding load-bearing exit codes in export and catalog tell the truth; and reconcile_file_metadata's pgrep guard is fail-closed on OSError like its own comment always promised.
  • The docs tell the truth: the spec's accreted cquarry-floor sentence (stale three times) collapses into a floor-plus-bumps list with the floor pinned against pyproject by a test; the --format row names md; the no-network claim scopes itself to the read surface and names run backfill as the one exception; and the README test-suite section says 542 tests across 26 files, naming every suite (it said 373 and named ten).
  • The version-sync guard is real now: tests enforce the spec's Version header, the spec's cquarry floor, and the roadmap's Updated as of stamp, alongside the existing code/VERSION/pyproject/patchnotes set -- the three carriers whose drift the audit caught live.
  • Housekeeping: .gitignore gains venv/, .venv_ci/, .pytest_cache/, .ruff_cache/, and .claude/ (the tool self-ignores were masking the gap); the dead [dependency-groups] pytest residue is deleted; docs/taxonomy.example.yaml ships a placeholder library_path; README's LICENSE and screenshot links are absolute URLs that render on PyPI; and REPORT-12-Sept.md moves to docs/ with its roadmap references updated.
  • Suite: 536 → 542 tests. Floor: cquarry >= 1.22.0.

3.43.0 (2026-09-15)

The record-integrity batch: the same-day manifest collision, the Z-Lib provenance ruling, the lossy-consent hole, and the exit-code unification

  • Two same-day batches can no longer destroy each other's manifests (the final audit's HIGH). Phase 1 saved every batch to a fixed {date}-batch.json, so a second batch on the same day silently overwrote the first, destroying the durable record (imported ids, decisions) that phase-2 resume and phase-3 consume. The newcomer now takes {date}-batch-2.json, the same exists()-loop the backup paths already used three times. Regression-tested on a real collision.
  • Provenance seeds Z-Lib, the ruling's spelling. The 2026-09-13 enum ruling renamed the #source value to Z-Lib (Brandon's entry spelling), but the seeder still emitted the 3.40-era Z-Library: every z-lib file in a 3.42-run manifest failed its phase-2 stamp until hand-corrected (19 files in the 2026-09-14 STEM run). The seeder, the test fixture's enum, and both import skills now track Z-Lib; manifests produced by 3.40-3.42 still carry Z-Library and need the hand-correction before signing.
  • Bindery's gate-accepted EPUB repairs land in the manifest's lossy records (observed the same run). Bindery's phase-1 report carried an apply_lossy decision naming gate-accepted repairs, and the manifest still wrote {flagged: false, repairs: []}: the seal bound nothing, so signing consented to repairs it never saw. _mirror_lossy now walks bindery's repair records: status accept/partial becomes flagged: true plus the named repairs and an applied flag, in both dry and --apply-lossy runs.
  • The "Calibre is running" refusal is lock-class exit 1 everywhere. Set mode exited 1 while the five run/integrate doors (phase 2, phase 3, dispatch_integrate) exited 2, and scripts branch on these codes; the recorded discipline ("usage problems exit 2, lock/write errors exit 1") now holds across all of them. dispatch_integrate also runs its usage guards before resolving the library, so run convert with a missing --ids is a usage error (exit 2) however resolvable the library is. The two 3.41-era guard pins were updated with the reasoning.
  • --export refuses honestly. --export --format md (md became a legal catalog format in 3.42.0) and a bad --show-custom printed their refusal and exited 0, reporting success while writing nothing; run_export now returns 2 and 1 respectively, exactly like the search-export path. The implicit --wing catalog fallback also passes --format through, so --wing W --format md renders Markdown instead of silently plain text.
  • A typo'd TUI scope no longer runs unrestricted silently. The _restricted parse-failure note was printed and then immediately erased by the next terminal reset; it now goes through the blocking notice, so the reader must acknowledge it.
  • --restrict ... --tags obeys the universe. --tags was the last read mode reading a global aggregation through the RestrictedView; tag counts are now recounted from the scoped book rows. Correction recorded against the final audit: its ratings-recount half was a misdiagnosis (the recount's key was and is correct; a test pins it so the suggested "fix" cannot land later).
  • Suite: 525 → 536 tests. Floor note: the vir-tui floor moved to >=2.5.0 on 2026-09-14 (upstream 2.4.0/2.5.0 are additive; every fresh resolution takes it automatically).

3.42.0 (2026-09-13)

The blitz candidates close: --health, the era split, Markdown catalogs, and the TUI's Phase 19 surfaces

  • --health, the one-shot digest. The audit's finding counts in one short screen: book issues with the top problems, duplicate groups, series gaps, conversion overrides, the metadata-quality trio, the filesystem tree, FTS coverage, and the pending OPF queue. Both renderers now consume one shared derivation (collect_issues), so the CSV and the digest cannot drift; --restrict scopes the book-level classes exactly as it scopes --audit, and --health always exits 0 (a dashboard, not the audit's CSV).
  • The reading-analytics era split. Days-from-added-to-finished used one median over all spans, and backfilled pre-library reads (finished before their added date) dominated it: the real library's median was -489, which said nothing about how cataloged books actually read. When negative spans exist the two populations report separately (library era / pre-library, each with median, mean, min, max); a library with no pre-library reads keeps the previous single line unchanged.
  • Markdown catalogs. --format md renders the catalog's Markdown shape: one # header with the same provenance content, ## per author, bulleted books with bold titles, an hr and a bold total. --catalog passes its format through, and the wing and saved-search sweeps name their files .md. The plain text form is untouched when no format is given.
  • The TUI joins Phase 19. Five new menu entries, each with a shared scope prompt that adopts the --restrict modifier per invocation (blank = whole library; an expression resolves once through the CLI's RestrictedView; a parse failure notifies and stays unrestricted): Content Search (FTS), Saved Search Catalogs, and Library Health in the first section; Reading Analytics and FTS Index Status under Analytics. The Catalog and Catalog Wings entries gain a Markdown prompt. The menu structure is testable now, with the Settings section pinned where the s/q aliases need it.
  • Cosmetic: detail-less plan lines (backfill, polish, cover, flush dry runs) no longer end in a trailing space (the recorded Matrix 3 note).
  • Suite: 512 → 525 tests.

3.41.0 (2026-09-13)

The six-lens audit batch: three HIGH integration-seam defects, the run-verb hardening, and the routed metadata-quality rows

  • run flush --apply no longer writes outside the queue. The verb joined each chunk's distinct ids into one hyphen range ("5-900"), and calibredb reads a range as EVERY book between the endpoints, so a non-contiguous dirtied set embedded metadata into books the queue never named. Ids now pass space-separated, and the regression test pins a 1,3 queue against touching the between book. A bad --ids/--search on flush is a usage error (exit 2) instead of a traceback.
  • run phase2's metadata fetch can actually succeed now. _fetch_metadata passed the output path after -o, but -o/--opf is a store flag whose OPF arrives on stdout: the path was a silently-ignored stray positional, the success gate always failed, every import queued a bogus metadata_download decision, and the ok branch, _apply_opf, and the clobber watch were dead code (the sibling seam 3.39.2 fixed in the backfill; this call site was missed). Stdout is staged to the temp file like the backfill does; the ambiguity sniff reads "multiple" only, since the no-result log's "No matches found" classified every empty lookup as ambiguous. The seam tests mock the subprocess, not the verb, so the revived path runs for real.
  • run convert --apply no longer overwrites an existing target. Converting to a format the book already has used to plan the conversion anyway: ebook-convert overwrote the file, then the registration raised uncaught. Like the merge verb's move list, the plan skips such books ("already has TARGET"), apply counts them already-so, and a registration failure is a failed report row instead of a traceback.
  • --restrict is refused with the run verbs (spec 3.4's contract, previously bypassed by dispatch order): --restrict EXPR run ... ran unrestricted; it now exits 2.
  • The integration-verb pgrep guard is fail-closed: a pgrep timeout (or an unrunnable pgrep) answers assumed-RUNNING and refuses --apply, matching run.py's recorded semantics instead of proceeding against a live Calibre.
  • run backfill failures are real: a book whose OPF fails to apply is a counted failure that fails the verb (exit 1) instead of "Applied 0, failed/skipped 0" at exit 0; a malformed OPF or refused write is a report row, not a traceback; a hung lookup times out into a failed row; --fields isbn prefers the opf:scheme=ISBN identifier and falls back to an ISBN shape through cquarry's to_isbn13 (the first dc:identifier could be a Goodreads id); the fetch no longer appends a stray --opf <path> positional.
  • The catalog sweeps report per-file failures: --all-wings discarded write failures entirely ("All wings written" at exit 0), and --all-saved-searches counted a catalog before the write ran. Both now count only files that exist, drop the stale file of a failed entry, warn per failure regardless of --quiet, report "N of M written", and exit nonzero.
  • Set mode's backup goes through the sqlite backup API (the last copy2 door among the write paths; a file copy of a database with a hot journal can snapshot a state its WAL would never replay into).
  • --audit renders the routed metadata-quality rows (the bindery ruling of 2026-09-12): cquarry 1.21's find_invalid_uuids, find_sentinel_pubdates, and find_bad_language_codes as advisory issue_type rows with the offending values in brackets and summary blocks, one class per commit with fixtures and false-positive notes. Real-library probe: all three classes are clean at the database level; bindery's 51 OPF-085 warnings were file-side (stale sidecar OPFs), not DB drift.
  • Test hygiene: the __main__ guards of three test files moved below the suites appended after them, so direct-file runs exercise the 3.39/3.40 regression classes again (discovery was never affected); and the phase-3/backfill guard tests now pin the closed-Calibre pgrep instead of assuming it, so a desktop Calibre that happens to be open no longer flips five outcomes.
  • Suite: 483 → 512 tests. Floor: cquarry >= 1.21.0 (the write-path fixes plus the predicates).

3.40.0 (2026-09-12)

The post-release decisions: whitelist, Z-Library provenance, and a retraction

  • The enum "bug" was a misdiagnosis (retracted). The 3.37.0 notes recorded that the contains form over a normalized enum column (#reading_status:Read) matching the whole library was a cquarry engine issue. Instrumented comparison says otherwise: contains is honest case-insensitive substring semantics, identical to upstream's CONTAINS_MATCH (query in t), and in this library every status value ("To Read", "Reading", "Read") contains the substring "read", so a full-library match is the correct answer. cquarry needed no fix. Prefer the exact form (=Read) for precise enum selection; that guidance stands.
  • The tree audit whitelists the library root's workspace furniture (the 2026-09-12 decision): dot-entries (.claude/, .ruff_cache/, .nomedia, .trackerignore) and the named doc/tool set (CLAUDE/AGENTS/MEMORY/README/roadmap/spec/patchnotes/refresh/TAXONOMY .md, taxonomy.json/taxonomy.yaml, validate_library.py) are never findings: a library that doubles as a working checkout carries them by design. On the real library this drops 13 findings to zero.
  • Z-Library provenance goes forward (the enum decision pending since 3.35): z-lib filename markers (z-library.sk, 1lib.sk, z-lib.sk) now seed #source "Z-Library" instead of "Other". Your one-time step: add the Z-Library value to the #source enum in Calibre BEFORE importing a z-lib batch; enum validation refuses unknown values, deliberately. Existing "Other" rows stay untouched.
  • run backfill live-verified under an authorized one-off network drill on a scratch library: the fetch, the OPF round-trip, and the title write all work end to end (the drill also caught that the verb's plugin restriction allowed no metadata source at all; fixed alongside).
  • Suite: 482 → 483 tests.

3.39.2 (2026-09-12)

The functional-pass matrix audited backfill and the shared target plumbing; six defects fixed

  • run backfill --apply was dead on arrival: the command builder orphaned its own --opf argument (a leftover slice dropped the temp-file path), so the metadata source could never deliver and every book failed. Fixed and proven end to end: a mocked source writes a real OPF, the fetched title lands through cquarry's write module.
  • _apply_backfill looked in the wrong namespace: real fetch-ebook-metadata OPFs use dc:-prefixed elements; the code searched the opf wrapper namespace and would have found nothing. Fixed; the end-to-end test pins the real shape.
  • --fields comments was accepted, planned, and then silently ignored while reporting "applied": comments is now refused with the available list. Overwriting a curated description from a publisher OPF is a curation decision, not a backfill.
  • run backfill --apply required no --backup-dir despite mutating metadata.db: it is in the backup set now.
  • --search silently beat --ids when both were given (the ids, including unknown ones, were never validated): the combination is refused (exit 2), which is what "exactly one target source" always meant.
  • A bad --search expression crashed with a raw traceback in every verb: it is now a clean usage error (exit 2).
  • --format json and --quiet work after run (subcommand position), matching --db's dual-position treatment; plans and reports carry {plan}/{results} JSON either way.
  • Backups moved to after plan validation: an invocation aborted by a usage error no longer writes a backup nobody needs (timestamped backups made this harmless, but it was undocumented behavior).
  • Found by the functional-pass matrix (inspection + dry-run probes) and closed with 6 new regression tests, including backfill's first. Suite: 476 → 482 tests.

3.39.1 (2026-09-12)

The live drills caught two calibredb seams

  • run flush passed metadata.db to --library, which wants the library DIRECTORY: calibredb died with apsw.CantOpenError on every call. The verb now passes the directory and its id spans match calibredb's documented grammar (space-separated ids, hyphen ranges).
  • run export passed --dont-save-opf, which does not exist (the real flag is --dont-write-opf): every export refused with a usage error. Fixed, and both verbs now regression-test the exact command line they build. Found by the live mutating drills on a scratch library, which is what they are for.
  • Skills sync: swept both import skills for the C verbs; they are library-maintenance surfaces with no import-flow teaching and nothing went stale. Suite: 474 → 476 tests.

3.39.0 (2026-09-12)

Phase 19 C: the integration verbs (subprocess-driven; no new dependencies; every verb dry-run first)

  • run convert: ebook-convert over a resolved set (--search/--ids); the source is --from-format or the largest other format, the output lands in the book's directory under Calibre naming and registers through WritableCalibreDB.add_format. Failures are per-book and listed.
  • run polish: ebook-polish with --polish-ops (smarten, unused-css, compress-images, subset-fonts, jacket, kepubify) on each book's EPUB; the post-verify re-syncs the polished file's size through set_format, so Calibre never sees a stale size.
  • run cover: places an image (--cover FILE) or removes the cover (--remove-cover) through cquarry 1.19's set_cover/remove_cover; closes the loop the coverless/low-res/aspect audits open.
  • run export: calibredb export --template per resolved ids into --dest; calibredb missing is a clean setup refusal, not a traceback.
  • run merge --keeper ID --duplicate ID: the duplicate's unique formats are copied into the keeper's directory and registered, then the duplicate's removal goes to cquarry 1.20's trash (.caltrash/b/<id>/, recoverable by hand until expired or emptied).
  • run flush: the headless OPF-queue flush: calibredb embed_metadata over the metadata_dirtied ids in --chunk batches; an empty queue is a clean exit 0.
  • run backfill: drives fetch-ebook-metadata per book (network only at --apply) and applies the requested --fields through cquarry writes, one batch per book.
  • Shared discipline across all seven: dry-run by default with a printed plan; --apply demands a closed Calibre (anchored pgrep) and, for metadata-mutating verbs, a timestamped --backup-dir outside the library; targets resolve read-only first and unknown hand-supplied ids abort (exit 2) before anything opens writable; one report shape with per-book results and --format json.
  • check_library subprocess parity: skipped by the recorded A.3 route decision (the tree audit is CQ-native since 3.37.0).
  • Erratum for 3.38.0: that release's note said the skills sweep found "nothing stale"; the sweep also ADDED teaching (phase-1's check_pdf depth note and phase-3's year_mismatch interpretation), which the released entry undersold.
  • run now also accepts --db after the subcommand (the facility-run document has been suggesting that shape all along). Suite: 461 → 474 tests.

3.38.0 (2026-09-12)

Phase 19 B: the audit-depth batch (one class per commit, each with its fixture and false-positive note)

  • Content-duplicate fingerprinting (scripts/audit_duplicates_content.py): 64-bit simhash over 3-word shingles of each book's spine text finds re-downloads filed under different metadata; a bottom-32 shingle-sketch candidate pass catches omnibus containment (an omnibus's simhash is not close to the standalone's, so Hamming distance alone never sees it), and candidates are classified exactly: near_duplicate vs omnibus_overlap. Front-matter-only files are excluded (under 200 shingles a simhash is noise) and formats without an extractor are skipped, not guessed; legitimate public-domain reissues are findings by charter. ~1s per book: scope with --search/--ids.
  • Truncation cross-checks (scripts/audit_truncation.py): the Count Pages plugin's page counts meet poppler's pdfinfo; more than 20% disagreement is page_count_mismatch, and stale_plugin_data (catalogued size drifted, the plugin's own needs_scan flag, post-scan mtime) is its own class. Only PDF rows are checked: the plugin's EPUB pages are word-count estimates by design.
  • Author-sort sanity in validate_metadata.py: AUTHOR_SORT_NOT_INVERTED (sort identical to a multi-word display name) and AUTHOR_SORT_ORPHAN (matching no legitimate shape of the book's authors). Hosted locally by decision: cquarry never boxed the predicate, and the promotion remains a future option. The first cut compared against display names only and flagged 7718 of 7842 real books; the real-library probe caught it within minutes and the fix accepts every legitimate Calibre shape (display name, the authors.sort column, mechanical inversion, the &-joined multi-author sort). Real probe after the fix: 0 orphans.
  • DB-level ISBN checksum in validate_metadata.py: INVALID_ISBN warns on any isbn identifier failing its check digit (cquarry's existing helper; no file open). Blank-ish values are "no ISBN", not invalid ones.
  • Cover aspect-ratio bands (scripts/audit_cover_aspect.py): covers outside w/h 0.55-0.80 report as advisory cover_aspect_narrow/cover_aspect_wide through cquarry's header-only image readers (no image library). Legitimate landscape art exists; the audit surfaces the distribution, the operator judges.
  • FTS coverage audit in --audit: the A.1 staleness classes render as fts_coverage CSV rows. With no sidecar the CSV stays silent (the rows would be the whole library) and the prose summary carries the absent-sidecar note instead.
  • PDF battery depth in check_pdf.py: text sampled at pages 1, middle, and last (the page-1-only sample passed OCR-once scans clean; a partial pass is now text_layer_partial), and image DPI parsed from the existing pdfimages -list output as an area-weighted mean (a small sharp logo cannot hide a full-page 72-dpi scan; below 150 dpi is a low_dpi advisory). Advisory classes only: the structural total and the run-verb seam contract are unchanged.
  • Copyright year vs pubdate in audit_isbns.py: the front-matter pass captures (c)-years (each copyright marker claims the years on its line), and an earliest-year-vs-pubdate gap over two years is a year_mismatch advisory with its own report section, riding the exit-1 findings contract. A reprint legitimately prints the original year; the report asks whether the pubdate describes this edition.
  • Skills sync: swept both import skills for the new audit surface; nothing teaches the old shapes, no staleness found. Suite: 433 → 461 tests.

3.37.0 (2026-09-12)

Phase 19 A: --restrict scopes everything; FTS content search; the tree audit; reading analytics

  • --restrict SEARCH scopes every read mode. Stats, audits, analytics, exports, catalogs, full-text search, and the rest compute over the books matching a search expression (a wing composes as vl:Name). The implementation is a scoping view over cquarry's connection: every predicate and stat is still derived by cquarry, only the inputs are scoped, and the two SQL-level aggregations (entity counts, format stats) are recounted so they stay honest. Book-level audit findings follow the restriction; library-shape ones (orphan dirs, root strays) always report globally. Write verbs and --book/--id refuse the combination (exit 2); a bad expression exits 1 like --search. Upstream precedent: --restrict-to on calibredb fts_search and restricted-id category counts.
  • --fts QUERY searches what the books say. Content search over Calibre's full-text-search.db sidecar (the plain books_text table; no FTS5 machinery), case- and accent-folded, every match naming its formats, --format json export through the output guard. Every run ends with an index-staleness summary in separate classes (never indexed; indexed empty; extraction errors; stale entries queued in dirtied_formats), and --fts-status reports just that. A missing sidecar (the common case until Calibre builds its index) degrades with a note, never an error.
  • --audit now walks the filesystem. Tree rows against the database: missing book dirs and format files, extra format files and unknown files inside book dirs, covers on disk the DB does not claim, orphan book dirs and author dirs, malformed book-dir names, stray root files, and unreadable directories. CQ-native by decision (no calibredb dependency; the walk composes with --restrict and the shared CSV shape); upstream check_library's class list is the completeness checklist. metadata.opf, any *.opf, cover files, and data/ are never extras; name comparison is case-insensitive.
  • --analytics reading (read-only). The #reading_status funnel in the column's configured enum order with (no status) last; recent finishes from #date_read, newest first; days from added to finished with median/mean/min/max and a stale-timestamp note. The NON-NEGOTIABLES write ban is untouched: an mtime-pinned test proves nothing writes. Missing columns degrade to a clear message.
  • --all-saved-searches. One catalog per saved search into --outdir (the --all-wings analog), each headed with the search's own expression; zero-hit searches write nothing and say so; an unresolvable search is skipped with a warning, never a dead sweep.
  • @Name user-category resolution: skipped by its own gate. The library's preferences carry zero user categories (read-only peek), so the box records not-in-use instead of building a resolver. If categories ever appear, the right home is the cquarry search engine.
  • Adopts cquarry 1.20.0 (floor bump ahead of Phase 19 C's integration batch; the trash lifecycle and create/delete_custom_column land in the shared layer this CLI's next release consumes).
  • Adopts cquarry 1.18.0 (floor bump only; no behavior change required here yet): the FTS sidecar reads, the search-parity honesty pass, and the write completions land in the shared layer this CLI consumes; Phase 19 consumes them for real.
  • Adopts cquarry 1.19.0 (floor bump only): the four approved write verbs land in the shared layer (rename_entity/remove_entity_everywhere, the misspelled-author fix; set_cover/remove_cover, which closes the low-res-cover audit loop; the verbatim sort setters; and save/restore_original_format). Phase 19's curation and cover verbs build directly on these.
  • Found while composing --restrict with the status column and recorded for the cquarry lane: the contains form over a normalized enum column (#reading_status:Read, quoted or not) currently matches the whole library; the exact form (#reading_status:=Read) is correct. Fixing it belongs upstream in the cquarry search engine.
  • Skills sync: phase-3-import's audit step names the new tree classes, runs the audit with --output outside the library root (a stray audit.csv there is now itself a tree finding), and uses --restrict to scope post-import verification to the batch's tag; phase-1-import swept, nothing stale. The README's full help dump was regenerated from the live parser. Suite: 385 → 433 tests.

3.36.0 (2026-09-10)

The audit mode absorbs conversion overrides

  • --audit reports manual conversion overrides. The last open box (promoted from the 3.34 sweep's scripts verdict): the check behind scripts/audit_conversion_overrides.py now renders inside the audit mode. Every book carrying a per-book conversion_options recipe gets a conversion_override row in the CSV (book id, title, author, the format, and the recipe blob's size) plus a summary block naming the affected books, with the pointer to Calibre's conversion dialog where those blobs are inspected or cleared. The rows consume cquarry's get_conversion_profiles (the predicate's library home; the frontend renders, never re-derives, and the pickles are never unpickled).
  • The standalone script stands. audit_conversion_overrides.py keeps its pipeable surface (--quiet prints only the ids; exit 1 when any are found, so the report feeds a repair workflow). The mode keeps --audit's exit-0 reporting contract; no exit-code change for existing users of either surface.
  • Housekeeping: the --audit help line names the new check and the README's full help dump was regenerated from the live parser (the dump had also drifted from the parser on --format-stats' position, which this regeneration fixes). Suite: 382 → 385 tests.

3.35.0 (2026-09-10)

The correctness batch: the stamp convention, provenance seeding, and the qpdf verdicts

  • The filename-stamp convention is decided (roadmap :902). The stamp writer is the authority: run.py's parser reads "Author - Title", the stamping path emits exactly those values, and the observed corpus (libgen.li's "[Series] Author - Title (year, publisher) - site" names, verified in the 2026-09-10 Redwall run) confirms the direction, so that convention stands. Calibre's own filename fallback guesses the opposite (probed on a metadata-less file: "Brian Jacques - Mossflower.pdf" imports as Title "Brian Jacques"). The decision is now recorded where the readers live: run.py states the convention, stamp_pdf's preview documents that it deliberately mirrors Calibre's opposite guess (that is what a preview of an unstamped import is for), and screen_duplicate names the screening gap metadata-less "Author - Title" files carry. A corpus regression test pins the direction.
  • run phase1 seeds the manifest's provenance from the filename (the second Redwall box). The field existed but was never populated, so phase 2's cc6 stamp fell back to a blanket "Anna's Archive" regardless of true source and phase 3 re-derived provenance from filenames every batch (observed 2026-09-08 and 2026-09-10). Phase 1 now derives it from the same filename evidence the review already worked with, mapped onto the #source enum's vocabulary: the "-- Anna's Archive" trailer seeds Anna's Archive, libgen.li seeds Library Genesis, z-library.sk/1lib.sk naming seeds Other (the recorded practice of both runs, pending Brandon's Z-Library enum decision), and a name with no marker seeds nothing. The review step corrects the value exactly like the stamps, and phase 2 keeps stamping cc6 from the reviewed value. The HMAC seal now binds provenance alongside the stamps and lossy flags: a post-sign source swap fails every load until re-signing.
  • check_pdf.py classifies qpdf exit 3 by its documented contract. qpdf writes its warnings (and the "operation succeeded with warnings" summary) to stderr, so the old stdout marker gate never matched and every warning-only file was recorded as errors in the battery report, the CLI summary, and the phase-1 manifest; the Redwall run filed two benign PDFs (unknown-token tolerance; linearization /E + hint-table drift) as structural damage. The class now reads the exit code alone (0 clean, 3 warnings, 2 errors) and a warning finding carries the first warning line as triage evidence. Re-triage confirmed both warning kinds benign; the Multics PDF re-checks clean today because the phase-1 stamp rewrite healed its linearization drift.
  • The parser-disagreement box (:902) and the Redwall run's two (the provenance field, the qpdf exit-3 class) are closed; the --audit promotion of audit_conversion_overrides is the one box still open, landing next. Suite: 373 → 382 tests.

3.34.0 (2026-09-10)

Batch C: the read surface, the tests, and the docs (Phase 18 closes)

  • The TUI respects the saved library. _resolve_db_for_tui consulted a hard-coded default list that started with a CWD-relative metadata.db and never read the saved config, so launching the TUI from any directory with a stray metadata.db overwrote the shared config and the next CLI run read the wrong library. The saved path wins whenever it exists; discovery binds only when nothing is saved.
  • Failures exit like failures. A bad --search expression exits 1 (matching --exportlt --search, which always did); an unknown --wing exits 2 instead of leaving a stale catalog file standing in for a fresh one; and exportlt's self-check verdict ("do not upload") reaches the exit code instead of dying behind an unconditional 0 (its raw-SQL half was already retired by the cquarry 1.17 export_rows adoption).
  • Read-mode papercuts. --export-annotations/--exportlt/ --format-stats join the mutually exclusive read-modes group (--format-stats had been declared in the write-verbs group; --untagged stays a --book modifier on purpose); negative --recent is refused with exit 2; a corrupt epoch renders raw instead of killing --reading-progress/--book; the TUI Entity Browser prints a readable refusal for an unknown kind instead of paging a traceback; plugin values ride along in json/csv/ai serialization instead of vanishing; and a directory export target that exists as a file is refused in prose.
  • The suite runs the tests, and the scripts run at all. run_tests.sh executes the hermetic unittest suite before the real-library smoke (following the documented command used to deliver smoke-only coverage). New tests/test_instruments.py drives the actual companion scripts through the actual seam adapters, and immediately caught a genuine bug: check_pdf.py referenced args.quiet without declaring the flag, so the PDF/DJVU battery crashed on every invocation, invisible for exactly as long as its callers ignored exit codes and empty reports. The flag is declared. Suite hygiene: the mid-file unittest.main() guards that silently truncated direct runs moved to true EOF (test_scripts.py had 1028 lines after its guard), and test_manifest.py's real-library path literal is synthetic.
  • Scripts housekeeping. fix_cq_lint.sh (the sweep's only delete) is gone; taxonomy.example.yaml moved to docs/ with the README pointer updated; comments_census.py's --json help no longer claims a runner that never consumed it. Two deferrals carry dated notes in the roadmap: the six drifting _SCHEMA fixtures still want a shared builder, and db_util's consolidation waits because the private connect_ro copies have genuinely drifted (reconcile needs Row rows and its own temp layout); the audit_conversion_overrides -> --audit promotion is now its own open box.
  • Docs truth. The README troubleshooting line claiming saved searches "match nothing" is corrected (they evaluate, cquarry 1.1+); the search engine is named where it lives (cquarry.search, not a nonexistent local file) and the "zero dependencies" claim is replaced with the real dependency set; the six botched "minimal-dependency (uses tqdm)" artifacts are cleaned. The spec absorbs the seven shipped modes its table never listed (--book, --entities, --reading-progress, --columns, --info, --exportlt, --format-stats), --set-pubdate/--clear-pubdate join the verb list, and §5's closing sentence names all four writer scripts. The README grows a run-verbs prose section (the file-side consents: --stamp, --apply-lossy, --quarantine; --bindery-report; --audience) and companion script sections for the five tools that had none, while the full help dump is regenerated from the live parser (--yes gone, --quarantine in). Housekeeping: Phase 17's seven malformed double checkboxes normalized, the patchnotes H1 moved to the top of the file, the pre-3.14 ## vX.Y.Z entry headings normalized to the dominant # X.Y.Z style, and the README test count made current.
  • Phase 18's 26 boxes are now 25 closed and 1 open (the audit_conversion_overrides promotion, opened from the sweep's scripts verdict). Suite: 362 → 373 tests.

3.33.0 (2026-09-09)

Batch B: the run verbs and the write path (Phase 18)

  • Phase 1 moves files only on true DRM and only with consent. Quarantine used to run for any DRM verdict that was not clean, sweeping in audit_drm's BENIGN (font obfuscation) and N/A (DJVU): every DJVU in a batch was moved and a manual_repair decision recorded for a file with no DRM at all. The classification is now audit_drm's own problem set (DRM, plus ERROR, a scan that could not verify), and the move is gated behind a new --quarantine flag: without it the verdict and decision are recorded and the file stays. Quarantine also no longer moves onto basename collisions (numbered siblings instead), stamp backups live in a dated temp directory outside the tree (which makes --stamp actually work, since stamp_pdf refused the in-tree backup dir it was handed), stamp failures print a WARNING instead of passing silently, and a rerun can no longer sweep _stamp_backups or _quarantine as books.
  • Phase 2 keeps its accounts. The resume record saves the moment the import batch commits, before the unguarded download segment, so a crash there no longer costs the imported ids. --audience with no flag stamps the documented default instead of the literal string 'None'. Backups are timestamped and taken through sqlite's backup API, so a second run keeps its own restore point and a hot journal cannot leave an inconsistent snapshot. The rollback message tells the truth (the batch rolled back; nothing was written). The dead --yes flag is gone, and phase 1's --apply-lossy is now actually wired to the bindery slice. add_book's orphaned directories on batch failure are closed upstream (cquarry 1.15's batch-scoped compensation), as is the byte-identity duplicate floor.
  • Phase 3 has the same rails as every other write path. A closed-Calibre guard runs before the answer gates; the answer file may not name #reading_status/status/date_read (checked against the shared NON-NEGOTIABLES tuple before anything opens writable); bindery/reconcile trouble (their exit 2) always prints, lands in the batch record, and fails the verb instead of hiding behind a clean validator; and the download segment re-checks the Calibre guard, deferring to phase 3 rather than racing a library that opened after the commit.
  • Run-verb papercuts. The PDF battery now covers the same recursive inventory everything else uses (not just the top level) and honors check_pdf's exit codes, with its report in a temp file; answer-file loading produces readable errors and warns on ids the manifest never imported instead of dropping them silently; a download timeout maps to failed, not ambiguous; quarantined files appear in files[] (the quarantined verdict is no longer writer-dead); the dead _existing_book_ids helper is deleted; and phase 2's calibredb set_metadata shell-out is replaced by _apply_opf, which applies a downloaded OPF through cquarry's write module in one batch (the no-calibredb constraint holds again).
  • The single-verb write path batches. run_write wraps every action in wdb.batch(), so an interrupt mid-verb cannot commit a torn edit invisible to Calibre's OPF sync (cquarry 1.15's __exit__ rollback is the upstream half of the same fix).
  • --commit-per-book does what it claims. Each book is its own outermost transaction; a book whose verbs failed rolls back alone, its entries read rolled_back/book_committed: false, and the pass continues. Partial rollbacks are reported as such; exit 0 only when every book committed.
  • Empty-string flag values are refused arguments (exit 2). --batch-set-title "" used to vanish through the truthiness gates while --batch-set-column audience "" was collected and cquarry treats '' as a clear, wiping the column on every targeted book with rc 0. Both are refused now, and _has_verbs counts value-bearing flags by presence so the refusal message is what the user sees.
  • The banned-columns refusal is a shared chokepoint. #reading_status/status/date_read were refused at set mode's door and open at the other two (single-book --set-column/--clear-column, the TUI). The refusal now lives in the writeops action builders (writeops.FORBIDDEN_COLUMNS, one tuple for every door), raising the argument-level exit 2; cquarry 1.17 deferred the policy upstream, so the frontend owns it.
  • Set/write papercuts. --batch-set-cover maybe exits 2 instead of tracebacking; the --batch-clear-rating manifest gate validates the file with the repo's own manifest module (a plain id file no longer unlocks a bulk rating clear); failure detail survives --quiet on stderr; the --remove-book dry run is a read-only connection, never WritableCalibreDB; set-write backups are timestamped; parse_book_id rejects int()'s '5_0'; the dry-run JSON carries the resolved ids; and the --backup-dir usage check precedes the pgrep probe.
  • Skill sync (same release): the phase-3-import skill's --batch-clear-rating line names the sealed-manifest gate; the phase-1-import skill names --quarantine among the file-side consents. Suite: 330 → 362 tests.

3.32.0 (2026-09-09)

The four P0s from the deep-dive backlog (Phase 18, batch A)

  • The phase-1 seams work on real input. run phase1 crashed on any non-empty directory: the duplicate screen's JSON report is a bare list of per-file records and the runner read it as a dict, and the same comprehension would have flagged every screened file as a duplicate (only records with library or within-batch hits count now; an unparseable report is a hard error instead of a silent pass). The screener is handed exactly the files its own extension set covers, so a djvu-only tree is a clean screen rather than the script's exit-2 setup error. The bindery seam now passes the --json FILE argument bindery always required and reads the report back from the file; bindery's exit contract (2 = trouble found, the report is still written; 1 = invocation problem, no report) replaces the raise-on-2 that fired on exactly the case phase 1 exists to surface, and a missing report names a stale entry point instead of sailing on. The seam tests pin the instruments' real payload shapes, including one run against the actual screen_duplicate.py.
  • The manifest signature is a real seal. sign() used to set a bare "signed": true in the same editable JSON, and a manifest whose rejected file was listed as approved passed phase 2 (proven end to end by the sweep). Signing now seals the approved set, the per-file stamps and lossy flags, and the decisions list with HMAC-SHA256 over canonical JSON; every load of a signed manifest recomputes it and refuses a mismatch with a re-sign hint. The seal is tamper-evidence, not secret authentication (the key is a schema constant); the honest re-sign path is the new cquarry run sign --manifest FILE verb, which checks structure but not the stale seal so deliberate edits can be re-approved. approve() refuses to list a file whose verdict is not approved_for_import, and validate() flags any list/verdict disagreement it sees, so the forged-approval attack is caught twice. Phase 2's own appends re-seal on save, keeping the retained manifest verifiable for phase 3. Manifests signed before this release must be re-signed (none exist outside tests: the verbs failed on real input until now).
  • A read mode can no longer overwrite metadata.db. No output writer compared its path to the database path: --export --output <library>/metadata.db replaced a fixture database with a JSON report, exit 0. New cquarry_cli/output.py is the shared last mile for every read-mode file output (--export, --search --output, --catalog, --audit, --export-annotations, --exportlt): the database and its sqlite sidecars are refused before anything opens, the directory exporters also refuse the library root itself (their stale-file sweep would write and delete inside the library), every file stages through a temp copy replaced into place only on clean close, and the refusal exits 2 as an argument error.
  • The TUI survives a malformed database. CalibreDB was constructed outside every exception boundary in the menu loop, so a corrupt or foreign sqlite file at the chosen path ended the session in a raw traceback, and Change Database validated only the filename suffix. The loop now probe-opens the database every iteration; a failure prints a prose notice and drops into the re-prompt, Change Database refuses a path that does not open and keeps the configured database, and a sqlite3.Error raised mid-session degrades the same way.
  • Skill sync (same release): the phase-1-import skill's known-defect paragraph is retired (the seams work as of this release; the scripts-tree caveat for wheel installs stays) and its sign instruction now names cquarry run sign with the seal semantics. The phase-3-import skill's manifest references were swept; none went stale.
  • Suite: 298 → 330 tests.

3.31.0 (2026-09-09)

Cascade: cquarry 1.17 adoption

  • The LibraryThing exporter runs on export_rows(). build_rows retired its hand-rolled correlated-subquery SQL for cquarry 1.17's flat row provider (cquarry 1.16's native custom-column lists underneath); same columns, same author_sort/title order, same sentinel and empty shaping, so the CSV stays byte-identical for the same library state.
  • Native-list display. With cquarry 1.16, multi-valued custom columns read back as native lists; the csv/ai export writers and the catalog renderer re-join them with the historical comma form so output shape is unchanged (JSON keeps the native list).
  • cquarry floor moves to >=1.17.0 (the promoted APIs this release adopts).
  • Test fixture fidelity: the phase-2 fixture's #source column now mirrors the real library (enumeration, normalized storage, the real enum values) instead of a text+direct shape no real Calibre schema creates -- cquarry 1.15's datatype dispatch refuses that shape, and the multi-file import test payloads are now distinct so add_book's byte-identity floor is exercised honestly.

3.30.0 (2026-09-06)

Phase 17 closes: run phase1/phase2/phase3 and the acquisition manifest

  • The manifest (acquisition-manifest/1). One JSON per batch in the library-local .claude/manifests/, RETAINED as the durable machine-readable record (the prose .claude/project_*.md files stay the human summary). Per-file verdicts, checks, lossy flags, filename-derived stamps, provenance, import outcomes, decisions_needed (fixed taxonomy), and approved_for_import. The six 2026-09-06 decisions are structural: provenance stamps #source, #audience is unconditional Brandon (no per-file column), a signed report is standing consent for the listed lossy repairs, refused duplicates and download failures are decisions while the batch continues, and phase 2 is non-interactive.
  • cquarry run phase1 DIR vets a downloads directory into a manifest: duplicate screening and the DRM scan drive the companion scripts, the PDF/DJVU battery rides the new scripts/check_pdf.py, bindery's phase-1 EPUB slice runs read-only (--bindery-report captures its JSON), and --stamp / --apply-lossy are the file-side consents. Read-only against metadata.db; DRM-locked files quarantine; clean files approve.
  • cquarry run phase2 --manifest FILE imports the signed manifest. Guards first (signed; no blocking decisions; Calibre closed; mandatory --backup-dir outside the library), then ONE batch(): add_book from the manifest stamps (cquarry 1.14), the pathway reset (clear tags + rating) scoped to the imported ids, #source and #audience stamped. Metadata downloads run after the commit via fetch-ebook-metadata + calibredb set_metadata; failures and ambiguities become decisions, and a post-download author clobber watch records what the download changed. Resumable: imported ids are written back and skipped on re-run.
  • cquarry run phase3 --manifest FILE curates the imports: the batch set is manifest ids crossed with find_untagged, dossiers come from --book --format json, and the decision gates (tags by precedent, description rewrite, field fixes) run as TTY prompts or from an --answer-file in ONE batch(). Then bindery run phase3 and reconcile_file_metadata.py --apply --repair-pdf drive, the validator re-runs (exit 1 unless clean), and the prose batch record lands beside the manifest.
  • --book --format json emits get_book_dossier dicts verbatim as a JSON array (raw comments included) - phase 3's structured input.
  • scripts/comments_census.py stands up the phase-3 skill's inline three-liner: a read-only census of the mechanical description defects (double hyphens, spaced-hyphen dashes, markdown bold, tag debris, body shape, soft hyphens, zero-width characters, mojibake, lost ligatures, exact-duplicate bodies) with --id scoping and a --json report.
  • Skill sync + the pathway amendment. The library CLAUDE.md carries the 2026-09-06 amendment (phase 2 tool-driven, sign-off = the signed manifest, the ratings ban as an anti-library-wide-predicate rule); the phase-1 skill names run phase1 (its manual commands stay the appendix), the phase-3 skill names run phase3 and the set forms.
  • Suite 280 -> 294.

3.29.0 (2026-09-06)

Phase 16: set-oriented write verbs, one target set at a time

  • One target set, many --batch-* verbs, one transaction. The write side catches up with Phase 15's batch-shaped read side: exactly one target source per invocation (--ids 1,2,3, --from-search 'EXPR' resolved read-only through the search engine, --from-untagged reusing find_untagged, or --from-manifest FILE with ids one per line or comma-separated) feeds id-less verbs: add/remove/clear tags, clear rating, set/clear column, add-column-value (append for multi-valued columns, via cquarry 1.13's add_custom_column_values), set/clear pubdate, set title/authors/publisher/languages/series (+--series-index), set/clear identifier, set cover, remove format. Hand-supplied ids are validated read-only before anything opens writable: unknown ids are reported and the run aborts (exit 2). Deletion has no set form: --remove-book stays per-book.
  • Dry-run by default. A set write prints the verbatim target, the resolved id count and list, and a per-verb preview; nothing opens WritableCalibreDB without --apply. Apply demands a closed Calibre (the anchored pgrep ^calibre guard, the fetch_library_codes.py precedent) and a mandatory --backup-dir outside the library directory (the stamp_pdf.py precedent), then commits the whole pass as ONE batch() transaction: any failure rolls everything back. --commit-per-book is the documented non-default escape hatch for very large sets.
  • The rating carve-out is mechanically encoded. The library NON-NEGOTIABLES ban bulk rating edits, so --batch-clear-rating is accepted ONLY with --from-manifest (the manifest proves which ids that run imported); --ids, --from-search, and --from-untagged are refused with exit 2. The column verbs refuse #reading_status, status, and date_read by label (case-insensitive), belt-and-braces on the absolute ban.
  • Honest reporting. The write path's action builders now carry cquarry's changed returns, so multi-verb summaries report applied: / already-so: per verb instead of a blanket ok: (the all-or-nothing rollback contract is untouched). Set mode reports per-verb applied/already-so/failed counts plus the per-id failure list on stderr, and --format json emits {target, verbs, results[{id, verb, status, detail}], committed, dry_run}; a failed row rolls the pass back and reports committed: false (exit 1). Exit 0 committed or dry-run, 1 failures or lock, 2 usage.
  • --set-rating ID 0 now clears (behavior change). It used to write a phantom 0-rating row that reads as unrated everywhere while polluting the ratings table; 0 now routes through the true clear (set_rating(id, None): link deleted, orphan pruned), matching Calibre's own 0-stars semantics. This amends Phase 16's no-single-verb-change non-goal for exactly this one case, Brandon's call (2026-09-06); the other three open questions (flag naming, manifest-only carve-out, --backup-dir) landed as specced.
  • Skill sync. The phase-3-import skill names the set forms where it taught per-id CLI loops, with the rating carve-out called out so rating changes there stay per-book.
  • Suite 225 → 257.

3.28.0 (2026-09-06)

--genre-depth N: descend the genre hierarchy

  • Second-level genres (and beyond) in the breakdown. --analytics genres now takes --genre-depth N (default 1, the previous root-only shape). Depth 2 renders each root's children indented beneath it, labeled by their last path segment (SciFi under Fic, not the full Fic.SciFi path); deeper flags go deeper. No cquarry change was needed: genre_distribution() already returns every node of the hierarchy tree-ordered, so this is renderer slicing only.
  • Every level stays a share of the whole library. A child's percentage is its fraction of all books, not of its parent, so children need not sum to their parent (a book with both Fic.Fantasy and Fic.SciFi counts once toward each). The TUI's Genre Breakdown entry prompts for the depth. --genre-depth below 1 is a usage error (exit 2).
  • Suite 216 → 219.

3.27.0 (2026-09-06)

--analytics genres: every genre's share of the library

  • Genre breakdown by percentage. --analytics genres renders cquarry 1.12's new genre_distribution(): every top-level genre (the root of the dot-path tag hierarchy) as a share of the whole library, biggest first, with a bar per genre. A book counts once per genre even when several of its tags share an ancestor; a book tagged into two roots lands in both, so the shares are honest fractions of the library and can sum over 100% (the output says so). Untagged books get their own row when present. The deeper levels of the hierarchy stay where they belong: --analytics tags remains the taxonomy tree.
  • TUI menu. The Analytics section gains a "Genre Breakdown" entry (between Tag Tree and Wing Overlap) backed by the same renderer through the pager. Shares and ordering are cquarry's math; percentages, bars, and the caveat are this repo's formatting, per the frontend-only split.
  • Suite 213 → 216.

3.26.0 (2026-09-02)

Phase 15 closes: batch dossiers, the clear verb, and the fetch script's single-pass rewrite

  • --book batch forms. --book takes a comma-separated id list (--book 8884,8885,8886) and composes the dossier renderer in a loop; a bare --book --untagged (no ids) selects every untagged book through cquarry's find_untagged() — the phase-3 curation entry state, which used to mean a hand-rolled get_book() loop per batch. An unknown id inside a list renders the books that exist and exits 1, matching the single-id behavior; conflicting or empty forms are usage errors (exit 2).
  • --show-custom speaks both names. cquarry 1.9's dual resolution flows straight through load_custom_column, so --show-custom now accepts the display name (Status), the bare label, or the #label (#reading_status); the docs' old "asymmetry" warnings are rewritten. No code change was needed beyond the cquarry bump — the flag routed through load_custom_column all along.
  • --clear-identifier BOOK_ID TYPE existed since 3.19.0 but reached the deletion as set_identifier(..., ""); it now calls cquarry 1.9's explicit clear_identifier() (same observable behavior, honest naming).
  • fetch_library_codes.py gains --all-codes and a misses worklist, and its write path moves inside the sanctioned one.
    • LCC and DDC arrive in the same SRU response, so --all-codes stores both in ONE pass and selects every book missing either code; the old flow (an LCC --apply, then --apply --write-ddc behind it) created exactly the concurrent-writer contention that bit the 2026-09-02 import.
    • Bug fix riding along: the DDC write was nested inside the LCC-hit branch, so a book the catalogue has DDC but no LCC for was silently dropped even under --write-ddc. The writes are now independent.
    • The writer is no longer a raw INSERT OR REPLACE: it goes through cquarry.write.WritableCalibreDB.set_identifier in one batch() transaction, so touched books land in metadata_dirtied (Calibre regenerates their sidecar .opfs) and last_modified moves — neither happened before. Retry/backoff over database is locked rides on top of the module's 30s busy timeout; identical values are honest no-ops, so a --refresh re-run reports the real change count.
    • Every book ending the pass with no LCC is written to a misses worklist (--misses-file, default fetch_library_codes_misses.txt) as id<TAB>isbn<TAB>ddc<TAB>title — the manual-research pass starts from a file instead of terminal scrollback (2026-08-27: three misses tracked by hand). The worklist is a report artifact and is written in dry runs too.
  • Version/docs re-sync. This roadmap's header catches up (was "as of v3.23.1"); README, spec's companion table, and CLAUDE.md carry the new surface. Phase 15 is fully ticked; the phase-3 skill names the batch forms and the one-pass fetch invocation.
  • Tests: 201 → 213 (batch/untagged forms and exit codes, --show-custom's both-forms catalog render, --all-codes selection semantics, the misses worklist, and write_identifiers' change-counting, dirtied queueing, and wait-out-a-concurrent-writer behavior).

3.25.0 (2026-08-30)

  • Feature: scripts/screen_duplicate.py — the phase-1 duplicate screen as a tool. Reads each download's embedded title/authors/ISBN with ebook-meta (filenames are display hints only; they lie), matches exact ISBN first, then normalized title + first author over cquarry's tight = search with Python-side normalization (accent/case fold, leading-article strip, subtitle/edition scrub — "Capital: Volume I" screens against "Capital: A Critique of Political Economy"), and screens within the batch too. Prints the comparison columns (existing id/title/authors/formats/size/pages) the skill says to report before judging keep/upgrade/re-source; --format json for the batch report; report-only, exit 0/1/2.
  • Feature: scripts/stamp_pdf.py — the phase-1 pre-stamp as a tool. Fixed exiftool field set (Title/XMP-dc:Title, Author/XMP-dc:Creator, XMP-dc:Publisher, ISBN via -Keywords=isbn:... — never -XMP-dc:Identifier, which Calibre maps to a bogus doi); multi---author joins with " & " (documented opposite of --set-authors' ;); dry-run by default with the filename-derivation preview beside the requested stamp; --apply requires --backup-dir, refused if it resolves inside any target's directory (a backup beside the file gets imported); verifies with ebook-meta and prints STAMP_FAILED on read-back disagreement instead of re-fighting a stubborn XMP store. Mechanics only — value choice stays the agent's research step (stated in the docstring).
  • Docs: spec §5's companion table gained both scripts; README gained a companion-scripts section; the phase-1 skill names both tools in its duplicate-screen and pre-stamp steps.
  • Dependency: Requires the latest cquarry (1.8.x).

3.24.0 (2026-08-30)

  • Upgrade: --book renders over cquarry 1.8's get_book_dossier() — the composed deep fetch lives in the library now, and the dossier gains a publication date in the facts line (published YYYY-MM-DD, skipped while the 0101 undefined-date sentinel stands). Output is otherwise byte-identical to 3.23.x.
  • Upgrade: --audit's per-book predicates (untagged, unrated, authorless, formatless, deprecated-format-only, coverless, missing cover file, low-res cover) are cquarry.integrity calls — one shared definition of "incomplete" across the ecosystem. Verified byte-identical CSV against 3.23.1 on the real library (7,641 rows). Duplicate grouping stays inline on purpose: the CSV joins ids in scan order and find_duplicate_books() sorts numerically.
  • Upgrade: --analytics and --stats render over cquarry.analytics (author_stats, addition_timeline, rating_distribution, and --analytics overlap over vl_overlap, unparseable wings still skipped); output verified identical on the real library for every subcommand. author_stats gained an additive rated_count key so the "(N rated)" rendering survives the move.
  • Upgrade: The LibraryThing exporter and scripts/audit_isbns.py use cquarry.helpers' ISBN family (isbn_normalize / isbn_check_digit_is_valid / to_isbn13) instead of two divergent local copies. The exporter maps None back to "" so its CSV stays byte-identical; audit_isbns.py now reports null for a stored identifier that is not a 10/13-digit number (the old copy passed normalized junk through — junk matched nothing anyway). Requires cquarry 1.8.
  • Docs: README's --book row notes the publication date; CLAUDE.md records the rendering-over-cquarry-modules contract.

3.23.1 (2026-08-30)

  • Fix: scripts/fetch_library_codes.py's "Calibre is running" guard now matches every Calibre process NAME (anchored pgrep ^calibre: the GUI, calibre-debug, and the calibre-parallel job workers, whose comm truncates to calibre-paralle and could never match the old exact-name -x "calibre"). The probe still never reads command lines, so a concurrent Bindery sweep mentioning "Calibre Library" in its args cannot trip it. The 2026-08-27 "args false-positive" record was a misdiagnosis (name-only matching has been in place since 2026-08-09); the real defect was the missed calibre-parallel workers. Verified live against a running Calibre with four workers.
  • Chore: scripts/fetch_library_codes.py.bak and the orphaned audit_epub*.pyc removed (flagged as litter by the repo's own cleanup standard).
  • Upgrade: tests/test_version.py now also pins the newest patchnotes.md heading to cquarry_cli.VERSION, catching the patchnotes-vs-code desync class by CI instead of by eye.
  • Docs: spec header (3.15.0) and roadmap header (v3.12.0) current; spec §5's companion table gained the missing audit_conversion_overrides.py row; spec/CLAUDE.md cquarry floors moved to the real 1.7; README's "completed software / no new features" note rewritten (it contradicted the roadmap's open phases).
  • Dependency: Requires the latest cquarry (1.7.x).

3.23.0 (2026-08-28)

  • Feature: --set-pubdate BOOK_ID DATE and --clear-pubdate BOOK_ID — publication-date writes through cquarry 1.7's set_pubdate, which stores Calibre's exact TEXT convention ('YYYY-MM-DD 00:00:00+00:00', naive input taken as UTC) and clears via the 0101-01-01 undefined-date sentinel. No-op honest: re-setting the same instant doesn't bump last_modified or queue OPF resync. The TUI Edit Book menu gains the matching Set Pubdate / Clear Pubdate pair. Retires the raw-SQL pubdate workaround that cost 8 linter errors on 2026-08-27.
  • Upgrade: Multi-verb batch mode, on top of cquarry 1.7's batch() transaction context — several write flags in one invocation (cquarry --set-title 42 "New" --set-pubdate 42 1991-10-01 --add-tag 42 Curated) now run in ONE WritableCalibreDB inside one batch() transaction: all-or-nothing (a failure anywhere rolls back every verb, metadata_dirtied included), committed exactly once, with a per-verb ok: summary printed after the commit. Repeated --add-tag/--remove-tag flags join the same transaction instead of opening one per flag. --remove-book refuses batch combinations (exit 2). Single-verb invocations behave exactly as before.
  • Upgrade: writeops.py internals — each verb's mutation now lives in an action_* builder so the CLI dispatcher and the TUI executors share one implementation; argument validation (ids, rating range, cover yes/no, format size) happens up front for the whole invocation before the database opens.
  • Upgrade: Requires the latest cquarry (1.7.x).

3.22.0 (2026-08-28)

  • Upgrade: The TUI now delegates to vir-tui 2.2.0's Phase-3 primitives — interactive_session() owns the curses session lifecycle (open, degrade-to-text, close, KeyboardInterrupt → exit 130), prompt_float() replaces the hand-rolled rating loop (blank now means 0/clear rather than cancel), prompt_path() powers the first-run and change-database flows (existence loop with a notice built in; the change-database cancel path keeps the "database unchanged" behavior), confirm(danger=True) gates the Remove Book flow, and report footers use the public out_note().
  • Chore: Dropped the private _USE_CURSES/_SCREEN module globals in favor of vir-tui's public text_mode().
  • Dependency: vir-tui tracks @main; requires 2.2.0.

3.21.0 (2026-08-27)

  • Feature: --book BOOK_ID — full single-book dossier composing cquarry's read APIs: identifiers, per-format files with catalogued sizes and on-disk paths (missing files flagged), cover resolution, comments rendered through strip_html, custom columns (via the search-engine field() hook), e-reader annotations, per-device reading positions, plugin data, and conversion overrides. Unknown ids exit 1 cleanly.
  • Feature: --entities {authors,series,publishers,tags,languages,ratings} — cquarry's get_entities() listing with per-entity book counts and the sort/link secondary columns for authors/series/publishers (ratings render as stars).
  • Feature: --reading-progress — every last_read_positions row across devices with progress-fraction bars, newest first.
  • Feature: --columns — the custom-column schema (label, datatype, editability, normalized flag, enum values, composite templates) via get_custom_columns().
  • Feature: --info — library dossier: identity UUID, virtual libraries with their defining expressions, saved searches, @Name user categories, grouped search terms, news feeds, conversion overrides, metadata_dirtied/annotations_dirtied queue depth, and tag-browser layout state.
  • Feature: Write-verb expansion completing cquarry ≥1.5's write module — --add-tag / --remove-tag (repeatable; orphaned entity rows pruned), --set-identifier / --clear-identifier (EAV upsert, empty VALUE deletes), --set-series + --series-index / --clear-series (clear nulls books.series_index and prunes), --set-publisher / --clear-publisher, --set-languages / --clear-languages (canonicalized through Calibre's language map), --add-format / --remove-format (metadata-only data rows), and --set-cover (has-cover flag). All funnel through the shared dispatcher: argument problems exit 2 before the DB opens, lock/validation errors exit 1.
  • Upgrade: All write plumbing (previously inline in cli.py) moved to cquarry_cli/writeops.py so the TUI reuses the exact same executors; the TUI gains Display/Info entries, Annotations + LibraryThing exports, an Edit Book submenu covering every write verb, and a dry-run-first Remove Book flow.
  • Upgrade: Dependency policy — cquarry (and vir-tui) now tracked explicitly at @main; the committed uv.lock (which pinned cquarry 1.0.0 at an old commit) is removed and gitignored so every pip/uv/CI install pulls the latest cquarry.
  • Upgrade: Requires the latest cquarry (1.6.x).

3.20.0 (2026-08-26)

  • Upgrade: --search inherits cquarry 1.6's @Name:query user-category location — Calibre's get_user_category_matches parity (exact member matching per member location, @Name:.query for subcategories, false inversion, upstream's @...: lexer word rule so spaced category names work). No flags needed; unknown @Names match nothing instead of degrading to an all: text sweep.
  • Upgrade: Inherited v1.6 read-side completeness — get_book() rows are now shape-identical to get_all_books() rows (both carry uuid, identifiers, size), languages follow books_languages_link.item_order like Calibre, get_entities() gained a ratings kind, and get_feeds() / get_annotations_dirtied_books() / get_tag_browser_counts() (Calibre's own tag_browser_* sidebar rollups with avg_rating) are available for future verbs.
  • Upgrade: Requires cquarry>=1.6.

3.19.0 (2026-08-26)

  • Feature: Write-verb surface completed on cquarry ≥1.5's expanded module — --set-authors ID "A; B" (semicolon-separated; author_sort recomputed), --set-rating ID STARS (0–5), --set-comments / --clear-comments, --set-column ID #LABEL VALUE / --clear-column (layout auto-detected, enumerations validated against configured values, non-editable columns refused), and --remove-book ID [--confirm-remove] (dry-run by default printing title+formats; irreversible with the flag). All verbs queue OPF regeneration via metadata_dirtied.
  • Feature: --format-stats prints per-format book counts and total bytes (get_format_stats()).
  • Upgrade: All write verbs funnel through a shared _run_write dispatcher with uniform lock/validation error handling (exit 1) and argument validation (exit 2).
  • Upgrade: Requires cquarry>=1.5.

3.18.0 (2026-08-26)

  • Feature: --show-author-details — opt-in enrichment for --catalog, --all-wings, --export, and structured --search output. Catalog lines gain a {sort; link} segment and JSON/CSV exports gain author_sorts/author_links fields, sourced from cquarry ≥1.4's entity secondary columns (each author's true sort key and author-page URL).
  • Upgrade: --search now resolves custom grouped-search terms (GroupName:query, from Calibre's grouped_search_terms preference, with upstream union/false-inversion semantics) and the new annotations: location (full-text over e-reader highlights) — both inherited automatically via cquarry 1.4's engine; no flags needed.
  • Upgrade: Requires cquarry>=1.4.

3.17.0 (2026-08-26)

  • Feature: Page counts flow through exports — --export JSON gains a pages key per book, CSV a pages column, and the AI format a <N>p segment — sourced from Calibre's native books_pages_link table via cquarry ≥1.3.
  • Feature: Library provenance stamping — text catalogs (--catalog, --all-wings) carry the library's identity UUID in their header line, and --audit's summary names it, so any output can be traced back to its source library after moves/restores (cquarry get_library_uuid()).
  • Upgrade: The audit's cover checks now resolve through cquarry's get_cover_path() (canonical cover.jpg/cover.png layout logic) instead of hand-built paths.
  • Upgrade: Requires cquarry>=1.3.

3.16.0 (2026-08-26)

  • Feature: First write flow — --set-title BOOK_ID TITLE renames a book through cquarry's separate WritableCalibreDB module (trigger-safe: registers Calibre's title_sort/uuid4 SQL functions, refreshes the sort key, bumps last_modified). Every mutation also records the book id in Calibre's metadata_dirtied queue (requires cquarry 1.2), which is what upstream consumes to regenerate the book's sidecar .opf and re-push metadata to wireless readers on its next startup — external edits finally propagate. Requires closing Calibre first; lock contention fails cleanly with exit code 1.
  • Feature: --audit output now includes a "Pending OPF sync" section listing the books queued in metadata_dirtied (via cquarry's read-only get_dirtied_books()), so you can see exactly what Calibre will resync next time it starts. Absent when the queue is empty.
  • Upgrade: Requires cquarry>=1.2.

3.15.0 (2026-08-25)

  • Feature: New --export-annotations command dumps e-reader highlights, bookmarks, and notes as JSON (via cquarry's annotations reader), optionally scoped to one book with --id <BOOK_ID>.
  • Feature: New --plugin-data NAME flag (works with --catalog, --all-wings, and --search) appends third-party plugin values — e.g. goodreads_id or wordcount from books_plugin_data — as <name: value> segments on each book line.
  • Upgrade: Migrated every mode off the retired string-typed get_all_books() fields. authors, tags, formats, and languages are consumed as the native list[str] arrays cquarry 1.1+ exposes; author names containing literal commas no longer risk splitting.
  • Upgrade: Requires cquarry>=1.1. Search gains saved-search interpolation (--search 'search:"Name"'), multi-valued count operators (tags:#>2), language canonicalization (languages:English), slash date separators, tristate boolean keywords (checked/blank/...), strict errors on unknown virtual libraries, and new size:/pages: locations — all available through the existing --search flag with no CLI changes.
  • Note: scripts/reconcile_file_metadata.py already fetches per-book records without a full-library scan, so the roadmap's single-record fast path was satisfied by design; no change was needed.
  • Feature: New scripts/audit_conversion_overrides.py lists books carrying manual conversion recipes — cquarry’s extractor, which never unpickles the blobs.

3.14.1 (2026-08-24)

  • Fix: Replaced hardcoded custom columns logic in librarything.py with dynamic ID resolution from the database schema.
  • Fix: Restored functionality in test_queries.sh by targeting cquarry_cli instead of the extracted cquarry module.
  • Fix: Migrated spot_check.py and audit_isbns.py to use connect_ro() with WAL/SHM fallback for safe read-only locking against active Calibre DBs.
  • Fix: Corrected NULL title bug in spot_check.py by coalescing absent titles.
  • Fix: Fixed logical gap check in display.py series output.
  • Fix: Reordered argument validation in export.py to prevent 0-byte file truncations on invalid format flags.
  • Fix: Removed .title() enforcement during duplication checks in audit.py to preserve original casing.

3.14.0 (2026-08-24)

  • Refactor: Adapted to vir-tui v2.0.0 public API and decoupled menu fallbacks.
  • Fix: The non-curses fallback text menu now functions properly for CalibreQuarry by passing custom letter_keys and aliases during initialization.

3.13.0 (2026-08-23)

Changed

  • TUI Extraction (vir-tui): Extracted the generic CLI formatting (core.py) and curses menu primitives (tui.py) into the vir-tui shared repository. CalibreQuarry now depends on vir-tui for all UI logic, ensuring perfect parity and centralized updates for all interactive prompts across the workspace.

3.12.0 (2026-08-23)

Changed

  • Shared Library Extraction (cquarry): Extracted the core Calibre database reading layer and search expression grammar into a new, standalone Python library (cquarry). CalibreQuarry now depends on this shared library for all data access and search resolution, ensuring 100% parity across all tools in the workspace (like Hermitage and Wings).
  • Project Renamed to calibrequarry: To prevent pip namespace collisions with the newly extracted cquarry shared library, the CLI project has been formally renamed to calibrequarry in pyproject.toml, and its internal modules have been moved to cquarry_cli. The terminal command remains cquarry.

UI Upgrade: CLI scripts now feature rich output (ANSI formatting, tqdm progress bars, and a clear summary block). The project is no longer strictly stdlib-only and now depends on tqdm.

3.10.1 (2026-08-14)

Fixes

CI Configuration & Code Closures. The GitHub Actions pipeline (ruff check) failed because of several B023 late-binding closures inside tui.py which were unnoticed by the global configuration. Fixed those closures and added a test (test_version.py) to prevent version drift between pyproject.toml, config.py, and VERSION files.

3.10.0 (2026-08-11)

Features

LibraryThing Export Integration. Ported the standalone export_librarything.py script natively into cquarry. You can now use the --exportlt flag to generate LT-formatted CSVs directly. Crucially, this can be combined with --search to export only specific subsets of your library (e.g., --search "date:>2026-08-05" --exportlt), making targeted updates significantly easier. The export fully handles LibraryThing data quirks, such as folding ISBN-10 to ISBN-13, expanding translators into individual tags, clearing sentinel dates, and breaking output into manageable 500-book chunks split by "Read" vs "Unread" statuses.

3.9.2 (2026-08-09)

A bug, maintenance and improvement sweep across the package and all seven companion scripts. Eight fixes, three additions, four cleanups, every one pinned by a regression test. The suite grows from 243 to 273 tests.

Fixes

Every database opener built its file: URI by raw interpolation. ? and # are URI syntax, so a library at Books #2/metadata.db resolved to a different path entirely and failed with the thoroughly unhelpful "no such table: books". This affected db.py and six of the seven scripts; fetch_library_codes.py already percent-encoded its path, which is the form the rest have adopted via a shared db_uri_ro helper. The package's helper is in helpers.py; the standalone scripts each carry their own copy, as they carry everything else.

--tag scoping in audit_isbns.py and fetch_library_codes.py matched one arbitrary tag per book. Both pulled a single tag through a LIMIT 1 subquery with no ORDER BY, so a book carrying two tags was filtered on whichever one SQLite happened to return: a book tagged both Fic.Fantasy and NonFic.Tech was invisible to --tag NonFic.Tech roughly half the time, and fetch_library_codes.py's per-branch hit-rate report was attributing books to a branch at random. Scoping now considers every tag, and a book is reported under a tag the filter actually selected on. This is latent on the reference library, where all 7,439 books carry exactly one tag; it was found by reading the query, reproduced with a two-tag fixture, and fixed as insurance rather than to change any result observed so far.

compress_pdf.py --out-dir pointed at the PDF's own directory destroyed the original. The output lands at out_dir/<same name>, which in that case is the source path, so the mode documented as leaving the original untouched moved the compressed temp over it, left no .pre-compress.pdf rollback, and then printed "Original untouched at:" followed by the path of the file it had just replaced. It is now refused (exit 2), with the comparison resolving both sides so a symlinked directory cannot slip past.

write_all_wings silently overwrote colliding wing catalogs. Sanitizing a wing name to a filename is lossy: "Tabletop: RPG" and "Tabletop RPG" both reduce to Tabletop_RPG, and a punctuation-only name reduces to nothing, yielding _Library.txt. One wing's catalog replaced another's with no warning, leaving a file that claimed to be a wing it was not. Names are now de-duplicated with a numeric suffix, and a name that sanitizes away gets a positional fallback.

write_catalog could not write into a directory that did not exist yet. run_audit and run_export both create the parent first; write_catalog opened directly, so --catalog --output reports/catalog.txt exited 1 on a FileNotFoundError. It now matches its siblings.

spot_check.py --review crashed on an empty sample. With no books selected (--n 0, or a --limit that selects nothing) no chunk files were written, and the closing instructions indexed written[0]. It now reports an empty sample.

audit_drm.py and audit_epub.py had no locked-database fallback. The package, validate_metadata.py and reconcile_file_metadata.py all read from a temporary snapshot when Calibre holds the lock; these two let the OperationalError escape as a traceback, so a library-mode run while Calibre was open simply crashed. Both now degrade the same way, pinned by tests that take a real BEGIN EXCLUSIVE lock rather than faking the error.

Two writers used the locale's encoding instead of UTF-8. spot_check.py's report and bundle, and audit_drm.py's CSV, were the only file writes in the repository without an explicit encoding=; a title outside the locale's character set raises UnicodeEncodeError partway through a long run. Reachable only under a genuinely non-UTF-8 locale (en_US.ISO-8859-1 and the like) rather than the far commoner LANG=C, which auto-enables Python's UTF-8 mode: this is consistency with the rest of the repository more than a crash anyone was hitting.

Additions

--audit reports cover_file_missing. A book whose database row says has_cover but whose cover file is gone from disk was silently clean: the audit flagged a missing cover record and a low-resolution cover, but not the case where the two sources of truth disagree. Every cover consumer hits that, Calibre's own grid included.

The curses prompt accepts non-ASCII input. It read one byte at a time and discarded anything outside 32-126, so Brontë could not be typed into a search query or an output path. It now reads whole characters (get_wch), which for a library full of translated fiction is the difference between the TUI being usable for search and not.

Every external tool call is bounded by a timeout. spot_check.py (exiftool, djvused), reconcile_file_metadata.py (calibredb, exiftool, qpdf, djvused, pgrep), fetch_library_codes.py (pgrep) and compress_pdf.py's inspection probes previously had none, so a single wedged tool on a single damaged file hung a whole-library run with no indication of which book it stopped on. Each failure path is handled rather than merely raised: a hung reader flags that book and the run continues. Ghostscript itself is deliberately left unbounded in compress_pdf.py, because a legitimate 1 GB sourcebook takes many minutes and killing that mid-convert would be the wrong answer.

Maintenance

audit_epub.py extracted each book's rendered text twice under all: emptytext built it per spine document and ocr rebuilt the same string, which is the expensive half of a pass whose entire purpose is touching each EPUB once. It is now computed once per book and shared. Two stale documentation references to validate_library.py, a script that is not in this repository, now point at validate_metadata.py; audit_epub.py's usage line said "all three audits" when there have been four since v3.6.0. audit_drm.py imported sqlite3 inside a function while importing everything else at module scope, and stats.py carried a conditional whose two branches computed the same value.

3.9.1 (2026-08-08)

audit_isbns.py counted any labelled ISBN as the book's own. Books quote other books' ISBNs constantly, and one citation is indistinguishable from a self-identification if you only count numbers, so a handful of famous false accusations followed: The Atrocity Archives names The New Hacker's Dictionary's ISBN in a glossary entry, Metamagical Themas lists one among Hofstadter's self-referential joke titles, and C++ Primer Plus advertises six other Sams books in its back matter.

The fix is to require corroboration rather than to enumerate the ways a citation can look. A copyright page never carries a bare number: it sits beside a copyright line, a rights reservation, a binding, a printing statement, or a CIP block. A citation carries none of that, so an ISBN now counts as the book's own only when such a marker appears within 260 characters.

Applying that symmetrically was itself a bug, caught by measuring before committing. The first version demanded self-identification in both directions and cost 639 confirmations across a real library while removing only 25 false findings. The two directions need different evidence: a book printing the same number the catalogue holds is conclusive whatever the surrounding prose says, because a citation coinciding with your own stored value does not happen. Only a different number needs to have been claimed. Restricting the test to the negative direction gives 46 fewer false findings with confirmations slightly up (2,974 to 2,979).

Both directions are now pinned by tests carrying the real passages. 243 tests.

3.9.0 (2026-08-08)

A new companion script, audit_isbns.py, and the first new capability since the v3.8 sweep. It answers a question nothing else in the Calibre ecosystem asks: does the ISBN stored against a book actually identify that book? Calibre downloads metadata but never re-examines what it stored, so a wrong ISBN stays invisible, and an ISBN is what other systems key on when you hand them a catalogue.

The motivating evidence: a four-source sweep of a 6,786-ISBN library that was already validator-clean found 51 identifiers pointing at a different book. The dominant shape is a same-publisher sibling, which is why the defect survives every existing check: the number is well-formed, the checksum passes, and only the book it names is wrong. Programming Clojure carried tmux 2's ISBN, Spelunky carried Super Mario Bros. 3's, and A Book on C carried 9782147483649, the 2147483649 integer-overflow constant dressed as an ISBN.

This release ships the offline half: verification against the ISBN each book prints on its own copyright page. That is the best authority available for exactly the books no bibliographic database has heard of (small-press RPGs, indie ebooks, print-on-demand reprints), and it needs no network, no credentials, and no new dependencies.

It reads body text only, never embedded metadata. reconcile_file_metadata.py writes the database's values into those metadata blocks, so comparing against them would be comparing the database with itself and would cheerfully confirm every error the tool exists to find. A test pins this: an OPF carrying a matching identifier must not produce a confirmation.

Not crying wolf is most of the work. Three benign things resemble a mismatch and are classified apart. A bibliography prints other books' ISBNs (The Art of UNIX Programming prints 49), so above --max-printed distinct numbers a file is read as a citing work. A bundle or series volume legitimately prints several ISBNs, reported AMBIGUOUS for a human to resolve rather than guessed at. A format variant (print versus ebook) differs only in its final digits, so a printed number sharing the stored one's registrant prefix is reported VARIANT rather than MISMATCH; a genuinely wrong ISBN almost always comes from a different publisher block entirely.

There is deliberately no --apply, and there will not be one. Single-source verdicts proved wrong often enough during the motivating sweep that an auto-fixer would have "corrected" Curse of Strahd, Cold Mountain, Kitchen and The Master and Margarita, every one of which was already right.

Scoping reuses fetch_library_codes.py's anchored-hierarchical --tag rule rather than inventing a virtual-library flag: a comma-separated prefix list covers a multi-root wing without coupling a standalone script to the package's VL resolver. Text extraction is stdlib zipfile for EPUB and optional pdftotext/djvutxt for PDF and DJVU, whose absence is reported rather than fatal. MOBI/AZW3 are skipped.

Three things were corrected before merge, all found by running it against a whole 6,783-ISBN library rather than a single wing.

The severe verdict was renamed SUSPECT to MISMATCH, because the old name overclaimed. On a real library the commonest reason a book prints a different ISBN than the catalogue holds is that the file is a different edition, which is worth knowing but is not a wrong book. An ISBN alone cannot separate the two, and a report that says "suspect" 100 times about mostly-benign edition drift trains its reader to ignore it.

The publisher-prefix comparison was too long. Registrant lengths run 2 to 7 digits and are inversely proportional to publisher size, so a nine-digit prefix (right for a one-book press) split HarperCollins from itself: Sabriel's stored 978-0-06-447183-1 and printed 978-0-06-000548-1 are one publisher and were being reported as rivals. Six digits reclassified 28 findings from the severe bucket to VARIANT.

UNREADABLE conflated two opposite things. Half of those books were MOBI/AZW3, which this tool skips by design; the rest were supported formats that yielded nothing, which may be damaged files. They are now SKIPPED and UNREADABLE respectively.

Also documented: the printed ISBN can itself be wrong, which is the limit of the whole premise. The TSR Forgotten Realms Campaign Setting prints a number belonging to The Jungles of Chult (a permanent typo), and Night Witches prints Durance's because a small press reused its previous copyright page. Both surfaced as VARIANT and both resolved in favour of the database, which is exactly why nothing is ever written automatically.

Runs: 121 books in 15s on one wing and 468 in 63s on another, both with zero severe findings; 6,783 in about 15 minutes across the library. The suite grows from 208 to 237 tests, with the real-world classifications pinned so a future refactor that silently reclassifies them fails loudly.

3.8.1 (2026-08-07)

A full-repository bug, maintenance, and documentation sweep: the package, all seven companion scripts, the tests, and every doc. The package core came out clean (one micro-refactor: _num_predicate in search.py returned a two-tuple whose second element nothing read; it now returns just the predicate). The scripts yielded nine real fixes, every one now pinned by a regression test. The suite grows from 195 to 208 tests, and fetch_library_codes.py and validate_metadata.py gain their first tests.

Fixes

audit_drm.py reported an unparseable PDF as CLEAN. qpdf --is-encrypted exits 0 for encrypted, 2 for not encrypted, and 3 for a file it cannot parse; the code collapsed 2 and 3 into "clean". A corrupted or truncated PDF (exactly the kind of loose file the tool exists to vet) therefore skipped the trailer /Encrypt fallback that was built for the couldn't-classify case, and a Standard-encrypted broken file read as definitively DRM-free. Exit 3 now defers to the fallback and reports BENIGN encrypted-unclassified when an /Encrypt dictionary is present.

audit_epub.py content contradicted itself on an injected signature in a declared-foreign book. The expected-foreign flag (declared language / NonFic.Language.* tags) was applied to injection-signature hits too, so a piracy notice in a legitimately-French book printed under a green "(expected-foreign)" label, reported "0 file(s) need review", and still exited 1. A signature is a defect regardless of language; it is now always counted and listed as needing review.

compress_pdf.py skipped verification when it mattered most. page_count() returns None both when pdfinfo is missing and when pdfinfo cannot parse the file, and the page-count comparison only ran when both counts were known. A Ghostscript output so broken pdfinfo could not read it (the strongest possible bad-conversion signal) therefore passed "verification" and replaced the original. When the original's page count is known and the output's is not, the run now aborts with the original untouched.

reconcile_file_metadata.py split author names on commas. The file-side author string was split on &, ;, and ,, but ebook-meta only ever joins authors with &; a comma belongs to the name itself. "Martin Luther King, Jr." parsed as two bogus authors, never matched the database, and the book reported as drifted forever, surviving every re-embed. Same shape as the v3.8.0 identifier-space bug, one field over. The split now uses & and ; only.

spot_check.py --record silently truncated comma-delimited notes. A verdict line's note field was parts[4] of an unbounded split, so a comma-mode note of "wrong author, should be Jane Doe" recorded as "wrong author" with the rest dropped and no error. The split now caps at five fields, keeping free-text notes intact.

spot_check.py --record no longer runs without --against. The id reconciliation is documented as the load-bearing part of review mode ("nothing is written unless the ids reconcile"), but --against was optional, and omitting it skipped the check entirely, accepting exactly the short-but-plausible verdict lists it exists to refuse. Recording without an ids file is now a setup error (exit 2).

fetch_library_codes.py clobbered its own backup. The pre---apply backup was stamped with the date only, so a second run the same day overwrote the first run's restore point, the copy that actually holds the pre-change database. The stamp now carries seconds plus a collision counter; no backup is ever overwritten. Also fixed: a missing metadata.db exited 1 via sys.exit(str) while the documented contract (and every other setup-error path) says 2.

validate_metadata.py inflated FORMAT_FICTION_PDF counts. The check joined through the tags table before filtering, producing one warning per (book, tag) pair; a crossover book tagged Fic.Fantasy and Fic.Horror was reported twice. Now one warning per book, listing every matching tag.

audit_epub.py's latin-1 decode fallback was dead code. bytes.decode("utf-8", "replace") can never raise, so the documented fallback was unreachable and the except path re-read the same corrupt entry just to fail again. A corrupted spine entry now reads as empty text through one honest path, and the docstring says so.

Maintenance

  • test_queries.sh exercised vl:"The Tabletop", a wing that no longer exists (split into the two "Tabletop:" wings in the live library), so the VL-resolution smoke was silently testing the unknown-VL path. It now targets Fantasy Wing.
  • All seven scripts are executable now (five carried a shebang without the bit; the documented python3 scripts/... invocation is unchanged).
  • Docs drift closed across the board: spec.md §5 gains the missing fetch_library_codes.py row and the --review half of spot_check.py (and now says three scripts write, not two); roadmap.md gains Phase 7 recording the v3.7.0–v3.8.0 companion work; the README's test-suite section describes all seven test files instead of the original two; CLAUDE.md's architecture tree adds audit_drm.py, spot_check.py, modes/tags.py, and the five test files it didn't list, and renames the long-gone audit_epub_content.py to audit_epub.py.

3.8.0 (2026-08-02)

New companion script fetch_library_codes.py: derive Library of Congress Classification codes from the LoC SRU catalogue and store them as identifiers. Written after the existing Calibre plugin for this job, "Library Codes - SRU", was diagnosed as unable to do it at all.

The plugin fails two independent ways. It refuses composite custom columns outright (library_codes_dialog.py validates cc_datatype != "text" and clears the active flag), which is fatal if you store the code as an identifier and project it with {identifiers:select(lcc)}. And its ISBN lookup queries the LCDB index dc.identifier, which the server has never supported: it answers with SRU diagnostic 1/16, 1007 (Bib-1 114 Unsupported Use attribute), "Unsupported index". The index that resolves an ISBN there is bath.isbn, confirmed live against lx2.loc.gov:210 alongside dc.title and dc.creator, which do work and are what the plugin's author/title fallback happens to use. So the plugin's primary path, the one its own description advertises, has been dead while the fallback quietly carried it.

Storing to identifiers rather than to a column value is the deliberate difference. Identifiers are the catalogue's canonical home for external keys, a composite column displays the value with no second copy to drift, and reconcile_file_metadata.py already carries identifiers into embedded file metadata. Books that already hold an lcc are skipped unless --refresh is passed, so the tool is naturally incremental against a growing library.

Two operational facts shape the design. The Library of Congress rate-limits harder than the plugin's 1.6s pacing suggests: at 0.6s between requests it began resetting connections after roughly twenty queries, and a clean 60-query run at 2.0s posted zero errors. Pacing therefore defaults to 2.0s with exponential backoff, the run aborts after eight consecutive failures rather than hammering a service that has stopped answering, and every result including a miss is cached to ~/.cache/cquarry/library_codes.json. That cache is what makes a multi-hour pass survivable: an interrupted run resumes for free, and a repeat of the same 12 books fell from 30s to 0.09s.

Coverage is partial and worth knowing before committing hours to it. Measured on a 7,362-book reference library, a clean 60-book random sample of ISBN-bearing books resolved 18, or 30%. The average is misleading, because the distribution is steep: NonFic.Tech hit 5/7, while Fic.Fantasy managed 3/16 and Fic.Contemporary, Fic.Horror, Fic.Translated and Gaming.TTRPG returned nothing at all across the sample. LoC catalogues academic, technical and canonical trade titles well and genre fiction, indie and small-press releases, translations and tabletop material poorly. The dry run is therefore the default and reports a per-branch hit rate, and --tag scopes a pass by anchored-hierarchical tag prefix so the dense half can be run without paying for the sparse half.

--apply backs metadata.db up to the sibling .backups directory before writing and refuses to run while Calibre is open. INSERT OR REPLACE is correct here specifically because identifiers is UNIQUE(book, type), unlike the enum custom-column link tables where the same statement would append a second row.

New: spot_check.py --review, a judgement mode for the checks no pattern can make. The mechanical lints in that script decide whether a field has the wrong shape. They cannot decide whether a title is the right title, whether an author field holds the person who wrote the book, or whether a description describes this book rather than another one. Review mode puts those three questions in front of a reader in a form that can be answered in bulk: full title, authors, context line and complete description, emitted in numbered chunks with a matching .ids file, answered with a verdict file, and accumulated in a ledger.

The ledger is what makes it worth doing more than once. Reviewed books drop out of later samples, so repeated runs converge on full coverage while the random sampling that makes a partial pass statistically meaningful is preserved. --worklist prints the BAD entries as a punch list; nothing is ever written back to metadata.db.

Recording refuses unless the verdict ids reconcile exactly with the ids emitted, in both directions, and refuses again on any verdict word that is not OK or BAD. This is the load-bearing part rather than a nicety: a reviewer working through hundreds of records drops some, and a short list looks identical to a complete one. On the reference library the same failure has now been recorded five times against agent output, once during the run that motivated this mode. Six tests cover the drop, invent, malformed, roundtrip, worklist, and chunking paths; the suite grows to 70.

Also fixed: reconcile_file_metadata.py truncated any identifier value containing a space, and reported the book as drifted forever. parse_identifiers split the ebook-meta output on [,\s]+, commas or whitespace, so lcc:BF637.S4 G63 2007 parsed as lcc = "bf637.s4" with the rest silently dropped. The parsed value never matched the database, so the book was reported as drifted no matter how many times it was successfully re-embedded. The split is now on commas alone, which is the format ebook-meta actually emits.

This is a latent bug rather than a new one, and it went unfound because nothing could reach it: every identifier type in use until now (isbn, goodreads, storygraph, google, amazon, oclc) is space-free. Library of Congress call numbers are the first that are not, so writing lcc identifiers is what exposed it. Measured on the reference library, a reconcile pass over 655 freshly embedded books reported 469 of them as still drifted before the fix and 1 after, that one being an AZW3, which genuinely cannot carry arbitrary identifier types. Tests grow to 69.

One deviation worth recording: the XML comes from a plain-HTTP endpoint and is parsed with xml.etree.ElementTree. defusedxml is the conventional hardening and is not an option in a stdlib-only project, so the response is instead capped at 8 MB before parsing, which bounds the entity-expansion exposure without a dependency. ElementTree does not resolve external entities, so there is no XXE path.

3.7.1 (2026-07-31)

spot_check.py reported a complete book as EPUB_EMPTY_SPINE when the package used the legacy OEB 1.0 namespace. check_epub resolved the manifest and spine with a hardcoded {"o": "http://www.idpf.org/2007/opf"}, so any package declaring http://openebook.org/namespaces/oeb-package/1.0/ instead (OverDrive-era conversions) matched nothing at all: no manifest, no spine, and therefore a HARD failure and a nonzero exit code on a book that opens perfectly. Found by the first full-library pass, which flagged exactly one hard failure across 7,339 books, #8048 Dying Inside: 31 content documents, 446,083 characters of body text, a spine listing every one of them, and a checker that could not see any of it.

Manifest items and spine itemrefs are now matched by local element name through a _by_local_name helper, so the package's declared namespace stops mattering. This is the same root cause as the known audit_epub.py emptytext false positive on legacy OEB files; that analyzer is untouched here and still carries it.

Tests grow to 64 (an OEB 1.0 package resolves its manifest and spine).

3.7.0 (2026-07-31)

Three spot_check.py correctness fixes and one new advisory flag, all found by running the checker against the 7,339-book reference library and then auditing what it did not catch.

Fixes

_MOJIBAKE missed the commonest lead byte of all, â. The pattern enumerated individual accented characters (Ã[©¨¤¶¼£±]), which covers é/ã but not â, the double-encoded form of a curly apostrophe and by far the most frequent mojibake in scraped blurbs. One live case was missed on the reference library (#2658 Observability Engineering, youâ??re doing). Replaced with [ÃÂ] followed by anything in U+0080-U+00BF: that band is never valid text, because a real Portuguese à is followed by an ASCII vowel (Ãvila) and never by latin-1 supplement punctuation. Verified against ten cases including São Paulo, café society and Ãvila, which must stay clean.

The comment, title, and author lints read raw markup instead of text. lint_comment stripped tags but never decoded HTML entities, so a description of &amp; repeated forty times measured 200 characters and cleared the 120-character stub gate on 40 characters of real content (one live case). The same blindness hid any mojibake stored in entity form. A new plain_text() helper strips tags, drops script/style bodies, decodes entities, and normalizes non-breaking spaces; the lints and the review bundle's blurb excerpt all route through it, so every check now sees what a reader sees.

OPF hrefs were resolved without URL-decoding. Manifest hrefs are percent-encoded per the EPUB spec, so a content document whose filename contains a space arrives as %20 and never matches the zip namelist, which reads as a missing spine item and therefore a HARD failure and a nonzero exit. No book in the reference library trips it (checked all 4,890 EPUBs: zero false positives cleared by decoding), so this is a latent correctness fix rather than an observed one, but the failure mode is silent and severe enough to close. Fragments are stripped before resolution as well.

New

Advisory COMMENT_TRUNCATED: a description that stops mid-word. The motivating defect, found 2026-07-30: a metadata download returned four of six Jacqueline Carey blurbs cut off mid-sentence (Blessed Elua founded Terre d'A, Raphael de Mereliot, her manipul). Nothing caught them, because every existing check passes: the text exists, is well-formed, and is long (977 to 2,265 characters). Only the final word gives it away. A whole-library sweep then found roughly fifteen more.

The check is deliberately gated and deliberately advisory. It needs a wordlist and reads /usr/share/dict/words on exactly the terms the PDF and DJVU checks already use for exiftool and djvused: used when the system has it, silently skipped when it does not, never a Python dependency. Descriptions legitimately end without punctuation all the time (blurb attributions, series lists, contents dumps), so the heuristic requires the final word to be lowercase, absent from the wordlist, ASCII, preceded by whitespace, in prose of at least fifteen tokens, and not part of a URL or a numbered contents line.

Measured honestly on the reference library: 22 flagged out of 7,339, of which 6 are confirmed truncations (27% precision, 43% recall against a hand-built set of 14). The residual false positives are systematic and worth knowing before trusting a flag: modern technical vocabulary absent from an old wordlist (microservices, autoscaling, lifecycle, asyncio), and truncations whose final fragment happens to be a real word (the co, on the st) are missed entirely. It is a lead generator for the judgment pass, not a verdict, which is why it is not in HARD and does not affect the exit code.

Tests grow to 63 in tests/test_scripts.py (mojibake lead-byte coverage with the Portuguese and French negatives, entity decoding and the stub-gate skew it caused, and truncation detection with its proper-noun, URL, and complete-prose negatives).

3.6.0 (2026-07-03)

New Features

audit_epub.py gains a fourth analyzer: ocr, flagging OCR/conversion-damaged prose. The defect class none of the existing three catches: a bad conversion splits paragraphs mid-sentence at line-wrap or page-break positions, so the text reflows broken ("could just make out the shape" / "of another boat"). The motivating case was a damaged Jingo EPUB that measured 80 such splits where a clean edition of the same text measured 0. The primary signal is exactly that split shape (a paragraph ending without terminal punctuation, the next starting lowercase, paired only within one spine document), reported as a per-book rate normalized by paragraph count. all runs it inside the same single decompression pass as the other three.

The hard problem was separating damage from intentional style, and rate alone cannot do it: deliberately unpunctuated literary prose (Fosse's Septology, Evaristo's Girl, Woman, Other, Kingsnorth's The Wake, Faulkner) posts split rates far above genuinely damaged books. The discriminator that works is where the fragment ends. Style breaks at clause boundaries ("...she started out in theatre"); damage breaks at line-wrap positions, which land on function words ("sat the disembodied" / "heads who were..."). On the reference library every style book measured at most 11% of splits ending on a function word and every hand-confirmed damage case at least 26%, so the flag gate requires 25%, alongside minimum-splits, minimum-paragraphs, and rate floors. The function-word set is a small closed list, the same spirit as the content analyzer's stopword votes, not a dictionary.

Five false-positive idioms found during validation are guarded explicitly: paragraphs interrupted by a rendered figure (an inline formula or card-diagram image reads as a split otherwise; _Blocks now records when an image falls between two blocks' text), display math set as text (mostly non-alphabetic fragments), back-of-book indexes rendered as paragraph blocks (the "See also" signature), epistolary sign-offs (a dangling short unterminated fragment), and the block-quotation idiom of academic prose (pairs into or out of a <blockquote>, plus attribution fragments ending on "that"). Secondary signals are reported but never gate: en-dashes embedded inside words (bottom–feedin'), doubled opening quotes (' 'Course, with the first quote required to follow whitespace so British single-quote dialogue's close-then-open sequences stay invisible), and space-stripped proper nouns recurring alongside their hyphenated form (AnkhMorpork vs Ankh-Morpork).

Hand-validated against the full 4,605-EPUB reference library, every flagged book inspected: 105 flagged, 104 confirmed damage, one borderline residue (display quotes publisher-styled as plain paragraphs, indistinguishable without CSS). Documented out of scope: character-substitution errors ("sonic" for "some") need a wordlist the stdlib-only contract rules out, and damage whose signature is word truncation or whitespace corruption rather than paragraph splitting. Tests grow to 60 in tests/test_scripts.py (split detection, dialogue-fragment and scene-break non-splits, image-interrupted pairs, style-vs-damage discrimination, threshold boundaries, and an all run including the new analyzer).

3.5.0 (2026-07-02)

The persistent-curses-screen rework, closing the last item in the roadmap's "Port from the Lattice TUI audit" section (Lattice T7, shipped there as v4.10.0). Purely a lifecycle change; no menu, prompt, or mode behavior differs.

Changes

One curses screen per session. Every menu, prompt, pause, and pager used to be its own curses.wrapper init/teardown, so multi-prompt flows visibly flashed to the shell between widgets. interactive_menu now opens the screen once and every widget draws into it (_with_screen); a widget invoked outside a session still gets its own one-shot wrapper, so nothing changes for direct callers. _reset_terminal becomes a no-op while the session screen lives (running stty sane under curses would undo cbreak/noecho beneath it). Measured under a pty: a full menu, prompt, Esc-cancel, mode, pager, quit session enters the terminal's alternate screen exactly once, where the same flow used to enter it once per widget.

One degradation path. A curses failure at session startup or mid-session (unknown terminal, capability lost) funnels through _degrade_to_text: the screen is suspended once and the rest of the session runs the text fallback. All the v3.3.2 guarantees hold (no stuck terminal, no silent exit 0); a pager that dies mid-display now also degrades and prints the mode's output as plain text instead of eating it. Ctrl-C at the menu (curses or text fallback alike) ends the session cleanly with exit code 130 and the screen restored, while EOF at the text menu stays a quiet Quit.

tests/test_tui.py grows to 29 cases; the session lifecycle was additionally verified end-to-end under a pty (alternate-screen count, degraded startup on an unknown TERM, and a full cancel-then-run flow against the live library).

3.4.0 (2026-07-02)

The rest of the Lattice TUI audit ports (roadmap section "Port from the Lattice TUI audit": T2, T4, T6, and the fallback-menu generation). Lattice shipped all of these in its v4.9.0; this release keeps the two shared curses skeletons aligned.

Changes

Esc in a prompt now cancels back to the menu instead of accepting the default (behavior change, port of Lattice T2). In menus Esc has always meant back; in prompts it meant "accept the default", so mis-selecting a mode and mashing Esc launched it with all defaults. A cancelled prompt now unwinds the whole prompt chain back to the menu with nothing launched; bare Enter still accepts the default, and Ctrl-C or EOF at a prompt cancels the same way (the text-fallback prompts used to exit the whole program on Ctrl-C). The hint bar now reads "Enter Accept, Esc Cancel, Ctrl-U Clear". Cancelling the first-run prompt exits without persisting anything (exit code 1, the CLI's no-database signal, so unattended runs never read as success), and cancelling "Change database path" leaves the saved path untouched.

Fixes

Output-file prompts expand ~ (port of Lattice T4). ~/reports/x.txt used to create a literal ./~/reports/ directory, since the TUI has no shell to expand it. Every output path the TUI collects now goes through a shared _prompt_out() that expands the tilde but does not absolutize, so relative paths keep meaning the current directory. The results pager also gains a footer naming the resolved absolute path, so "where did my report go" answers itself; the footer is suppressed when the mode errored or was cancelled, so it never claims a file that was not written.

The no-curses fallback menu is generated from the same sections the arrow-key menu renders (port of Lattice's v4.8.1 fix). The numbered listing and its key map were hand-maintained twins of _MAIN_SECTIONS, exactly the pattern that silently desynced in Lattice (fallback keys dispatching the wrong modes). _build_fallback now derives both from the sections, the word aliases ("catalog", "stats", ...) stay as an explicit supplemental dict, and tests pin that every entry is reachable with its number matching its label.

Widget paper cuts (port of Lattice T6). A bad answer at a number prompt re-asks with a "not a number, try again" note instead of silently using the default. On terminals shorter than the menu, the box shifts so the selected row stays visible instead of clipping blind below the bottom edge. The pager gains horizontal panning (arrow keys or h/l) with ellipsis markers on lines that continue off-screen, computes its width once instead of on every keypress, and keeps the scroll position valid across resizes. Ctrl-U clears a prompt field, so editing a long pre-filled path no longer means backspacing through all of it. And the export format prompt validates its answer against json/csv/ai, re-asking instead of accepting any string.

tests/test_tui.py grows from 7 to 24 cases, pinning all of the above.

3.3.2 (2026-07-01)

Fixes

TUI hardening ported from the Lattice TUI audit (2026-07-01). cquarry/tui.py shares its curses skeleton with Lattice's; the two carry-overs from that audit's high-severity findings land here (roadmap section "Port from the Lattice TUI audit", items H7 and H6's exception-boundary half). First: a curses init failure no longer reads as Quit. On capability-poor terminals (TERM=vt100, dumb terminals) the color setup or curs_set raised curses.error, which the menu loop treated as the user quitting, so the TUI silently exited 0 even though the text fallback menu works. Cosmetic capabilities are now non-fatal (a monochrome TUI beats a dead one), and a real curses.wrapper failure flips the session to the text fallback menu instead of exiting. Second: _run_with_capture now has an exception boundary. A mode error used to escape as a raw traceback and lose the captured output; it is now paged under an [Error] heading with the traceback plus whatever was captured, and Ctrl-C pages a [Cancelled] notice the same way. New tests/test_tui.py (7 cases) pins both behaviors. The remaining Lattice carry-overs (Esc-cancels-prompt, ~ expansion in output prompts, generated fallback menu) stay on the roadmap until their Lattice counterparts land.

3.3.1 (2026-06-30)

Fixes

audit_drm.py no longer flags a freed EPUB on a leftover marker file. v3.3.0 treated the mere presence of META-INF/rights.xml (Adobe ADEPT) or sinf.xml (Apple FairPlay) as DRM. But those are token/voucher files, not the lock itself: the actual lock is content encryption, which a DRM'd EPUB records in encryption.xml against its XHTML. When a book is freed, the content is decrypted but the marker can stay behind, so a bare marker with no content encryption is a residual artifact, not a locked book; it reads and embeds fine. The first whole-library sweep surfaced exactly one such case (Warhammer Helsreach: a sinf.xml, no encryption.xml, 37 plain-XHTML chapters), which is the same residual-artifact shape as the PDF that motivated the tool. EPUB classification now keys on actual content encryption (encryption.xml with non-font entries) and names the scheme from whichever marker is present; a standalone marker is reported BENIGN as a "residual DRM marker". PDFs are unchanged: a residual handler dictionary there still breaks metadata embedding, so it is still flagged. With this fix the library's real DRM count is 48 (all recoverable PDF ADEPT dictionaries), with the lone FairPlay EPUB correctly cleared. Tests extended to 21 cases (residual markers benign; markers plus encrypted content still DRM).

3.3.0 (2026-06-30)

New Features

audit_drm.py: a cross-format DRM scanner (read-only). The metadata and structural audits never look at encryption, so a DRM-locked file can pass epubcheck, report its page count, and even import, yet silently refuse to let its embedded metadata be rewritten. The case that prompted this was a z-library PDF carrying a residual Adobe ADEPT EBX_HANDLER dictionary that qpdf --check and pdfinfo both reported as "not encrypted" while exiftool choked on it, failing the reconcile embed. The new script classifies EPUB, PDF, and Kindle (MOBI/AZW3) files; DJVU has no DRM scheme and is reported N/A. It runs in library mode (formats and paths from metadata.db, opened strictly mode=ro) or directory mode (a recursive scan of loose files before import, for the pre-import battery), with the usual exit codes (0 clean, 1 DRM found or scan error, 2 setup error).

The design priority was not detecting encryption; it was not crying wolf. Two benign things look like DRM to a crude check and are explicitly cleared. Font obfuscation: an EPUB META-INF/encryption.xml that scrambles only the embedded fonts (the IDPF or Adobe #RC algorithms) is not a content lock; because publishers often name obfuscated fonts fonts/00001.dat with no font extension, an entry is cleared when it uses a font-scrambling algorithm OR targets a font resource (extension or a fonts/ path), which an algorithm-only or extension-only check gets wrong. Permission flags: a PDF "encrypted" with the Standard handler and an empty user password opens with no password and is only flagged against printing or copying, so it is classed with qpdf as PERMISSIONS, not a lock. Real DRM is rights.xml (Adobe ADEPT) or sinf.xml (Apple FairPlay) in an EPUB, a content-encrypting encryption.xml, a non-Standard PDF security handler found by a streaming byte scan (so a residual or inactive dictionary is still caught, which is exactly the Irodov case), a password-locked PDF, or a non-zero Mobipocket encryption-type field.

The first whole-library sweep (6,651 files) found 50 DRM-locked files (49 Adobe ADEPT, 1 Apple FairPlay), with 135 benign font-obfuscation/permission cases correctly cleared and one early false positive fixed before release: five EPUBs whose Adobe #RC font obfuscation targets fonts/*.dat were initially misread as encrypted content, which is what drove the algorithm-or-target rule above. Ships with a unittest suite (tests/test_audit_drm.py, 19 cases) building zip, PDF-byte, and PalmDB fixtures for each format and verdict.

3.2.1 (2026-06-25)

Fixes

reconcile_file_metadata.py --repair-pdf now deletes the .~qpdf-orig backup qpdf leaves behind. qpdf --replace-input, used to rebuild a broken cross-reference table before re-embedding, writes the pre-repair original to <name>.~qpdf-orig beside the file and never removes it. Across many reconcile passes these full-size copies accumulated inside the library tree, which is the worst place for them: Calibre scans that tree, and each one is a complete duplicate PDF. A sweep of one library turned up 21 such files totalling 403 MB. embed_pdf now unlinks the backup as soon as qpdf reports success (return code 0 or 3), before retrying the embed; a missing backup is a no-op, so the change is safe whether or not qpdf wrote one. Regression tests mock the exiftool/qpdf boundary to assert the backup is removed after a successful repair and that an absent backup does not raise. Pre-existing strays from older runs are not cleaned by the tool; remove them once with fd -H '\.~qpdf-orig$' "<library>" -X rm.

3.2.0 (2026-06-23)

New Features

audit_epub.py emptytext now flags partial / placeholder exports (new PARTIAL verdict). The whole-book character count missed a failure mode: a DRM-locked or sample export where most chapters are an identical "content unavailable" placeholder while one or two real chapters carry enough text to clear the THIN floor, so the book validates, repairs clean, and reads as full-length to the old total-char check. The canonical case was a BookShout export of Johannes Cabal: The Fear Institute, where 15 of 17 chapters were the same 138-char "something went wrong loading... bookshout.com" stub; even after structural repair to zero epubcheck fatals it stayed a 2-chapter sample. The analyzer now flags PARTIAL when a known DRM-placeholder signature appears anywhere in the spine, or when the same short stub (12 to 600 chars) repeats across at least 3 spine documents and at least 30% of the spine. PARTIAL is a real defect (counts as FOUND, exit 1, needs re-sourcing), distinct from the advisory THIN. The false-positive guard is the per-document distribution: a well-made book full of small but DISTINCT section dividers does not trip it (only repeated-identical stubs do), so the three full novels in the batch that surfaced this stayed OK. Library and directory modes both report it.

3.1.1 (2026-06-23)

Fixes

audit_epub.py now percent-decodes spine hrefs, fixing false EMPTY verdicts. OPF manifest hrefs are IRIs, so a content document whose archive filename contains a reserved character (commonly !, written %21; Sigil and calibre emit these routinely) was matched against the raw zip namelist undecoded, failed to resolve, and dropped out of the spine. A text-full book whose every chapter file had such a name resolved to zero readable spine documents and was reported EMPTY: the exact false positive hit on Martha Wells's The Serpent Sea (every split_NNN.html was named CR!RT...). The resolver now decodes the percent-encoding (UTF-8, with multi-byte runs decoded together) and strips any #fragment before matching the namelist. Stdlib-only via a small re-based decoder (_pct_decode); no urllib dependency added. The fix lands in the shared spine resolver, so all three analyzers (content, pagenumbers, emptytext) benefit. Regression tests cover the decoder (reserved char, multi-byte UTF-8, invalid escape) and an end-to-end encoded-spine EPUB.

3.1.0 (2026-06-20)

Changes

The three EPUB-content audits are merged into one scripts/audit_epub.py. audit_epub_content.py, audit_epub_pagenumbers.py, and audit_epub_emptytext.py shared the same spine resolution, library/directory dual-mode, read-only contract, and exit codes, and differed only in the per-book verdict; they are now three analyzers behind one tool, selected by subcommand: audit_epub.py content|pagenumbers|emptytext|all [directory]. The detection logic of each is unchanged (same thresholds, same results). Two wins beyond removing the duplicated scaffolding: all opens each EPUB once and runs all three analyzers in a single decompression pass (the expensive part is decompression, so this is much faster than three separate full-library runs), and there is now one spine resolver to maintain instead of three slightly-diverging copies. The old script names are removed; update any caller to audit_epub.py <mode>. The --min-chars / --thin-chars knobs (emptytext) carry over.

3.0.3 (2026-06-20)

New Features

scripts/audit_epub_emptytext.py: find empty / no-body-text EPUBs. Catches the failure mode every other audit misses: a content-less stub that still validates. The canonical case is the "Bookmate" export, where the archive holds only cover and promo images plus a tiny HTML placeholder, the OPF spine points only at that placeholder, and the book itself is absent. Such a file passes epubcheck, "repairs" clean in a structural repairer (its one referenced document is well-formed), and shows no foreign text to audit_epub_content.py because there is no text at all; the metadata looks perfect. The detector resolves the spine, drops <script>/<style>, strips tags, decodes entities, and counts the rendered characters: EMPTY (<=2000, --min-chars) is a real defect to re-source, THIN (<20000, --thin-chars) is advisory because a genuine short story or a publisher sample can also land there. A Bookmate origin is not itself a defect; most Bookmate exports carry their full text, so the flag is on empty content, not provenance. Library mode (DB-driven, mode=ro) and directory mode (vet downloads before import), mirroring the other two EPUB audits. First full-library run flagged four real defects (a Bujold stub, two image-only scans with no text layer, and a Draft2Digital sample of a full novel hiding in the THIN tier) against three genuinely short stories left alone.

Fixes

spot_check.py no longer flags OCaml and NCurses as case garble. The intercaps allowlist (_CASE_OK) now includes OCaml and NCurses alongside SQLite, QBasic, and the rest, so legitimate library titles stop tripping the advisory case-garble heuristic.

3.0.2 (2026-06-14)

New Features

scripts/audit_epub_pagenumbers.py: find print page numbers baked into EPUB body text. Bad PDF/OCR-to-EPUB conversions capture the print page number (and often the running header) as a literal paragraph in the flow instead of real EPUB pagination, so it reflows into the middle of a sentence ("where the hay cart 16 was taking him"). The detector reads each book's blocks in spine order and flags a number only when it genuinely interrupts prose: a lowercase continuation after it, a word split across it (the previous block ends in a hyphen), or it abuts a repeated running header/footer. Section and chapter numbers (which open the next block with a capital) and endnote/footnote numbers and chronology years are left alone. Library mode (DB-driven, mode=ro) and directory mode (vet downloads before import), mirroring audit_epub_content.py. Validated by hand against the full reference library: 21 flagged, every one a true positive; the false-positive tail (an experimental footnote-poem, a scraped web-serial's vote counts, placeholder section labels) all fell under the hit-count or book-span floors. It also surfaces piracy watermarks and bad OCR scans that ride along with the page-number cruft.

3.0.1 (2026-06-12)

New Features

scripts/spot_check.py: randomized metadata + file-integrity audit. Samples N random books (reproducible with --seed) and checks what pattern-based sweeps miss: title corruption and mojibake, junk author entries, missing or stub descriptions, EPUB archive integrity (CRC, container/OPF sanity, spine completeness, text volume), PDF header/page count, and DJVU page count. Emits a machine-readable flag report plus a review bundle (title/author/tag/series/blurb per sampled book) for a human or LLM judgment pass. Read-only against metadata.db; validator-owned checks are not duplicated. First 600-book run against the live library caught a wrong-book description, a truncated description, three mojibake descriptions, and an EPUB with ten dangling spine references.

Fixes

reconcile_file_metadata.py PDF embeds no longer fail silently. Two related defects: exiftool refuses to rewrite XMP packets containing duplicate properties (seen in the wild: a doubled prism:doi) unless -m is passed, and it exits 0 with "files unchanged", which the tool read as success; The embed now passes -m. (An interim attempt to also write a bare -Publisher tag was reverted: on PDFs exiftool maps it to the same XMP dc:publisher bag, so double assignment appended duplicate entries and broke the round-trip.)

write_catalog no longer corrupts the shared book cache. It sorted the list returned by get_all_books() in place, silently reordering the session cache for every later consumer. It now sorts a copy; a regression test pins the behavior (tests/test_modes.py).

compress_pdf.py cannot clobber a rollback original. If a .pre-compress rollback file already exists, an in-place run now aborts instead of overwriting the only copy of the true original with an already-compressed file.

audit_epub_content.py finds the library from either home. The library root resolves to wherever metadata.db sits: next to the script (the copy living inside the library) or the current working directory (running the repo copy from a library), in that order.

reconcile --id rejects malformed id lists cleanly via a dedicated parser instead of an unhandled ValueError.

Internals

Exception chaining (raise ... from) throughout search.py and helpers.py; unused loop variable removed in wing-overlap analytics; re.Scanner access satisfied for type checkers; new test coverage for the backup guard, library-root resolution, and cache isolation (suite: 87 tests).

3.0.0 (2026-05-26)


Breaking

Python 3.14+. The supported floor moved from 3.9 to 3.14 to match the development environment. The code does not depend on bleeding-edge syntax, but only 3.14+ is tested and supported.

New Features

Comprehensive search engine. The search expression parser was rebuilt as a dedicated, stdlib-only engine (src/cquarry/search.py) that ports Calibre's grammar and matching semantics. It now resolves field locations beyond tags and authors: series, publisher, rating, formats, languages, pubdate/date/last_modified, identifiers/isbn, comments, cover, id, uuid, and #custom columns, in addition to tags, authors, all, and vl:. It supports contains/=exact/~regex/^accent match kinds, numeric and date relational operators (rating:>=4, pubdate:>2015, date:30daysago), and field:true/field:false presence tests. Boolean grouping, implicit AND, quotes, and escapes follow Calibre's grammar, evaluated with its candidate-set semantics. The previous build only handled tags:, author(s):, and vl:; other prefixes silently matched nothing.

--search prints to stdout. With no --output, results stream to the terminal instead of forcing a file. --format json|csv|ai emits the matching books in that structured shape; otherwise a plain-text listing is produced. An empty query (--search '') returns the whole library, matching Calibre.

Deeper cover audit. Cover dimension reading no longer stops at the first 1 KB, so a JPEG whose SOF marker sits behind a large EXIF/ICC block is measured correctly; PNG covers are now read too.

Fixes

Interactive TUI analytics no longer crashes. Selecting any item under the ANALYTICS menu (Author statistics, Reading pace, Tag tree, Wing overlap) raised a NameError because those functions were never imported into tui.py. They now work.

Half-star ratings are visibly distinct. A 2.5 rating rendered identically to 2.0 (both ★★☆☆☆); it now shows ★★½☆☆ using the universally available ½ glyph.

Series "complete" is computed correctly. Completeness now means "no missing integer volumes" rather than "book count equals the top index", so a series with novellas (0.5) or duplicate editions is no longer wrongly marked incomplete.

Portability of the series rollup. get_all_series was rewritten to aggregate in Python instead of using GROUP_CONCAT(... ORDER BY ...), which required SQLite 3.44+.

Normalized custom columns now load. Single-valued text and enumeration custom columns (e.g. a "Status" / reading-state column) are stored by Calibre in a value table plus a books_custom_column_N_link table, exactly like multi-valued columns. The loader keyed off is_multiple and tried to read a book column straight from the value table, which errored and silently returned nothing. It now detects the link table, so --show-custom and #column searches work for every custom-column type.

Documentation & Tests

Honest parity claims. The README and spec no longer claim "100% parity"; they document exactly which locations and operators are supported and the deliberate, dependency-bound deviations (stdlib re instead of the regex module, unicodedata folding instead of ICU, no GPM templates or saved-search references, and cquarry's anchored hierarchical tags: match).

Companion scripts. scripts/compress_pdf.py (Ghostscript-based PDF shrinking with verify-or-rollback; writes files and metadata.db) and scripts/audit_epub_content.py (read-only EPUB content auditor) are now versioned alongside the toolkit and fully documented, explicitly outside the read-only cquarry package contract.

Portable test suite. tests/test_search.py and tests/test_helpers.py cover the parser grammar (adapted from Calibre's own tests), the matcher against an in-memory provider, a full-stack integration test on a temporary SQLite fixture, and the rating/series/image helpers, all without needing a live Calibre library.

2.6.0 (2026-05-03)


New Features

Tag Dump (--tags). A flat, alphabetized list of every tag in the library with its book count, written to stdout. Drop-in replacement for the noisy calibredb list_categories -r tags shell pipeline — pipe it to a file with cquarry --tags > tags.txt. Also reachable as "Tag dump" under LISTS in the interactive TUI. Honors --quiet (suppresses header/footer, leaves the body intact for scripting). Distinct from --analytics tags, which renders the hierarchical tree.


2.5.0 (2026-04-21)


New Features & Fixes

Full Parity Search Engine. Refactored the --search expression parser to achieve 100% parity with Calibre's native syntax.

  • Author Searching: Added full support for author: and authors: prefix tokens.
  • Fallback Text Search: Un-prefixed terms (e.g. cquarry --search "author:Anne Rice") now correctly fall back to searching anywhere across book titles, authors, and tags. This accurately mimics Calibre's implicit boolean AND handling of unquoted spaces.
  • Complex Grouping: Verified and documented support for nested boolean logic, parenthetical grouping, and negative lookaheads (e.g., NOT(tags:Fic.Romance OR tags:Fic.Contemporary) or tags:"Fic.Fantasy.Grimdark" AND author:"Phil Tucker").
  • Test Suite. Added an automated test suite (tests/test_search.py) mapped against Calibre's actual SearchQueryParser behavior to guarantee ongoing expression fidelity.

2.4.1 (2026-04-16)


Completed Software & Bug Fixes

Completed Software Status. Phase 4 has been concluded, and CalibreQuarry is now considered feature-complete and stable. It has undergone rigorous end-to-end testing against real-world Calibre databases.

Bug Fixes:

  • Fixed a NameError crash in --audit mode where the newly introduced color helper was not imported, preventing the final summary from printing when issues were found.

2.4.0 (2026-04-16)


New Features

Custom Column Support. Added a --show-custom "Column Name" flag that extracts data from user-defined custom Calibre columns. The values are automatically appended to text catalogs and are natively included in JSON, CSV, and AI exports. Color CLI Output. Introduced simple, lightweight ANSI color formatting for headers, warnings, and error highlights across CLI modes to improve readability when bypassing the interactive pager.


2.3.0 (2026-04-16)


New Features

Extended Audit Checks. The --audit mode has been significantly expanded to include three new checks:

  • Duplicate Detection: Identifies books with identical titles and primary authors across the library.
  • Cover Quality Audit: Scans the actual JPEG cover files on disk (without external library dependencies) and flags covers with low resolution (below 500px on their longest edge).
  • Format Migration Report: Flags books that are only available in deprecated legacy formats (MOBI, LIT, LRF, DJVU, PDB, AZW).

2.2.0 (2026-04-16)


New Features

Extended Analytics. Added a new --analytics argument with four detailed reporting modes: author (per-author breakdowns of formats, ratings, and series), pace (books added per month/year trend), tags (hierarchical taxonomy tree visualization), and overlap (virtual library wing overlap analysis). These are also accessible via a new ANALYTICS section in the interactive TUI.


2.1.0 (2026-04-16)


New Features

Search Query Export. You can now pass arbitrary Calibre search expressions directly to the CLI via --search "query" to export matching books. The results are written to a plain text file. This feature is also accessible via the interactive TUI under the OUTPUT menu.

AI-Readable Export. Added a new ai format to the --export option. This format outputs the library data as a highly token-efficient, flat text list designed specifically for LLM ingestion and recommendation prompts (e.g., Title by Author [Tags] - Rating/5).


2.0.1 (2026-04-12)


New Features

Top Authors and Top Tags in --stats. The statistics output now includes a "Top authors" section (10 most prolific by book count) and a "Top tags" section (15 most-used tags), inserted between the ratings distribution and the tag taxonomy breakdown.


2.0.0 (2026-04-12)


Major Overhaul: Package Restructure & TUI Upgrades

CalibreQuarry has been refactored from a single ~1450-line monolithic script (cquarry.py) into a proper Python package architecture, with TUI improvements modeled after the Lattice project.

Layer-Based Package Design. The codebase now lives in src/cquarry/ and is split by logical functionality: config.py, db.py, helpers.py, cli.py, tui.py, and a modes/ directory for individual feature operations (catalog.py, stats.py, audit.py, display.py, export.py). The monolithic script is gone.

Modern Build System (Hatch). CalibreQuarry now uses pyproject.toml managed by Hatch. Install via pip install . or pipx install . and the cquarry command is available globally. Also runnable via python -m cquarry.

Persistent Database Configuration. Both CLI and TUI now share a unified database resolution chain: explicit --db flag, saved config (~/.config/cquarry/config.json), default search paths, then an interactive prompt if running in a TTY. The path is saved on first successful resolution, eliminating the need to pass --db in future sessions. A "Change database path" option under the SETTINGS section in the TUI main menu allows updating the stored path.

Calibre Lock Handling. When metadata.db is locked by a running Calibre instance, CalibreQuarry now automatically copies the database (including WAL/SHM journal files) to a temporary snapshot and reads from that instead. A notice is printed to stderr, and the temp files are cleaned up on exit. Previously, a locked database would produce an unhandled sqlite3.OperationalError.

Fully Immersive TUI. All operations now run through a _run_with_capture() wrapper that intercepts stdout and stderr via io.StringIO buffers. Output is displayed within the scrollable curses pager rather than dropping the user back to raw terminal output. This matches the immersive TUI pattern established in Lattice v4.1.2.

Styled Curses Pause. The post-operation "Press Enter to continue" prompt now renders inside a styled Unicode box within the curses session (_tui_pause), instead of falling through to a raw input() call. Accepts Enter, q, or Esc to dismiss.

Null Byte Sanitization. The scrollable pager now strips null bytes from captured output before rendering, preventing ValueError: embedded null character crashes on corrupted data.


1.0.4 (2026-04-08)


New Features

Curses TUI. Running the script with no arguments now launches a full-screen, arrow-key navigable terminal UI utilizing the curses library (matching the interface of getMusic). Non-TTY environments or systems without curses will gracefully fall back to a styled text-based box menu. The TUI features a custom scrollable pager that intercepts standard output, allowing you to comfortably read and navigate command outputs directly within the interface.

Bug Fixes

Implicit AND in VL expressions ignored subsequent tags. Calibre's parser evaluates adjacent tags like tags:Fic.Fantasy tags:Magic as an implicit AND. The _parse_and method previously discarded tags after the first one unless the AND keyword was explicitly written out. It now correctly intersects all implicit constraints.

Exact tag matching (=) was case-sensitive. Calibre's tag matches are always case-insensitive. The exact match SQL query (WHERE t.name = ?) was missing COLLATE NOCASE, causing tags:"=Fic.Fantasy" to fail if capitalization varied.

Duplicate author headers in --primary-only mode. Generating a catalog with --primary-only caused highly fragmented author groups. The script relied solely on the SQL ORDER BY b.author_sort (which sorts by the full multi-author string). Books are now presorted in Python natively by their derived primary-only display key.

Non-deterministic GROUP_CONCAT output. The metadata fields built via GROUP_CONCAT(DISTINCT ...) (like authors and tags) returned unpredictably ordered results depending on SQLite's internal row execution. This occasionally resulted in the wrong primary author being selected. The SQL query has been rewritten to use correlated subqueries with explicit ORDER BY clauses for deterministic structure.

Double NOT cascading crashes. Expressions combining consecutive exclusions (e.g., NOT NOT vl:Name) failed because _parse_not routed directly to _parse_atom for the inner operand. This has been updated to recursively call _parse_not to handle complex nested negations gracefully.

Fractional series indices ignored in recent display. show_recent dropped series identifiers completely if the index contained a decimal (e.g., 1.5) due to a missing fallback formatting block.


1.0.3 (2026-04-04)


Bug Fixes

_read_value crashed on unmatched quotes in VL expressions. expr.index('"') raised ValueError if a virtual library search expression had an opening quote with no closing quote. A single malformed VL definition took down the entire tool. Now consumes the rest of the string as the value.

Series index 0.0 silently dropped. if idx and idx == int(idx) treated 0.0 as falsy — any book at series position zero lost its series display. Hit in both write_catalog and show_recent. Changed to explicit is not None checks.

Division by zero in show_stats. Empty library crashed on three separate lines: format bar chart (count * 40 // total), rating bar chart (max(rating_counts.values())), and unrated percentage (unrated * 100 / total). All three guarded.

max_index None dereference in show_series. s['max_index'] == int(s['max_index']) threw TypeError when max_index was None. Same fix propagated into detect_series_gaps.

format_stars produced garbage on corrupt ratings. A DB value of 12 (6.0 stars) yielded a negative empty count. Python silently returns "" for "☆" * -1, so no crash — but the display was meaningless. Rating now clamped to 0–5.

CSV export blanked zero-valued fields. stars or '' and series_index or '' used the or pattern, which treats 0.0 as falsy. Changed to explicit is not None checks.

JSON export had leading whitespace in split fields. b['authors'].split(',') on GROUP_CONCAT output produced ["Author1", " Author2"]. All split fields now strip.

Tokenizer keyword boundary missed underscores. isalnum() doesn't match _, so a hypothetical tag starting with or_ or not_ would misparse as a boolean operator. Boundary check now includes underscore.

_prompt_str displayed [None] in interactive prompts. When called with default=None, the user saw the literal text [None]. Now shows empty brackets.

Performance

get_all_books() results cached. The 8-JOIN metadata query was called once per wing in --all-wings mode — 18 times against a 3,800-book library. Now fires once and returns the cached list.

_get_all_book_ids() cached for NOT operations. The VL parser queried SELECT id FROM books on every NOT clause. Multiple NOT expressions in a single VL definition hammered the DB. Cached on first call.

count_books() uses warm caches. If the books or IDs cache is already populated, returns len() instead of hitting SQLite.

Parser built once in main(). Was constructing build_parser() to parse args, then building it again on the help-output fallthrough path. Stored the reference.

New Features

--version flag. Uses action="version" so argparse handles it during parse_args() — works even when no database is present. Version also shown in the interactive menu banner.

Code Hygiene

Unused imports removed. defaultdict and Path — imported, never referenced.

f-string with no placeholders. f" Languages:" → " Languages:".

show_wings caught bare Exception. Narrowed to ValueError, which is what resolve_vl actually raises.

quiet parameter wired up everywhere. show_recent, show_series, and show_stats all accepted quiet but ignored it. Now suppresses headers and decorative output when passed.

main() catches PermissionError. A read-only DB with wrong filesystem permissions previously produced an unhandled traceback.

_match_tags docstring corrected. Said "containing" for the non-exact case but the SQL does prefix match, not substring. Added a note that regex patterns (tags:~regex) are unsupported.

prog="cquarry.py" added to ArgumentParser. Version and help output now show the script name consistently regardless of invocation path.