_stamps_from_embeddedcuts the author line at the[bracket BEFORE splitting on " & " (the 2026-09-26 HFT 2nd-ed EPUB): calibre renders N authors asA & B & C & D [SortA & SortB & SortC & SortD], the bracket wrapping the WHOLE list, and the seeder split the display line before cutting it, so a correct 4-author stamp seeded 7 authors: the four display names plus sort-form fragments riding along as phantom authors in the manifest seeds. The OPF and the stamp_epub verifier were both correct in the live case, and the hand-corrected signed manifest stands. The fix is the display-segment shape stamp_epub's_display_authoralready compares for single authors: the single-authorName [Sort, Form]form seeds exactly 1, and the Unknown-placeholder drop still works in both its bare and bracketed forms.- stamp_pdf recovers the exiftool parse-crash class mechanically, exactly once (the 2026-09-26 C++23 STL Cookbook PDF): when the exiftool write exits nonzero WITH a Perl-space parse error in its output (today's shape:
Can't find Root objectoff a malformed catalog Names array, on a fileqpdf --checkpasses), the script attempts oneqpdf --replace-inputrebuild whose page count must be preserved (qpdf missing, an unreadable side, or a changed count restores the backup copy already made and no retry runs), then retries the stamp ONCE, re-reading the raw calibre author_sort property after any rebuild because the rebuild rewrote the XMP packet. A retry that still fails lands in the existing STAMP_FAILED path, and the stubborn-XMP class (write reports success, read-back disagrees) is untouched and keeps failing immediately. The live file was rebuilt by hand first, 496 pages preserved, the retry verified. - Docs and tests: both roadmap boxes are ticked with the fix notes; CLAUDE.md gains the 3.54.0 contract-notes section; the phase-1 library skill names the one-recovery exception in its stamp_pdf prose. New tests (9): the seeder's 4-author bracket cut, single-author suffix cut, and Unknown drop; and the parse-crash battery (rebuild verified and retried once, retry-still-failing STAMP_FAILED with no second rebuild, a page-count change restoring the backup byte-for-byte and failing without retry, missing qpdf and non-Perl failures never recovering, and the
_is_parse_crashshape table). 681 tests, one skip on ebook-meta-less environments.
scripts/stamp_epub.pyships the EPUB sibling of stamp_pdf.py (the open roadmap box; the 2026-09-22 fiction wave ran the EPUB half on a one-off batch driver, 46 EPUBs stamped and three verifier iterations to learn the identifier spellings): oneebook-metainvocation writes the fixed field set (--title/--authors/--publisher/--isbn); Calibre's own writer rewrites the OPF in place and REPLACES any existing isbn identifier, which IS the correction mechanism for wrong embedded ISBNs, so dc: metadata is never hand-edited. Dry-run by default;--applyrequires a--backup-dirrefused when it resolves inside any target file's directory; repeated--authorflags join with " & " (ebook-meta's own separator;cquarry --set-authorssplits on;instead, a CLI-layer difference only). The tool-boundary checksum gate rides stamp_pdf 3.52.0's shape: an--isbnfailing cquarry'sisbn_check_digit_is_validis refused (exit 2) in dry-run and apply alike, before anything writes.- The verify contract is the EPUB-specific one the live exercise cost three iterations to learn. Title/authors/publisher verify from the
ebook-metaread-back, the author compared as the display segment before the[bracket (live files renderName [Sort, Form]). ebook-meta's read-back never displays ISBN, so the ISBN verifies by reading the OPF directly (zipfile -> the container's rootfile entry ->dc:identifierelements) under VALUE EQUALITY: each identifier's value AND id attribute is normalized by dropping an optional leading urn:/isbn: prefix and then every non-alphanumeric character, and the stamp verifies when any candidate equals it. Bareisbn:978...values,opf:scheme="ISBN", namespacedns2:scheme="ISBN", dashed values, and the ISBN carried in the element id all verify; a file that ALREADY carried the right ISBN makes--isbnan honest no-op (the writer never destroys a matching identifier, it may add its own canonical one alongside) and verifies, never fails; when no--isbnwas requested the ISBN check is skipped entirely. - A failed stamp STOPS the list (STAMP_FAILED, exit 1, never re-fought): the per-file report names the mismatched fields, every not-yet-stamped remainder is named, and the
--json FILEmachine report records every target exactly once (dry_run/stamped_verified/stamp_failed/left_unstamped), the bijection record the batch driver used to assert by hand. A target named twice is refused at the boundary (exit 2): one file, one stamp. - Docs and tests: the roadmap box is ticked with the live notes; spec § 5 gains the stamp_epub row and the write-capable closing list names it; CLAUDE.md gains the 3.53.0 contract-notes section. New tests (29): the verify contract (bracket display segment, value equality over every live shape, wrong-ISBN failure, no-ISBN skip), the normalizer table, the OPF reader (value + id attribute as candidates, the container rootfile honored over the entry name), the guard set (checksum refusals dry and apply, backup-dir refusals, dry-run invokes nothing, missing/non-EPUB/duplicate targets), the stop-and-report flow, and a real-ebook-meta seam class (skips on ebook-meta-less machines) pinning the wrong-ISBN replace, the four already-right no-op shapes, and the clean-file gain (672 tests, one skip). The phase-1 library skill's EPUB half now drives this tool instead of raw
ebook-meta; the skill update ships in the same release. Erratum on the 3.52.0 entry below: its closing NOTE said the phase-1 skill's duplicate-screen prose was left for Brandon; it was in fact rewritten by ZCode that same evening, so nothing is owed there.
The duplicate-screen semantics fix: the 2026-09-21 waves' false refusals and the silent masked duplicate, plus the ISBN and pubdate guards
- screen_duplicate classifies title pairs instead of only comparing scrubbed equality (roadmap 2026-09-16 and 2026-09-19 boxes, driven by the 2026-09-21 mixed wave):
classify_titlesreturns a verdict per pair, and the four failing shapes now each get the right one. Differing REAL subtitles over one shared base aredistinct("Introduction to Computer Organization: ARM" vs "An Under-the-Hood Look at Hardware and x86-64 Assembly", two books the scrub collapsed onto each other and refused twice in one day). Differing DECLARED volume annotations (a token word plus an arabic or valid roman numeral on both sides) arevolume_sibling: the Moral Letters vols, the roman "Vol I" vs "Vol II" pair that was a false duplicate until now, and "Monster Vault" vs "Monster Vault 2" (separated by a trailing-ordinal check over the post-colon subtitle). Containment at a colon boundary (one full title equal to the other's pre-colon base or post-colon subtitle) isrelated: the Arcana base-vs-companion shape within the batch, and the masked DUPLICATE the old equality silently passed against the library ("Mothership: Wages of Sin" vs the bare "Wages of Sin" download, caught on 09-21 only by a manual collision sweep). Related candidates surface in the report (batch_related,related_hits), count toward exit 1, and are recorded in the phase-1 manifest underchecks.related_works; volume sets land asbatch_volumes/checks.volume_siblings, informational only. Neither class ever sets a verdict: the screen advises, the reviewer judges. The regression floor is untouched:normalize_titlekeeps the Capital scrub exactly, the roman-only annotation still matches the full-subtitle edition, and the 3.25.0 arabic-signature gate stays byte-for-byte (a "Book 8" next to the bare series title remains a silent gap-fill, so the 19-candidate Wandering Inn flood stays impossible). - stamp_pdf refuses an
--isbnfailing its check digit at the tool boundary (roadmap 2026-09-19 box):_check_isbnruns cquarry'sisbn_check_digit_is_valid(10- or 13-form, separators tolerated) and exits 2 before anything stamps, in dry-run and apply alike, so the guard lives where the write happens instead of in every caller's batch script. Correction found while pinning the test: the Hewitt "valid form" 1-59863-503-5 named in the roadmap box ALSO fails its checksum and is refused like every other bad number; that ISBN needs re-sourcing, not re-typing. - The phase-2 metadata download writes pubdate at date-only precision (roadmap 2026-09-16 box, recurred 2026-09-21):
_apply_opftruncates every OPF date toYYYY-MM-DDbeforeset_pubdate, because downloaded metadata carries no trustworthy time-of-day and the downloader was landing2010-05-25 01:49:00.233821+00:00-shaped values (42 of 47 books in one batch shared the same minute-band, all normalized by hand in phase 3). Bare years and exotic forms ride through unchanged, failingset_pubdateexactly as before; reconcile is untouched. - Docs and tests: the four roadmap boxes are ticked with fix notes; spec § 5's screen_duplicate and stamp_pdf rows name the new behavior; CLAUDE.md gains the 3.52.0 contract-notes section. New tests: the classifier case table (all three 2026-09-21 cases plus the Moral Letters, Arcana, and roman-volume shapes), the masked-duplicate and Monster Vault library fixtures, within-batch cross-reference classes, the seam's advisory notes, phase-1's advisory recording without refusal, the ISBN guard's refusals and passes, and the pubdate truncation end to end through
_apply_opf(643 tests, one skip on ebook-meta-less machines). NOTE FOR THE LIBRARY SKILLS: the phase-1 skill's duplicate-screen prose describes the old normalizer behavior and needs one line updated (related candidates and volume sets now surface in the report and manifest instead of nothing) — left for Brandon, outside this repo.
The bindery structured-fix records adoption: lossy consent reads data, with a version gate on the PATH binary
_mirror_lossyclasses repairs from bindery's structured fix records, not summary substrings (the 2026-09-18 box on bindery-cli's roadmap, fulfilled by its v0.45.0): every per-book record in bindery's phase-1 report now carries thefixesdict plusncx_uid_synced/watermark_refusalsas data, so the lossy/structural split is judged byfixes.get(marker)over the same five keys bindery's own gate treats as lossy strips. The rendered summary stays in the sealed record as the display line thelossy_consentdetail quotes. The practical gain: the class no longer depends on bindery's summary vocabulary staying stable; a summary that names no marker string still classes correctly from its data.- A PATH bindery below 0.45.0 is a hard error, never a silent downgrade:
_bindery_phase1probesbindery --versionand refuses below_BINDERY_MIN_VERSION(0.45.0) with the upgrade command named, because the data-driven classing against an older report would class every strip as structural (the vacuous-consent hole again). The probe result is also why this is a minor release: the consent gate's contract now includes a versioned report shape. - Landed from the preservation branch: the work was written in a working tree a parallel release lane had to sweep aside mid-run on 2026-09-21; it survived on
bindery-version-gate-wip(snapshot 9135493) and was cherry-picked onto main after 3.50.0 shipped. That branch's WIP message warnedtest_dry_phase1_emits_lossy_consent_decisionswas broken; the warning was stale (the pre-fix state from the collision window), the snapshot carries all five migrated fixtures, and the full suite runs 622 green against the branch tree, verified in an isolated checkout before landing. The branch is deleted; the content lives on main.
The calibre-touched PDF fix: stamp_pdf erases the stale author_sort that failed verify and would have poisoned imports
- stamp_pdf clears a pre-existing XMP
calibre:author_sort(roadmap 2026-09-21, found by the Effective C wave): any PDF Calibre has ever produced or touched carries that extension, ebook-meta renders the author asDisplay [Sort]from it, and the read-back disagreed on author even when every written field took (De Anima, How Software Works in that batch). The failure was never merely cosmetic: calibre'screate_book_entryhonors an embeddedmi.author_sortverbatim instead of recomputing from the stamped authors, so the stale sort would have landed in the library on import; the verify-side tolerance idea was discarded for exactly that reason. exiftool cannot write the calibre namespace (not in its tables, and a user-defined-configtable never associates with the packet parse), so the erase rides calibre's own writer: when the stamp sets authors andexiftool -s3 -Author_sortshows the property, a follow-up value-preservingebook-metapass with--author-sort ""(empty is null, so calibre writes no sort of its own) regenerates the XMP packet without ANY calibre-namespaced element, stale sorts, timestamps, and ratings included; the docinfo Keywords that carry the ISBN survive the rewrite, verified end to end on a calibre-touched repro (red before the fix, green after, clean XMP read). The erase re-carries exactly the stamped values, runs after the exiftool write so a failed write still leaves the file untouched, and its own failure is a STAMP_FAILED, not a blind continue._verifystays strict: the file is made clean, the comparison is not loosened. - Docs and tests: tests pin the erase argv (stamped values re-carried, the null sort, the file path last), the skip when no sort is present, and the erase-failure path; the two exiftool call-count assertions now filter for the write invocation since detection shells out too (619 tests). README gains a "How this compares" section, and the mixed-wave roadmap findings were recorded without code changes (screen_duplicate title-prefix containment, the stamp_pdf ISBN checksum idea, reviewer verdict-flip propagation, an EVERY_BOOK_COVER allowlist) and stay open. A stopped agent's bindery version-gate work was preserved unreleased on the
bindery-version-gate-wipbranch.
The roadmap findings batch: honest DJVU coverage, embedded-metadata stamp seeds, an in-batch cc6 check, and a lossy consent that only gates real strips
- audit_drm's N/A verdicts reach the runner (the DJVU box, observed 2026-09-17): the directory- and library-scan loops
continued past N/A verdicts (DJVU has no DRM scheme in practice) BEFORE the CSV append, so the manifest recordedchecks.drm = "unscanned"for a format that was judged. N/A rows now land in the CSV (still out of the counters and the summary), the manifest saysN/A, and N/A remains no quarantine reason. The instrument test is re-pinned from the old "unscanned" contract to the N/A one, and the phase-1 skill's coverage note matches the tool again. - Phase-1 stamps seed from the file's embedded metadata (the seeding box):
_stamps_from_embeddedreadsebook-meta's display output and merges title/authors/publisher/pubdate/language per-field over the filename parse; ISBN is deliberately never seeded (embedded identifiers are misidentification-prone). Guards that keep a bad seed from failing phase 2 or beating the decided conventions: a pubdate that does not parse as ISO is dropped (cquarry's add_book raises on a bare year), a read whose stderr carries calibre's traceback returns nothing (calibre exits 0 on unparseable files and prints its own reversed-order filename guess),Unknownplaceholders are dropped, a missing ebook-meta binary degrades to filename-only with a warning, and a per-file read failure is recorded in the entry'srepairs. The skill's step 5c is rewritten around the new seeding; the correction pass is now usually a verify. - The phase-2
#sourcestamp is verified in-batch (the cc6 boxes): inside the onebatch(), a fresh book's firstset_custom_columnmust report changed, or the import rolls back and the verb exits 1 with the library unwritten. The accompanying audit is an erratum on the observations: both signed manifests carry ZEROimported_ids (phase 2 saves the resume record immediately after the import), so neither batch went through the verb; the DB actually shows the reviewed provenance landed correctly on 77 of 79 books, with the two mismatches being the manifest-None files (one picked up Z-Lib, one Anna's Archive) and no all-Anna's-Archive state present. The out-of-verb door is outside the verb's reach; the effect-flag check is the in-verb half (a true read-back would need a cquarry custom-column reader and is recorded as a future option). - lossy_consent fires only on real lossy strips (the consent box, observed 2026-09-17 and 09-19):
_mirror_lossystill records every gate-accepted bindery repair in the sealed record, but each repair now names its class ("lossy": true/false, judged by_LOSSY_REPAIR_MARKERS, the five fix keys bindery's own gate treats as lossy: stripped_pagination, stripped_broken_tags, stripped_watermarks, dropped_marker, stub_docs_dropped), and only lossy-marker repairs set the file'sflagged. Structural-only files stop firing vacuous consent decisions, and the decision's detail carries the lossy summaries. Phase 2's consent re-drive flips and refuses over every recorded repair, lossy or structural, because bindery's re-drive applies all of them. The cleaner long-term home (a structuredfixesdict in bindery's phase-1 report) is recorded on bindery-cli's roadmap. - Docs and tests: CLAUDE.md gains the 3.49.0 contract-notes section; the roadmap's five findings boxes are ticked with the audit results; the phase-1/phase-3 library skills and bindery-cli's roadmap are updated in the same release (the sweep rule). New tests: the ebook-meta seam parse and phase-1 seeding merge/fallback paths, the in-batch stamp rollback, the structural-record flip and refusal on consent re-drive, the lossy class mirrors, and the real audit_drm CSV round-trip (615 tests).
The wave-2 refactor and consent batch: one dest-list source, one backup helper, and the lossy consent moves into the manifest
- The parallel write-dest lists have one source (roadmap "new work noticed", L2.14):
src/cquarry_cli/dests.pynow holdsSINGLE_BOOK_DESTS(writeops), set mode's target sources and--batch-*verbs (setwrite), and the restrict-refusal aggregateWRITE_FLAG_DESTS(restrict). The three consumers import their slices, and tests/test_dests.py pins every member against build_parser(), so a dest that exists only in a list (or only in the parser) cannot rot. - One backup helper:
backups.make_backupreplaces the triplicated_backup_db/_make_backup(run phase 2, the integrate verbs, set mode's--apply). The recorded error-mapping decision: a--backup-dirinside the library is a usage problem (exit 2) at every door, so the shared helper raises the stdlib-neutral ValueError and each dispatcher keeps its own usage path; unwritable destinations and sqlite failures raise ValueError too (set mode already wrapped them into its usage path; run/integrate previously propagated a traceback). Same timestamped sqlite-API backup, same on-disk shape. - bindery's
manual_watermark_repairdecisions mirror into the manifest'sdecisions_neededasmanual_repairentries (_mirror_bindery_decisions, the sibling of 3.43.0's_mirror_lossy): books bindery refuses to auto-strip used to vanish from the durable record entirely. - The lossy-pending double-manifest wrinkle is closed: a dry phase 1 now emits a
lossy_consentdecision per lossy-flagged file, and the reviewer resolves it in the manifest (set the decision's"resolution"to"apply", re-sign). Phase 2 then drivesbindery run phase1 --apply-lossyitself before the import batch, flips the lossy records to applied, and consumes the decisions, so consent no longer requires the phase-1 re-run that minted a second manifest and orphaned the first. Consent is all-or-nothing (a partial resolution refuses before anything runs); a failed strip fails the verb with the library unwritten; an unresolved lossy_consent still blocks like any open decision. - Docs truth: CLAUDE.md's contract-note sections are back in chronological order (newest first; 3.38/3.39 came in from the basement) with a 3.48.0 section on top, and the spec's script table gains
check_pdf.py,comments_census.py, anddb_util.py. Note: the 3.46.0 entry claimed this docs item; the work actually lands here. - Flag-coverage backfill:
--show-tags,--show-id,--primary-only,--plugin-data,--show-author-details,--set-comments, and--clear-commentshad zero tests; tests/test_flag_coverage.py pins all seven against the real render and write paths (including the comments table's id column and the metadata_dirtied queue). - The tag tree renders real libraries again: 3.47.0's rolled-up counts re-derived the arithmetic inline and crashed on depth-3 subtrees (a dict child recursed as an addend; the real library's Fic → Classic → African tripped it in run_tests.sh, the suite's fixtures stopped at depth two) and dropped a parent's own direct books wherever it also had children. The renderer now consumes cquarry's
tag_rollupdirectly (the engine derives, the frontend renders), so every node shows its true subtree total, and a regression test pins the real shape. - Suite: 575 → 601 tests.
The engine predicates adopted: identifierless listing, tag-tree rollups, pace year buckets, engine-driven duplicates
--identifierless: books carrying no identifiers at all, listed as[id] titlethrough cquarry'sfind_identifierless(L4 rank 11's curation queue for Calibre-Companion-style lookups); a fully identified library reports clean.- Tag-tree rolled-up counts: every node of
--tags --tree's taxonomy now shows its rolled-up book count -- a node's total is the sum of the leaf tags beneath it (cquarry'stag_rolluparithmetic, rendered in the tree). --analytics pace --pace-granularity year(L4 rank 10): the addition timeline buckets by year instead of month.- Duplicate detection is the engine's predicate now: the audit's hand-rolled (title, primary author) grouping is replaced by cquarry 1.8's
find_duplicate_books-- same key shape, but the ids inside a duplicate row are sorted numerically where book-iteration order used to decide (the one visible CSV difference). - The stale series-clear pin updated to the cquarry 1.23.1 fixed semantics: a cleared series resets
series_indexto 1.0, never NULL on real schemas.
The L4 headliners land -- Markdown search results, a machine-readable health digest, a TUI that finally covers the whole read surface, and honest format refusals
--search QUERY --format md(L4 rank 1): the Markdown catalog emitter over any query's match set -- headings per author, bulleted bold titles, the library UUID provenance header, and the query recorded in a scope note. Composes with--restrictfree (the view is already the database), and--outputnames the file as anywhere else; the default output issearch_results.md. Delegates towrite_catalog, so the emitter cannot drift from the catalogs.--health --format json(L4 rank 6, first half): the digest's counts as a machine-readable payload -- issue counts, problem tallies, metadata-quality rows, tree/FTS findings, pending OPF sync, and the annotations-dirtied count riding beside the OPF line (cquarry'sget_annotations_dirtied_books, unconsumed since 1.23).--fail-on-findingsis the second half: opt-in, flips the digest's exit from the standing 0 to 1 when anything was found, turning the dashboard into a gate. Quiet mode still gates on the exit, so scripts get both.- Per-mode
--formatcorners refuse instead of ignoring (roadmap 1351):--book --format mdused to render plain text silently;--fts --format csv/ai/mddid the same;--health --format <not-json>ignored the flag entirely. All three now exit 2 naming the one format the mode supports. - The TUI menu covers Format Stats and the Trash Listing (roadmap 1312's residue plus the 3.45.0 gap): the README's "menu covers every read mode" claim is true again, and the 3.45.0 trash surface has a menu door.
- The stale series-clear pin updated:
test_set_series_with_index_then_clearstill expectedseries_index = NULLafter a clear -- the exact IntegrityError the cquarry 1.23.1 fix closed on real schemas (REAL NOT NULL DEFAULT 1.0). The test now pins the fixed semantics (clear resets to 1.0, the link row gone). - Docs truth: the spec's script table gains
check_pdf.py,comments_census.py, anddb_util.py; CLAUDE.md's contract-note headers return to chronological order.
The curation verbs land over the 1.19.0 riders, the trash gets a surface, and the live prose layer cleans up
--rename-entity KIND OLD NEW: the fix-the-misspelled-name verb over cquarry 1.19'srename_entity, which had ridden the floor for two releases with zero consumers. Renames a tag, author, series, or publisher everywhere; a rename into an existing name MERGES the rows (links move to the survivor, duplicate links drop); a no-match old name is a clean exit 1 and an unknown kind a usage exit 2. Author renames recomputeauthor_sortand re-lay paths; series merges renumber incoming books.--set-author-sort/--set-title-sort BOOK SORT: the passthrough sort setters, storing hand-tuned corrections verbatim; a later--set-authors/--set-titledeliberately recomputes over them.- The trash surface:
--trashlists the library's.caltrashentries (category, book id, age, files) through a pure read-only filesystem inventory, andrun trash --empty/--expire DAYSowns the lifecycle through cquarry 1.20's verbs, dry-run listing by default with--applyexecuting. The apply half opensmetadata.dbwritable to reach the upstream verbs, so the closed-Calibre guard applies; no backup is owed because the database itself never changes.--format jsoncarries the plan/results shape. - restrict's refusal list,
SINGLE_BOOK_DESTS, and the setwrite combination guard keep step with the three new write dests, with a test pinning that they must; the README's full-help dump is regenerated byte-identical from the live parser. - The live em-dash layer is gone: all eleven rendered CLI separators (including the catalog header frozen into README's sample output), the README/spec titles, and every live README/spec line are recast; the patchnotes 3.40.0 entry's ASCII-dash sites are recast; the "not crying wolf" echo now appears once; the
CalibreQuarry (cquarry-cli)rename injections are down to one first-use mention (fixing the possessive grammar break). Comments keep their dashes by convention. - Comment truths:
_fetch_metadata's docstring names thefailedreturn; the download-segment comment no longer namescalibredb; integrate's docstring says which verbs actually drive external programs; setwrite's JSON report shape names theidskey; treeaudit's tolerance wording matches itsextra_cover_fileclass;audit_isbns' exit contract names the advisory rider;spot_check's 99 means "99 or more failures";compress_pdf's workflow names both synced size records; the three "Stdlib only" headers that import vir_tui say so; cli.py and tui.py carry module docstrings. - Polish: the five dead
[DRY RUN]conditionals are gone; audit's dead nested quiet-guard is unwound;run_covercounts already-so rows likerun_convert;run_flushhonors--format json; the FTS extraction-errors preview gains its tail; therun.py/setwrite.py/fetch_library_codes.pypgrep guards fail closed on OSError like integrate and reconcile already did. - Housekeeping: scripts/ exec bits normalized (every shebang'd script executable; the shared
db_util.pyhelper and the JSON template not), and pyproject takes the PEP 639 shape (SPDXlicense = "MIT"+license-files, the License classifier dropped,Python :: 3 :: Onlyadded; wheel metadata verified). GitHub topics drop the duplicatepython3and near-zeropython-314. - New-work boxes recorded on the roadmap: the dest-list constants module, the
_backup_dbtriplication, the per-mode--formatcorners, bindery's unmirroredmanual_watermark_repairdecisions, the lossy double-manifest wrinkle, and the UI-only pypi tag policy. - Suite: 542 → 559 tests.
The truth-and-hardening batch: the raw-SQL reads retire, every lying header tells the truth, and the publish path locks down
- The frontend tier's last unrecorded raw-SQL reads are gone.
_precedent_tags(the phase-3 prompt's four-table JOIN) is promoted to cquarry 1.22'sCalibreDB.precedent_tags-- released upstream under the cross-repo grant and consumed here with the floor bump (cquarry >= 1.22.0) -- and_remove_book_dry_runreads through cquarry's CalibreDB instead of raw sqlite3.modes/fts.py's sidecar read remains the one recorded exception. Both reads are pinned end to end by tests. - The publish path hardens (the workspace batch, byte-identical across cquarry and CalibreQuarry; vir-tui verified as already shipped): SHA-pinned actions (the pypa publish action was a moving branch holding
id-token: write), a top-levelcontents: readpermissions block, a no-cancel concurrency group, the CI's pinned ruff gates before the suite, a strict twine check plus a wheel smoke-install, and acreate-releasejob that mints the GitHub Release from the tag's verbatim message. Tag protection ships as arelease-tags-protectedruleset (deletion blocked forrefs/tags/v*), and an actions-only dependabot keeps the pins current. One piece is recorded as reverted: REST-created environment deployment policies are branch-type only and reject tag deployments outright (cquarry's v1.22.0 publish proved it live); a UI-applied tag policy is the reopen item. - The comment/header truth batch: the
--fieldshelp no longer advertisescomments(backfill refuses it by design);validate_metadata's header lists all fourteen checks (three were missing); librarything's module header documents the realcquarry --exportltsurface instead of a standalone CLI that never existed; integrate's convert skip comment names both shapes it swallows;SINGLE_BOOK_DESTSactually carries every single-book dest (add_tag/remove_tag included, and setwrite drops its duplicate inline pair); the-> Noneannotations hiding load-bearing exit codes in export and catalog tell the truth; andreconcile_file_metadata's pgrep guard is fail-closed on OSError like its own comment always promised. - The docs tell the truth: the spec's accreted cquarry-floor sentence (stale three times) collapses into a floor-plus-bumps list with the floor pinned against pyproject by a test; the
--formatrow namesmd; the no-network claim scopes itself to the read surface and namesrun backfillas the one exception; and the README test-suite section says 542 tests across 26 files, naming every suite (it said 373 and named ten). - The version-sync guard is real now: tests enforce the spec's Version header, the spec's cquarry floor, and the roadmap's
Updated as ofstamp, alongside the existing code/VERSION/pyproject/patchnotes set -- the three carriers whose drift the audit caught live. - Housekeeping:
.gitignoregainsvenv/,.venv_ci/,.pytest_cache/,.ruff_cache/, and.claude/(the tool self-ignores were masking the gap); the dead[dependency-groups]pytest residue is deleted;docs/taxonomy.example.yamlships a placeholderlibrary_path; README's LICENSE and screenshot links are absolute URLs that render on PyPI; and REPORT-12-Sept.md moves todocs/with its roadmap references updated. - Suite: 536 → 542 tests. Floor: cquarry >= 1.22.0.
The record-integrity batch: the same-day manifest collision, the Z-Lib provenance ruling, the lossy-consent hole, and the exit-code unification
- Two same-day batches can no longer destroy each other's manifests (the final audit's HIGH). Phase 1 saved every batch to a fixed
{date}-batch.json, so a second batch on the same day silently overwrote the first, destroying the durable record (imported ids, decisions) that phase-2 resume and phase-3 consume. The newcomer now takes{date}-batch-2.json, the same exists()-loop the backup paths already used three times. Regression-tested on a real collision. - Provenance seeds
Z-Lib, the ruling's spelling. The 2026-09-13 enum ruling renamed the #source value toZ-Lib(Brandon's entry spelling), but the seeder still emitted the 3.40-eraZ-Library: every z-lib file in a 3.42-run manifest failed its phase-2 stamp until hand-corrected (19 files in the 2026-09-14 STEM run). The seeder, the test fixture's enum, and both import skills now trackZ-Lib; manifests produced by 3.40-3.42 still carryZ-Libraryand need the hand-correction before signing. - Bindery's gate-accepted EPUB repairs land in the manifest's lossy records (observed the same run). Bindery's phase-1 report carried an
apply_lossydecision naming gate-accepted repairs, and the manifest still wrote{flagged: false, repairs: []}: the seal bound nothing, so signing consented to repairs it never saw._mirror_lossynow walks bindery's repair records: status accept/partial becomesflagged: trueplus the named repairs and anappliedflag, in both dry and--apply-lossyruns. - The "Calibre is running" refusal is lock-class exit 1 everywhere. Set mode exited 1 while the five run/integrate doors (phase 2, phase 3, dispatch_integrate) exited 2, and scripts branch on these codes; the recorded discipline ("usage problems exit 2, lock/write errors exit 1") now holds across all of them. dispatch_integrate also runs its usage guards before resolving the library, so
run convertwith a missing--idsis a usage error (exit 2) however resolvable the library is. The two 3.41-era guard pins were updated with the reasoning. --exportrefuses honestly.--export --format md(md became a legal catalog format in 3.42.0) and a bad--show-customprinted their refusal and exited 0, reporting success while writing nothing; run_export now returns 2 and 1 respectively, exactly like the search-export path. The implicit--wingcatalog fallback also passes--formatthrough, so--wing W --format mdrenders Markdown instead of silently plain text.- A typo'd TUI scope no longer runs unrestricted silently. The
_restrictedparse-failure note was printed and then immediately erased by the next terminal reset; it now goes through the blocking notice, so the reader must acknowledge it. --restrict ... --tagsobeys the universe.--tagswas the last read mode reading a global aggregation through the RestrictedView; tag counts are now recounted from the scoped book rows. Correction recorded against the final audit: its ratings-recount half was a misdiagnosis (the recount's key was and is correct; a test pins it so the suggested "fix" cannot land later).- Suite: 525 → 536 tests. Floor note: the vir-tui floor moved to >=2.5.0 on 2026-09-14 (upstream 2.4.0/2.5.0 are additive; every fresh resolution takes it automatically).
The blitz candidates close: --health, the era split, Markdown catalogs, and the TUI's Phase 19 surfaces
--health, the one-shot digest. The audit's finding counts in one short screen: book issues with the top problems, duplicate groups, series gaps, conversion overrides, the metadata-quality trio, the filesystem tree, FTS coverage, and the pending OPF queue. Both renderers now consume one shared derivation (collect_issues), so the CSV and the digest cannot drift;--restrictscopes the book-level classes exactly as it scopes--audit, and--healthalways exits 0 (a dashboard, not the audit's CSV).- The reading-analytics era split. Days-from-added-to-finished used one median over all spans, and backfilled pre-library reads (finished before their added date) dominated it: the real library's median was -489, which said nothing about how cataloged books actually read. When negative spans exist the two populations report separately (library era / pre-library, each with median, mean, min, max); a library with no pre-library reads keeps the previous single line unchanged.
- Markdown catalogs.
--format mdrenders the catalog's Markdown shape: one#header with the same provenance content,##per author, bulleted books with bold titles, an hr and a bold total.--catalogpasses its format through, and the wing and saved-search sweeps name their files.md. The plain text form is untouched when no format is given. - The TUI joins Phase 19. Five new menu entries, each with a shared scope prompt that adopts the
--restrictmodifier per invocation (blank = whole library; an expression resolves once through the CLI's RestrictedView; a parse failure notifies and stays unrestricted): Content Search (FTS), Saved Search Catalogs, and Library Health in the first section; Reading Analytics and FTS Index Status under Analytics. The Catalog and Catalog Wings entries gain a Markdown prompt. The menu structure is testable now, with the Settings section pinned where the s/q aliases need it. - Cosmetic: detail-less plan lines (backfill, polish, cover, flush dry runs) no longer end in a trailing space (the recorded Matrix 3 note).
- Suite: 512 → 525 tests.
The six-lens audit batch: three HIGH integration-seam defects, the run-verb hardening, and the routed metadata-quality rows
run flush --applyno longer writes outside the queue. The verb joined each chunk's distinct ids into one hyphen range ("5-900"), and calibredb reads a range as EVERY book between the endpoints, so a non-contiguous dirtied set embedded metadata into books the queue never named. Ids now pass space-separated, and the regression test pins a 1,3 queue against touching the between book. A bad--ids/--searchon flush is a usage error (exit 2) instead of a traceback.run phase2's metadata fetch can actually succeed now._fetch_metadatapassed the output path after-o, but-o/--opfis a store flag whose OPF arrives on stdout: the path was a silently-ignored stray positional, the success gate always failed, every import queued a bogusmetadata_downloaddecision, and the ok branch,_apply_opf, and the clobber watch were dead code (the sibling seam 3.39.2 fixed in the backfill; this call site was missed). Stdout is staged to the temp file like the backfill does; the ambiguity sniff reads "multiple" only, since the no-result log's "No matches found" classified every empty lookup as ambiguous. The seam tests mock the subprocess, not the verb, so the revived path runs for real.run convert --applyno longer overwrites an existing target. Converting to a format the book already has used to plan the conversion anyway: ebook-convert overwrote the file, then the registration raised uncaught. Like the merge verb's move list, the plan skips such books ("already has TARGET"), apply counts them already-so, and a registration failure is a failed report row instead of a traceback.--restrictis refused with the run verbs (spec 3.4's contract, previously bypassed by dispatch order):--restrict EXPR run ...ran unrestricted; it now exits 2.- The integration-verb pgrep guard is fail-closed: a pgrep timeout (or an unrunnable pgrep) answers assumed-RUNNING and refuses
--apply, matching run.py's recorded semantics instead of proceeding against a live Calibre. run backfillfailures are real: a book whose OPF fails to apply is a counted failure that fails the verb (exit 1) instead of "Applied 0, failed/skipped 0" at exit 0; a malformed OPF or refused write is a report row, not a traceback; a hung lookup times out into a failed row;--fields isbnprefers theopf:scheme=ISBNidentifier and falls back to an ISBN shape through cquarry'sto_isbn13(the firstdc:identifiercould be a Goodreads id); the fetch no longer appends a stray--opf <path>positional.- The catalog sweeps report per-file failures:
--all-wingsdiscarded write failures entirely ("All wings written" at exit 0), and--all-saved-searchescounted a catalog before the write ran. Both now count only files that exist, drop the stale file of a failed entry, warn per failure regardless of--quiet, report "N of M written", and exit nonzero. - Set mode's backup goes through the sqlite backup API (the last copy2 door among the write paths; a file copy of a database with a hot journal can snapshot a state its WAL would never replay into).
--auditrenders the routed metadata-quality rows (the bindery ruling of 2026-09-12): cquarry 1.21'sfind_invalid_uuids,find_sentinel_pubdates, andfind_bad_language_codesas advisoryissue_typerows with the offending values in brackets and summary blocks, one class per commit with fixtures and false-positive notes. Real-library probe: all three classes are clean at the database level; bindery's 51 OPF-085 warnings were file-side (stale sidecar OPFs), not DB drift.- Test hygiene: the
__main__guards of three test files moved below the suites appended after them, so direct-file runs exercise the 3.39/3.40 regression classes again (discovery was never affected); and the phase-3/backfill guard tests now pin the closed-Calibre pgrep instead of assuming it, so a desktop Calibre that happens to be open no longer flips five outcomes. - Suite: 483 → 512 tests. Floor: cquarry >= 1.21.0 (the write-path fixes plus the predicates).
- The enum "bug" was a misdiagnosis (retracted). The 3.37.0 notes recorded that the contains form over a normalized enum column (
#reading_status:Read) matching the whole library was a cquarry engine issue. Instrumented comparison says otherwise: contains is honest case-insensitive substring semantics, identical to upstream's CONTAINS_MATCH (query in t), and in this library every status value ("To Read", "Reading", "Read") contains the substring "read", so a full-library match is the correct answer. cquarry needed no fix. Prefer the exact form (=Read) for precise enum selection; that guidance stands. - The tree audit whitelists the library root's workspace furniture (the 2026-09-12 decision): dot-entries (
.claude/,.ruff_cache/,.nomedia,.trackerignore) and the named doc/tool set (CLAUDE/AGENTS/MEMORY/README/roadmap/spec/patchnotes/refresh/TAXONOMY.md,taxonomy.json/taxonomy.yaml,validate_library.py) are never findings: a library that doubles as a working checkout carries them by design. On the real library this drops 13 findings to zero. - Z-Library provenance goes forward (the enum decision pending since 3.35): z-lib filename markers (
z-library.sk,1lib.sk,z-lib.sk) now seed #source "Z-Library" instead of "Other". Your one-time step: add theZ-Libraryvalue to the #source enum in Calibre BEFORE importing a z-lib batch; enum validation refuses unknown values, deliberately. Existing "Other" rows stay untouched. run backfilllive-verified under an authorized one-off network drill on a scratch library: the fetch, the OPF round-trip, and the title write all work end to end (the drill also caught that the verb's plugin restriction allowed no metadata source at all; fixed alongside).- Suite: 482 → 483 tests.
run backfill --applywas dead on arrival: the command builder orphaned its own--opfargument (a leftover slice dropped the temp-file path), so the metadata source could never deliver and every book failed. Fixed and proven end to end: a mocked source writes a real OPF, the fetched title lands through cquarry's write module._apply_backfilllooked in the wrong namespace: real fetch-ebook-metadata OPFs usedc:-prefixed elements; the code searched the opf wrapper namespace and would have found nothing. Fixed; the end-to-end test pins the real shape.--fields commentswas accepted, planned, and then silently ignored while reporting "applied": comments is now refused with the available list. Overwriting a curated description from a publisher OPF is a curation decision, not a backfill.run backfill --applyrequired no--backup-dirdespite mutating metadata.db: it is in the backup set now.--searchsilently beat--idswhen both were given (the ids, including unknown ones, were never validated): the combination is refused (exit 2), which is what "exactly one target source" always meant.- A bad
--searchexpression crashed with a raw traceback in every verb: it is now a clean usage error (exit 2). --format jsonand--quietwork afterrun(subcommand position), matching--db's dual-position treatment; plans and reports carry{plan}/{results}JSON either way.- Backups moved to after plan validation: an invocation aborted by a usage error no longer writes a backup nobody needs (timestamped backups made this harmless, but it was undocumented behavior).
- Found by the functional-pass matrix (inspection + dry-run probes) and closed with 6 new regression tests, including backfill's first. Suite: 476 → 482 tests.
run flushpassed metadata.db to--library, which wants the library DIRECTORY: calibredb died withapsw.CantOpenErroron every call. The verb now passes the directory and its id spans match calibredb's documented grammar (space-separated ids, hyphen ranges).run exportpassed--dont-save-opf, which does not exist (the real flag is--dont-write-opf): every export refused with a usage error. Fixed, and both verbs now regression-test the exact command line they build. Found by the live mutating drills on a scratch library, which is what they are for.- Skills sync: swept both import skills for the C verbs; they are library-maintenance surfaces with no import-flow teaching and nothing went stale. Suite: 474 → 476 tests.
Phase 19 C: the integration verbs (subprocess-driven; no new dependencies; every verb dry-run first)
run convert:ebook-convertover a resolved set (--search/--ids); the source is--from-formator the largest other format, the output lands in the book's directory under Calibre naming and registers throughWritableCalibreDB.add_format. Failures are per-book and listed.run polish:ebook-polishwith--polish-ops(smarten, unused-css, compress-images, subset-fonts, jacket, kepubify) on each book's EPUB; the post-verify re-syncs the polished file's size throughset_format, so Calibre never sees a stale size.run cover: places an image (--cover FILE) or removes the cover (--remove-cover) through cquarry 1.19'sset_cover/remove_cover; closes the loop the coverless/low-res/aspect audits open.run export:calibredb export --templateper resolved ids into--dest; calibredb missing is a clean setup refusal, not a traceback.run merge --keeper ID --duplicate ID: the duplicate's unique formats are copied into the keeper's directory and registered, then the duplicate's removal goes to cquarry 1.20's trash (.caltrash/b/<id>/, recoverable by hand until expired or emptied).run flush: the headless OPF-queue flush:calibredb embed_metadataover themetadata_dirtiedids in--chunkbatches; an empty queue is a clean exit 0.run backfill: drivesfetch-ebook-metadataper book (network only at--apply) and applies the requested--fieldsthrough cquarry writes, one batch per book.- Shared discipline across all seven: dry-run by default with a printed plan;
--applydemands a closed Calibre (anchored pgrep) and, for metadata-mutating verbs, a timestamped--backup-diroutside the library; targets resolve read-only first and unknown hand-supplied ids abort (exit 2) before anything opens writable; one report shape with per-book results and--format json. - check_library subprocess parity: skipped by the recorded A.3 route decision (the tree audit is CQ-native since 3.37.0).
- Erratum for 3.38.0: that release's note said the skills sweep found "nothing stale"; the sweep also ADDED teaching (phase-1's check_pdf depth note and phase-3's year_mismatch interpretation), which the released entry undersold.
runnow also accepts--dbafter the subcommand (the facility-run document has been suggesting that shape all along). Suite: 461 → 474 tests.
Phase 19 B: the audit-depth batch (one class per commit, each with its fixture and false-positive note)
- Content-duplicate fingerprinting (
scripts/audit_duplicates_content.py): 64-bit simhash over 3-word shingles of each book's spine text finds re-downloads filed under different metadata; a bottom-32 shingle-sketch candidate pass catches omnibus containment (an omnibus's simhash is not close to the standalone's, so Hamming distance alone never sees it), and candidates are classified exactly:near_duplicatevsomnibus_overlap. Front-matter-only files are excluded (under 200 shingles a simhash is noise) and formats without an extractor are skipped, not guessed; legitimate public-domain reissues are findings by charter. ~1s per book: scope with--search/--ids. - Truncation cross-checks (
scripts/audit_truncation.py): the Count Pages plugin's page counts meet poppler'spdfinfo; more than 20% disagreement ispage_count_mismatch, andstale_plugin_data(catalogued size drifted, the plugin's own needs_scan flag, post-scan mtime) is its own class. Only PDF rows are checked: the plugin's EPUB pages are word-count estimates by design. - Author-sort sanity in
validate_metadata.py:AUTHOR_SORT_NOT_INVERTED(sort identical to a multi-word display name) andAUTHOR_SORT_ORPHAN(matching no legitimate shape of the book's authors). Hosted locally by decision: cquarry never boxed the predicate, and the promotion remains a future option. The first cut compared against display names only and flagged 7718 of 7842 real books; the real-library probe caught it within minutes and the fix accepts every legitimate Calibre shape (display name, the authors.sort column, mechanical inversion, the&-joined multi-author sort). Real probe after the fix: 0 orphans. - DB-level ISBN checksum in
validate_metadata.py:INVALID_ISBNwarns on any isbn identifier failing its check digit (cquarry's existing helper; no file open). Blank-ish values are "no ISBN", not invalid ones. - Cover aspect-ratio bands (
scripts/audit_cover_aspect.py): covers outside w/h 0.55-0.80 report as advisorycover_aspect_narrow/cover_aspect_widethrough cquarry's header-only image readers (no image library). Legitimate landscape art exists; the audit surfaces the distribution, the operator judges. - FTS coverage audit in
--audit: the A.1 staleness classes render asfts_coverageCSV rows. With no sidecar the CSV stays silent (the rows would be the whole library) and the prose summary carries the absent-sidecar note instead. - PDF battery depth in
check_pdf.py: text sampled at pages 1, middle, and last (the page-1-only sample passed OCR-once scans clean; a partial pass is nowtext_layer_partial), and image DPI parsed from the existingpdfimages -listoutput as an area-weighted mean (a small sharp logo cannot hide a full-page 72-dpi scan; below 150 dpi is alow_dpiadvisory). Advisory classes only: the structural total and the run-verb seam contract are unchanged. - Copyright year vs pubdate in
audit_isbns.py: the front-matter pass captures (c)-years (each copyright marker claims the years on its line), and an earliest-year-vs-pubdate gap over two years is ayear_mismatchadvisory with its own report section, riding the exit-1 findings contract. A reprint legitimately prints the original year; the report asks whether the pubdate describes this edition. - Skills sync: swept both import skills for the new audit surface; nothing teaches the old shapes, no staleness found. Suite: 433 → 461 tests.
--restrict SEARCHscopes every read mode. Stats, audits, analytics, exports, catalogs, full-text search, and the rest compute over the books matching a search expression (a wing composes asvl:Name). The implementation is a scoping view over cquarry's connection: every predicate and stat is still derived by cquarry, only the inputs are scoped, and the two SQL-level aggregations (entity counts, format stats) are recounted so they stay honest. Book-level audit findings follow the restriction; library-shape ones (orphan dirs, root strays) always report globally. Write verbs and--book/--idrefuse the combination (exit 2); a bad expression exits 1 like--search. Upstream precedent:--restrict-toon calibredb fts_search and restricted-id category counts.--fts QUERYsearches what the books say. Content search over Calibre'sfull-text-search.dbsidecar (the plainbooks_texttable; no FTS5 machinery), case- and accent-folded, every match naming its formats,--format jsonexport through the output guard. Every run ends with an index-staleness summary in separate classes (never indexed; indexed empty; extraction errors; stale entries queued indirtied_formats), and--fts-statusreports just that. A missing sidecar (the common case until Calibre builds its index) degrades with a note, never an error.--auditnow walks the filesystem. Tree rows against the database: missing book dirs and format files, extra format files and unknown files inside book dirs, covers on disk the DB does not claim, orphan book dirs and author dirs, malformed book-dir names, stray root files, and unreadable directories. CQ-native by decision (no calibredb dependency; the walk composes with--restrictand the shared CSV shape); upstream check_library's class list is the completeness checklist.metadata.opf, any*.opf, cover files, anddata/are never extras; name comparison is case-insensitive.--analytics reading(read-only). The#reading_statusfunnel in the column's configured enum order with(no status)last; recent finishes from#date_read, newest first; days from added to finished with median/mean/min/max and a stale-timestamp note. The NON-NEGOTIABLES write ban is untouched: an mtime-pinned test proves nothing writes. Missing columns degrade to a clear message.--all-saved-searches. One catalog per saved search into--outdir(the--all-wingsanalog), each headed with the search's own expression; zero-hit searches write nothing and say so; an unresolvable search is skipped with a warning, never a dead sweep.@Nameuser-category resolution: skipped by its own gate. The library's preferences carry zero user categories (read-only peek), so the box records not-in-use instead of building a resolver. If categories ever appear, the right home is the cquarry search engine.- Adopts cquarry 1.20.0 (floor bump ahead of Phase 19 C's integration batch; the trash lifecycle and create/delete_custom_column land in the shared layer this CLI's next release consumes).
- Adopts cquarry 1.18.0 (floor bump only; no behavior change required here yet): the FTS sidecar reads, the search-parity honesty pass, and the write completions land in the shared layer this CLI consumes; Phase 19 consumes them for real.
- Adopts cquarry 1.19.0 (floor bump only): the four approved write verbs land in the shared layer (rename_entity/remove_entity_everywhere, the misspelled-author fix; set_cover/remove_cover, which closes the low-res-cover audit loop; the verbatim sort setters; and save/restore_original_format). Phase 19's curation and cover verbs build directly on these.
- Found while composing
--restrictwith the status column and recorded for the cquarry lane: the contains form over a normalized enum column (#reading_status:Read, quoted or not) currently matches the whole library; the exact form (#reading_status:=Read) is correct. Fixing it belongs upstream in the cquarry search engine. - Skills sync: phase-3-import's audit step names the new tree classes,
runs the audit with
--outputoutside the library root (a stray audit.csv there is now itself a tree finding), and uses--restrictto scope post-import verification to the batch's tag; phase-1-import swept, nothing stale. The README's full help dump was regenerated from the live parser. Suite: 385 → 433 tests.
--auditreports manual conversion overrides. The last open box (promoted from the 3.34 sweep's scripts verdict): the check behindscripts/audit_conversion_overrides.pynow renders inside the audit mode. Every book carrying a per-bookconversion_optionsrecipe gets aconversion_overriderow in the CSV (book id, title, author, the format, and the recipe blob's size) plus a summary block naming the affected books, with the pointer to Calibre's conversion dialog where those blobs are inspected or cleared. The rows consume cquarry'sget_conversion_profiles(the predicate's library home; the frontend renders, never re-derives, and the pickles are never unpickled).- The standalone script stands.
audit_conversion_overrides.pykeeps its pipeable surface (--quietprints only the ids; exit 1 when any are found, so the report feeds a repair workflow). The mode keeps--audit's exit-0 reporting contract; no exit-code change for existing users of either surface. - Housekeeping: the
--audithelp line names the new check and the README's full help dump was regenerated from the live parser (the dump had also drifted from the parser on--format-stats' position, which this regeneration fixes). Suite: 382 → 385 tests.
- The filename-stamp convention is decided (roadmap :902). The stamp writer is the authority: run.py's parser reads "Author - Title", the stamping path emits exactly those values, and the observed corpus (libgen.li's "[Series] Author - Title (year, publisher) - site" names, verified in the 2026-09-10 Redwall run) confirms the direction, so that convention stands. Calibre's own filename fallback guesses the opposite (probed on a metadata-less file: "Brian Jacques - Mossflower.pdf" imports as Title "Brian Jacques"). The decision is now recorded where the readers live: run.py states the convention, stamp_pdf's preview documents that it deliberately mirrors Calibre's opposite guess (that is what a preview of an unstamped import is for), and screen_duplicate names the screening gap metadata-less "Author - Title" files carry. A corpus regression test pins the direction.
run phase1seeds the manifest'sprovenancefrom the filename (the second Redwall box). The field existed but was never populated, so phase 2's cc6 stamp fell back to a blanket "Anna's Archive" regardless of true source and phase 3 re-derived provenance from filenames every batch (observed 2026-09-08 and 2026-09-10). Phase 1 now derives it from the same filename evidence the review already worked with, mapped onto the #source enum's vocabulary: the "-- Anna's Archive" trailer seeds Anna's Archive, libgen.li seeds Library Genesis, z-library.sk/1lib.sk naming seeds Other (the recorded practice of both runs, pending Brandon's Z-Library enum decision), and a name with no marker seeds nothing. The review step corrects the value exactly like the stamps, and phase 2 keeps stamping cc6 from the reviewed value. The HMAC seal now binds provenance alongside the stamps and lossy flags: a post-sign source swap fails every load until re-signing.- check_pdf.py classifies qpdf exit 3 by its documented contract.
qpdf writes its warnings (and the "operation succeeded with warnings"
summary) to stderr, so the old stdout marker gate never matched and
every warning-only file was recorded as
errorsin the battery report, the CLI summary, and the phase-1 manifest; the Redwall run filed two benign PDFs (unknown-token tolerance; linearization /E + hint-table drift) as structural damage. The class now reads the exit code alone (0 clean, 3 warnings, 2 errors) and a warning finding carries the first warning line as triage evidence. Re-triage confirmed both warning kinds benign; the Multics PDF re-checks clean today because the phase-1 stamp rewrite healed its linearization drift. - The parser-disagreement box (:902) and the Redwall run's two (the
provenance field, the qpdf exit-3 class) are closed; the
--auditpromotion ofaudit_conversion_overridesis the one box still open, landing next. Suite: 373 → 382 tests.
- The TUI respects the saved library.
_resolve_db_for_tuiconsulted a hard-coded default list that started with a CWD-relativemetadata.dband never read the saved config, so launching the TUI from any directory with a straymetadata.dboverwrote the shared config and the next CLI run read the wrong library. The saved path wins whenever it exists; discovery binds only when nothing is saved. - Failures exit like failures. A bad
--searchexpression exits 1 (matching--exportlt --search, which always did); an unknown--wingexits 2 instead of leaving a stale catalog file standing in for a fresh one; and exportlt's self-check verdict ("do not upload") reaches the exit code instead of dying behind an unconditional 0 (its raw-SQL half was already retired by the cquarry 1.17export_rowsadoption). - Read-mode papercuts.
--export-annotations/--exportlt/--format-statsjoin the mutually exclusive read-modes group (--format-statshad been declared in the write-verbs group;--untaggedstays a--bookmodifier on purpose); negative--recentis refused with exit 2; a corrupt epoch renders raw instead of killing--reading-progress/--book; the TUI Entity Browser prints a readable refusal for an unknown kind instead of paging a traceback; plugin values ride along in json/csv/ai serialization instead of vanishing; and a directory export target that exists as a file is refused in prose. - The suite runs the tests, and the scripts run at all.
run_tests.shexecutes the hermetic unittest suite before the real-library smoke (following the documented command used to deliver smoke-only coverage). Newtests/test_instruments.pydrives the actual companion scripts through the actual seam adapters, and immediately caught a genuine bug:check_pdf.pyreferencedargs.quietwithout declaring the flag, so the PDF/DJVU battery crashed on every invocation, invisible for exactly as long as its callers ignored exit codes and empty reports. The flag is declared. Suite hygiene: the mid-fileunittest.main()guards that silently truncated direct runs moved to true EOF (test_scripts.pyhad 1028 lines after its guard), andtest_manifest.py's real-library path literal is synthetic. - Scripts housekeeping.
fix_cq_lint.sh(the sweep's only delete) is gone;taxonomy.example.yamlmoved todocs/with the README pointer updated;comments_census.py's--jsonhelp no longer claims a runner that never consumed it. Two deferrals carry dated notes in the roadmap: the six drifting_SCHEMAfixtures still want a shared builder, anddb_util's consolidation waits because the privateconnect_rocopies have genuinely drifted (reconcile needsRowrows and its own temp layout); theaudit_conversion_overrides->--auditpromotion is now its own open box. - Docs truth. The README troubleshooting line claiming saved searches
"match nothing" is corrected (they evaluate, cquarry 1.1+); the search
engine is named where it lives (
cquarry.search, not a nonexistent local file) and the "zero dependencies" claim is replaced with the real dependency set; the six botched "minimal-dependency (uses tqdm)" artifacts are cleaned. The spec absorbs the seven shipped modes its table never listed (--book,--entities,--reading-progress,--columns,--info,--exportlt,--format-stats),--set-pubdate/--clear-pubdatejoin the verb list, and §5's closing sentence names all four writer scripts. The README grows a run-verbs prose section (the file-side consents:--stamp,--apply-lossy,--quarantine;--bindery-report;--audience) and companion script sections for the five tools that had none, while the full help dump is regenerated from the live parser (--yesgone,--quarantinein). Housekeeping: Phase 17's seven malformed double checkboxes normalized, the patchnotes H1 moved to the top of the file, the pre-3.14## vX.Y.Zentry headings normalized to the dominant# X.Y.Zstyle, and the README test count made current. - Phase 18's 26 boxes are now 25 closed and 1 open (the
audit_conversion_overridespromotion, opened from the sweep's scripts verdict). Suite: 362 → 373 tests.
- Phase 1 moves files only on true DRM and only with consent.
Quarantine used to run for any DRM verdict that was not clean, sweeping
in audit_drm's BENIGN (font obfuscation) and N/A (DJVU): every DJVU in a
batch was moved and a manual_repair decision recorded for a file with no
DRM at all. The classification is now audit_drm's own problem set (DRM,
plus ERROR, a scan that could not verify), and the move is gated behind
a new
--quarantineflag: without it the verdict and decision are recorded and the file stays. Quarantine also no longer moves onto basename collisions (numbered siblings instead), stamp backups live in a dated temp directory outside the tree (which makes--stampactually work, since stamp_pdf refused the in-tree backup dir it was handed), stamp failures print a WARNING instead of passing silently, and a rerun can no longer sweep_stamp_backupsor_quarantineas books. - Phase 2 keeps its accounts. The resume record saves the moment the
import batch commits, before the unguarded download segment, so a
crash there no longer costs the imported ids.
--audiencewith no flag stamps the documented default instead of the literal string 'None'. Backups are timestamped and taken through sqlite's backup API, so a second run keeps its own restore point and a hot journal cannot leave an inconsistent snapshot. The rollback message tells the truth (the batch rolled back; nothing was written). The dead--yesflag is gone, and phase 1's--apply-lossyis now actually wired to the bindery slice. add_book's orphaned directories on batch failure are closed upstream (cquarry 1.15's batch-scoped compensation), as is the byte-identity duplicate floor. - Phase 3 has the same rails as every other write path. A
closed-Calibre guard runs before the answer gates; the answer file may
not name
#reading_status/status/date_read(checked against the shared NON-NEGOTIABLES tuple before anything opens writable); bindery/reconcile trouble (their exit 2) always prints, lands in the batch record, and fails the verb instead of hiding behind a clean validator; and the download segment re-checks the Calibre guard, deferring to phase 3 rather than racing a library that opened after the commit. - Run-verb papercuts. The PDF battery now covers the same recursive
inventory everything else uses (not just the top level) and honors
check_pdf's exit codes, with its report in a temp file; answer-file
loading produces readable errors and warns on ids the manifest never
imported instead of dropping them silently; a download timeout maps to
failed, not ambiguous; quarantined files appear in
files[](the quarantined verdict is no longer writer-dead); the dead_existing_book_idshelper is deleted; and phase 2'scalibredb set_metadatashell-out is replaced by_apply_opf, which applies a downloaded OPF through cquarry's write module in one batch (the no-calibredb constraint holds again). - The single-verb write path batches.
run_writewraps every action inwdb.batch(), so an interrupt mid-verb cannot commit a torn edit invisible to Calibre's OPF sync (cquarry 1.15's__exit__rollback is the upstream half of the same fix). --commit-per-bookdoes what it claims. Each book is its own outermost transaction; a book whose verbs failed rolls back alone, its entries readrolled_back/book_committed: false, and the pass continues. Partial rollbacks are reported as such; exit 0 only when every book committed.- Empty-string flag values are refused arguments (exit 2).
--batch-set-title ""used to vanish through the truthiness gates while--batch-set-column audience ""was collected and cquarry treats '' as a clear, wiping the column on every targeted book with rc 0. Both are refused now, and_has_verbscounts value-bearing flags by presence so the refusal message is what the user sees. - The banned-columns refusal is a shared chokepoint.
#reading_status/status/date_readwere refused at set mode's door and open at the other two (single-book--set-column/--clear-column, the TUI). The refusal now lives in the writeops action builders (writeops.FORBIDDEN_COLUMNS, one tuple for every door), raising the argument-level exit 2; cquarry 1.17 deferred the policy upstream, so the frontend owns it. - Set/write papercuts.
--batch-set-cover maybeexits 2 instead of tracebacking; the--batch-clear-ratingmanifest gate validates the file with the repo's own manifest module (a plain id file no longer unlocks a bulk rating clear); failure detail survives--quieton stderr; the--remove-bookdry run is a read-only connection, never WritableCalibreDB; set-write backups are timestamped;parse_book_idrejects int()'s'5_0'; the dry-run JSON carries the resolved ids; and the--backup-dirusage check precedes the pgrep probe. - Skill sync (same release): the phase-3-import skill's
--batch-clear-ratingline names the sealed-manifest gate; the phase-1-import skill names--quarantineamong the file-side consents. Suite: 330 → 362 tests.
- The phase-1 seams work on real input.
run phase1crashed on any non-empty directory: the duplicate screen's JSON report is a bare list of per-file records and the runner read it as a dict, and the same comprehension would have flagged every screened file as a duplicate (only records with library or within-batch hits count now; an unparseable report is a hard error instead of a silent pass). The screener is handed exactly the files its own extension set covers, so a djvu-only tree is a clean screen rather than the script's exit-2 setup error. The bindery seam now passes the--json FILEargument bindery always required and reads the report back from the file; bindery's exit contract (2 = trouble found, the report is still written; 1 = invocation problem, no report) replaces the raise-on-2 that fired on exactly the case phase 1 exists to surface, and a missing report names a stale entry point instead of sailing on. The seam tests pin the instruments' real payload shapes, including one run against the actual screen_duplicate.py. - The manifest signature is a real seal.
sign()used to set a bare"signed": truein the same editable JSON, and a manifest whose rejected file was listed as approved passed phase 2 (proven end to end by the sweep). Signing now seals the approved set, the per-file stamps and lossy flags, and the decisions list with HMAC-SHA256 over canonical JSON; every load of a signed manifest recomputes it and refuses a mismatch with a re-sign hint. The seal is tamper-evidence, not secret authentication (the key is a schema constant); the honest re-sign path is the newcquarry run sign --manifest FILEverb, which checks structure but not the stale seal so deliberate edits can be re-approved.approve()refuses to list a file whose verdict is notapproved_for_import, and validate() flags any list/verdict disagreement it sees, so the forged-approval attack is caught twice. Phase 2's own appends re-seal on save, keeping the retained manifest verifiable for phase 3. Manifests signed before this release must be re-signed (none exist outside tests: the verbs failed on real input until now). - A read mode can no longer overwrite metadata.db. No output writer
compared its path to the database path:
--export --output <library>/metadata.dbreplaced a fixture database with a JSON report, exit 0. Newcquarry_cli/output.pyis the shared last mile for every read-mode file output (--export,--search --output,--catalog,--audit,--export-annotations,--exportlt): the database and its sqlite sidecars are refused before anything opens, the directory exporters also refuse the library root itself (their stale-file sweep would write and delete inside the library), every file stages through a temp copy replaced into place only on clean close, and the refusal exits 2 as an argument error. - The TUI survives a malformed database. CalibreDB was constructed outside every exception boundary in the menu loop, so a corrupt or foreign sqlite file at the chosen path ended the session in a raw traceback, and Change Database validated only the filename suffix. The loop now probe-opens the database every iteration; a failure prints a prose notice and drops into the re-prompt, Change Database refuses a path that does not open and keeps the configured database, and a sqlite3.Error raised mid-session degrades the same way.
- Skill sync (same release): the phase-1-import skill's known-defect
paragraph is retired (the seams work as of this release; the
scripts-tree caveat for wheel installs stays) and its sign instruction
now names
cquarry run signwith the seal semantics. The phase-3-import skill's manifest references were swept; none went stale. - Suite: 298 → 330 tests.
- The LibraryThing exporter runs on
export_rows().build_rowsretired its hand-rolled correlated-subquery SQL for cquarry 1.17's flat row provider (cquarry 1.16's native custom-column lists underneath); same columns, same author_sort/title order, same sentinel and empty shaping, so the CSV stays byte-identical for the same library state. - Native-list display. With cquarry 1.16, multi-valued custom columns read back as native lists; the csv/ai export writers and the catalog renderer re-join them with the historical comma form so output shape is unchanged (JSON keeps the native list).
- cquarry floor moves to >=1.17.0 (the promoted APIs this release adopts).
- Test fixture fidelity: the phase-2 fixture's
#sourcecolumn now mirrors the real library (enumeration, normalized storage, the real enum values) instead of a text+direct shape no real Calibre schema creates -- cquarry 1.15's datatype dispatch refuses that shape, and the multi-file import test payloads are now distinct so add_book's byte-identity floor is exercised honestly.
- The manifest (
acquisition-manifest/1). One JSON per batch in the library-local.claude/manifests/, RETAINED as the durable machine-readable record (the prose.claude/project_*.mdfiles stay the human summary). Per-file verdicts, checks, lossy flags, filename-derived stamps, provenance, import outcomes,decisions_needed(fixed taxonomy), andapproved_for_import. The six 2026-09-06 decisions are structural: provenance stamps#source,#audienceis unconditionalBrandon(no per-file column), a signed report is standing consent for the listed lossy repairs, refused duplicates and download failures are decisions while the batch continues, and phase 2 is non-interactive. cquarry run phase1 DIRvets a downloads directory into a manifest: duplicate screening and the DRM scan drive the companion scripts, the PDF/DJVU battery rides the newscripts/check_pdf.py, bindery's phase-1 EPUB slice runs read-only (--bindery-reportcaptures its JSON), and--stamp/--apply-lossyare the file-side consents. Read-only againstmetadata.db; DRM-locked files quarantine; clean files approve.cquarry run phase2 --manifest FILEimports the signed manifest. Guards first (signed; no blocking decisions; Calibre closed; mandatory--backup-diroutside the library), then ONEbatch():add_bookfrom the manifest stamps (cquarry 1.14), the pathway reset (clear tags + rating) scoped to the imported ids,#sourceand#audiencestamped. Metadata downloads run after the commit viafetch-ebook-metadata+calibredb set_metadata; failures and ambiguities become decisions, and a post-download author clobber watch records what the download changed. Resumable: imported ids are written back and skipped on re-run.cquarry run phase3 --manifest FILEcurates the imports: the batch set is manifest ids crossed withfind_untagged, dossiers come from--book --format json, and the decision gates (tags by precedent, description rewrite, field fixes) run as TTY prompts or from an--answer-filein ONEbatch(). Thenbindery run phase3andreconcile_file_metadata.py --apply --repair-pdfdrive, the validator re-runs (exit 1 unless clean), and the prose batch record lands beside the manifest.--book --format jsonemitsget_book_dossierdicts verbatim as a JSON array (raw comments included) - phase 3's structured input.scripts/comments_census.pystands up the phase-3 skill's inline three-liner: a read-only census of the mechanical description defects (double hyphens, spaced-hyphen dashes, markdown bold, tag debris, body shape, soft hyphens, zero-width characters, mojibake, lost ligatures, exact-duplicate bodies) with--idscoping and a--jsonreport.- Skill sync + the pathway amendment. The library
CLAUDE.mdcarries the 2026-09-06 amendment (phase 2 tool-driven, sign-off = the signed manifest, the ratings ban as an anti-library-wide-predicate rule); the phase-1 skill namesrun phase1(its manual commands stay the appendix), the phase-3 skill namesrun phase3and the set forms. - Suite 280 -> 294.
- One target set, many
--batch-*verbs, one transaction. The write side catches up with Phase 15's batch-shaped read side: exactly one target source per invocation (--ids 1,2,3,--from-search 'EXPR'resolved read-only through the search engine,--from-untaggedreusingfind_untagged, or--from-manifest FILEwith ids one per line or comma-separated) feeds id-less verbs: add/remove/clear tags, clear rating, set/clear column, add-column-value (append for multi-valued columns, via cquarry 1.13'sadd_custom_column_values), set/clear pubdate, set title/authors/publisher/languages/series (+--series-index), set/clear identifier, set cover, remove format. Hand-supplied ids are validated read-only before anything opens writable: unknown ids are reported and the run aborts (exit 2). Deletion has no set form:--remove-bookstays per-book. - Dry-run by default. A set write prints the verbatim target, the
resolved id count and list, and a per-verb preview; nothing opens
WritableCalibreDBwithout--apply. Apply demands a closed Calibre (the anchoredpgrep ^calibreguard, thefetch_library_codes.pyprecedent) and a mandatory--backup-diroutside the library directory (thestamp_pdf.pyprecedent), then commits the whole pass as ONEbatch()transaction: any failure rolls everything back.--commit-per-bookis the documented non-default escape hatch for very large sets. - The rating carve-out is mechanically encoded. The library
NON-NEGOTIABLES ban bulk rating edits, so
--batch-clear-ratingis accepted ONLY with--from-manifest(the manifest proves which ids that run imported);--ids,--from-search, and--from-untaggedare refused with exit 2. The column verbs refuse#reading_status,status, anddate_readby label (case-insensitive), belt-and-braces on the absolute ban. - Honest reporting. The write path's action builders now carry
cquarry's
changedreturns, so multi-verb summaries reportapplied:/already-so:per verb instead of a blanketok:(the all-or-nothing rollback contract is untouched). Set mode reports per-verb applied/already-so/failed counts plus the per-id failure list on stderr, and--format jsonemits{target, verbs, results[{id, verb, status, detail}], committed, dry_run}; a failed row rolls the pass back and reportscommitted: false(exit 1). Exit 0 committed or dry-run, 1 failures or lock, 2 usage. --set-rating ID 0now clears (behavior change). It used to write a phantom 0-rating row that reads as unrated everywhere while polluting the ratings table; 0 now routes through the true clear (set_rating(id, None): link deleted, orphan pruned), matching Calibre's own 0-stars semantics. This amends Phase 16's no-single-verb-change non-goal for exactly this one case, Brandon's call (2026-09-06); the other three open questions (flag naming, manifest-only carve-out,--backup-dir) landed as specced.- Skill sync. The
phase-3-importskill names the set forms where it taught per-id CLI loops, with the rating carve-out called out so rating changes there stay per-book. - Suite 225 → 257.
- Second-level genres (and beyond) in the breakdown.
--analytics genresnow takes--genre-depth N(default 1, the previous root-only shape). Depth 2 renders each root's children indented beneath it, labeled by their last path segment (SciFiunderFic, not the fullFic.SciFipath); deeper flags go deeper. No cquarry change was needed:genre_distribution()already returns every node of the hierarchy tree-ordered, so this is renderer slicing only. - Every level stays a share of the whole library. A child's
percentage is its fraction of all books, not of its parent, so
children need not sum to their parent (a book with both
Fic.FantasyandFic.SciFicounts once toward each). The TUI's Genre Breakdown entry prompts for the depth.--genre-depthbelow 1 is a usage error (exit 2). - Suite 216 → 219.
- Genre breakdown by percentage.
--analytics genresrenders cquarry 1.12's newgenre_distribution(): every top-level genre (the root of the dot-path tag hierarchy) as a share of the whole library, biggest first, with a bar per genre. A book counts once per genre even when several of its tags share an ancestor; a book tagged into two roots lands in both, so the shares are honest fractions of the library and can sum over 100% (the output says so). Untagged books get their own row when present. The deeper levels of the hierarchy stay where they belong:--analytics tagsremains the taxonomy tree. - TUI menu. The Analytics section gains a "Genre Breakdown" entry (between Tag Tree and Wing Overlap) backed by the same renderer through the pager. Shares and ordering are cquarry's math; percentages, bars, and the caveat are this repo's formatting, per the frontend-only split.
- Suite 213 → 216.
--bookbatch forms.--booktakes a comma-separated id list (--book 8884,8885,8886) and composes the dossier renderer in a loop; a bare--book --untagged(no ids) selects every untagged book through cquarry'sfind_untagged()— the phase-3 curation entry state, which used to mean a hand-rolledget_book()loop per batch. An unknown id inside a list renders the books that exist and exits 1, matching the single-id behavior; conflicting or empty forms are usage errors (exit 2).--show-customspeaks both names. cquarry 1.9's dual resolution flows straight throughload_custom_column, so--show-customnow accepts the display name (Status), the bare label, or the#label(#reading_status); the docs' old "asymmetry" warnings are rewritten. No code change was needed beyond the cquarry bump — the flag routed throughload_custom_columnall along.--clear-identifier BOOK_ID TYPEexisted since 3.19.0 but reached the deletion asset_identifier(..., ""); it now calls cquarry 1.9's explicitclear_identifier()(same observable behavior, honest naming).fetch_library_codes.pygains--all-codesand a misses worklist, and its write path moves inside the sanctioned one.- LCC and DDC arrive in the same SRU response, so
--all-codesstores both in ONE pass and selects every book missing either code; the old flow (an LCC--apply, then--apply --write-ddcbehind it) created exactly the concurrent-writer contention that bit the 2026-09-02 import. - Bug fix riding along: the DDC write was nested inside the LCC-hit branch,
so a book the catalogue has DDC but no LCC for was silently dropped even
under
--write-ddc. The writes are now independent. - The writer is no longer a raw
INSERT OR REPLACE: it goes throughcquarry.write.WritableCalibreDB.set_identifierin onebatch()transaction, so touched books land inmetadata_dirtied(Calibre regenerates their sidecar .opfs) andlast_modifiedmoves — neither happened before. Retry/backoff overdatabase is lockedrides on top of the module's 30s busy timeout; identical values are honest no-ops, so a--refreshre-run reports the real change count. - Every book ending the pass with no LCC is written to a misses worklist
(
--misses-file, defaultfetch_library_codes_misses.txt) asid<TAB>isbn<TAB>ddc<TAB>title— the manual-research pass starts from a file instead of terminal scrollback (2026-08-27: three misses tracked by hand). The worklist is a report artifact and is written in dry runs too.
- LCC and DDC arrive in the same SRU response, so
- Version/docs re-sync. This roadmap's header catches up (was "as of v3.23.1"); README, spec's companion table, and CLAUDE.md carry the new surface. Phase 15 is fully ticked; the phase-3 skill names the batch forms and the one-pass fetch invocation.
- Tests: 201 → 213 (batch/untagged forms and exit codes,
--show-custom's both-forms catalog render,--all-codesselection semantics, the misses worklist, andwrite_identifiers' change-counting, dirtied queueing, and wait-out-a-concurrent-writer behavior).
- Feature:
scripts/screen_duplicate.py— the phase-1 duplicate screen as a tool. Reads each download's embedded title/authors/ISBN withebook-meta(filenames are display hints only; they lie), matches exact ISBN first, then normalized title + first author over cquarry's tight=search with Python-side normalization (accent/case fold, leading-article strip, subtitle/edition scrub — "Capital: Volume I" screens against "Capital: A Critique of Political Economy"), and screens within the batch too. Prints the comparison columns (existing id/title/authors/formats/size/pages) the skill says to report before judging keep/upgrade/re-source;--format jsonfor the batch report; report-only, exit 0/1/2. - Feature:
scripts/stamp_pdf.py— the phase-1 pre-stamp as a tool. Fixed exiftool field set (Title/XMP-dc:Title, Author/XMP-dc:Creator, XMP-dc:Publisher, ISBN via-Keywords=isbn:...— never-XMP-dc:Identifier, which Calibre maps to a bogus doi); multi---authorjoins with " & " (documented opposite of--set-authors';); dry-run by default with the filename-derivation preview beside the requested stamp;--applyrequires--backup-dir, refused if it resolves inside any target's directory (a backup beside the file gets imported); verifies withebook-metaand prints STAMP_FAILED on read-back disagreement instead of re-fighting a stubborn XMP store. Mechanics only — value choice stays the agent's research step (stated in the docstring). - Docs: spec §5's companion table gained both scripts; README gained a companion-scripts section; the phase-1 skill names both tools in its duplicate-screen and pre-stamp steps.
- Dependency: Requires the latest cquarry (1.8.x).
- Upgrade:
--bookrenders over cquarry 1.8'sget_book_dossier()— the composed deep fetch lives in the library now, and the dossier gains a publication date in the facts line (published YYYY-MM-DD, skipped while the0101undefined-date sentinel stands). Output is otherwise byte-identical to 3.23.x. - Upgrade:
--audit's per-book predicates (untagged, unrated, authorless, formatless, deprecated-format-only, coverless, missing cover file, low-res cover) arecquarry.integritycalls — one shared definition of "incomplete" across the ecosystem. Verified byte-identical CSV against 3.23.1 on the real library (7,641 rows). Duplicate grouping stays inline on purpose: the CSV joins ids in scan order andfind_duplicate_books()sorts numerically. - Upgrade:
--analyticsand--statsrender overcquarry.analytics(author_stats,addition_timeline,rating_distribution, and--analytics overlapovervl_overlap, unparseable wings still skipped); output verified identical on the real library for every subcommand.author_statsgained an additiverated_countkey so the "(N rated)" rendering survives the move. - Upgrade: The LibraryThing exporter and
scripts/audit_isbns.pyusecquarry.helpers' ISBN family (isbn_normalize/isbn_check_digit_is_valid/to_isbn13) instead of two divergent local copies. The exporter mapsNoneback to""so its CSV stays byte-identical;audit_isbns.pynow reportsnullfor a stored identifier that is not a 10/13-digit number (the old copy passed normalized junk through — junk matched nothing anyway). Requires cquarry 1.8. - Docs: README's
--bookrow notes the publication date; CLAUDE.md records the rendering-over-cquarry-modules contract.
- Fix:
scripts/fetch_library_codes.py's "Calibre is running" guard now matches every Calibre process NAME (anchoredpgrep ^calibre: the GUI,calibre-debug, and thecalibre-paralleljob workers, whose comm truncates tocalibre-paralleand could never match the old exact-name-x "calibre"). The probe still never reads command lines, so a concurrent Bindery sweep mentioning "Calibre Library" in its args cannot trip it. The 2026-08-27 "args false-positive" record was a misdiagnosis (name-only matching has been in place since 2026-08-09); the real defect was the missedcalibre-parallelworkers. Verified live against a running Calibre with four workers. - Chore:
scripts/fetch_library_codes.py.bakand the orphanedaudit_epub*.pycremoved (flagged as litter by the repo's own cleanup standard). - Upgrade:
tests/test_version.pynow also pins the newestpatchnotes.mdheading tocquarry_cli.VERSION, catching the patchnotes-vs-code desync class by CI instead of by eye. - Docs: spec header (3.15.0) and roadmap header (v3.12.0) current; spec §5's companion table gained the missing
audit_conversion_overrides.pyrow; spec/CLAUDE.md cquarry floors moved to the real 1.7; README's "completed software / no new features" note rewritten (it contradicted the roadmap's open phases). - Dependency: Requires the latest cquarry (1.7.x).
- Feature:
--set-pubdate BOOK_ID DATEand--clear-pubdate BOOK_ID— publication-date writes through cquarry 1.7'sset_pubdate, which stores Calibre's exact TEXT convention ('YYYY-MM-DD 00:00:00+00:00', naive input taken as UTC) and clears via the0101-01-01undefined-date sentinel. No-op honest: re-setting the same instant doesn't bumplast_modifiedor queue OPF resync. The TUI Edit Book menu gains the matching Set Pubdate / Clear Pubdate pair. Retires the raw-SQL pubdate workaround that cost 8 linter errors on 2026-08-27. - Upgrade: Multi-verb batch mode, on top of cquarry 1.7's
batch()transaction context — several write flags in one invocation (cquarry --set-title 42 "New" --set-pubdate 42 1991-10-01 --add-tag 42 Curated) now run in ONEWritableCalibreDBinside onebatch()transaction: all-or-nothing (a failure anywhere rolls back every verb,metadata_dirtiedincluded), committed exactly once, with a per-verbok:summary printed after the commit. Repeated--add-tag/--remove-tagflags join the same transaction instead of opening one per flag.--remove-bookrefuses batch combinations (exit 2). Single-verb invocations behave exactly as before. - Upgrade:
writeops.pyinternals — each verb's mutation now lives in anaction_*builder so the CLI dispatcher and the TUI executors share one implementation; argument validation (ids, rating range, cover yes/no, format size) happens up front for the whole invocation before the database opens. - Upgrade: Requires the latest cquarry (1.7.x).
- Upgrade: The TUI now delegates to vir-tui 2.2.0's Phase-3 primitives —
interactive_session()owns the curses session lifecycle (open, degrade-to-text, close, KeyboardInterrupt → exit 130),prompt_float()replaces the hand-rolled rating loop (blank now means 0/clear rather than cancel),prompt_path()powers the first-run and change-database flows (existence loop with a notice built in; the change-database cancel path keeps the "database unchanged" behavior),confirm(danger=True)gates the Remove Book flow, and report footers use the publicout_note(). - Chore: Dropped the private
_USE_CURSES/_SCREENmodule globals in favor of vir-tui's publictext_mode(). - Dependency:
vir-tuitracks@main; requires 2.2.0.
- Feature:
--book BOOK_ID— full single-book dossier composing cquarry's read APIs: identifiers, per-format files with catalogued sizes and on-disk paths (missing files flagged), cover resolution, comments rendered throughstrip_html, custom columns (via the search-enginefield()hook), e-reader annotations, per-device reading positions, plugin data, and conversion overrides. Unknown ids exit 1 cleanly. - Feature:
--entities {authors,series,publishers,tags,languages,ratings}— cquarry'sget_entities()listing with per-entity book counts and the sort/link secondary columns for authors/series/publishers (ratings render as stars). - Feature:
--reading-progress— everylast_read_positionsrow across devices with progress-fraction bars, newest first. - Feature:
--columns— the custom-column schema (label, datatype, editability, normalized flag, enum values, composite templates) viaget_custom_columns(). - Feature:
--info— library dossier: identity UUID, virtual libraries with their defining expressions, saved searches,@Nameuser categories, grouped search terms, news feeds, conversion overrides,metadata_dirtied/annotations_dirtiedqueue depth, and tag-browser layout state. - Feature: Write-verb expansion completing cquarry ≥1.5's write module —
--add-tag/--remove-tag(repeatable; orphaned entity rows pruned),--set-identifier/--clear-identifier(EAV upsert, empty VALUE deletes),--set-series+--series-index/--clear-series(clear nullsbooks.series_indexand prunes),--set-publisher/--clear-publisher,--set-languages/--clear-languages(canonicalized through Calibre's language map),--add-format/--remove-format(metadata-onlydatarows), and--set-cover(has-cover flag). All funnel through the shared dispatcher: argument problems exit 2 before the DB opens, lock/validation errors exit 1. - Upgrade: All write plumbing (previously inline in
cli.py) moved tocquarry_cli/writeops.pyso the TUI reuses the exact same executors; the TUI gains Display/Info entries, Annotations + LibraryThing exports, an Edit Book submenu covering every write verb, and a dry-run-first Remove Book flow. - Upgrade: Dependency policy — cquarry (and vir-tui) now tracked explicitly at
@main; the committeduv.lock(which pinned cquarry 1.0.0 at an old commit) is removed and gitignored so everypip/uv/CI install pulls the latest cquarry. - Upgrade: Requires the latest cquarry (1.6.x).
- Upgrade:
--searchinherits cquarry 1.6's@Name:queryuser-category location — Calibre'sget_user_category_matchesparity (exact member matching per member location,@Name:.queryfor subcategories,falseinversion, upstream's@...:lexer word rule so spaced category names work). No flags needed; unknown@Namesmatch nothing instead of degrading to anall:text sweep. - Upgrade: Inherited v1.6 read-side completeness —
get_book()rows are now shape-identical toget_all_books()rows (both carryuuid,identifiers,size),languagesfollowbooks_languages_link.item_orderlike Calibre,get_entities()gained aratingskind, andget_feeds()/get_annotations_dirtied_books()/get_tag_browser_counts()(Calibre's owntag_browser_*sidebar rollups withavg_rating) are available for future verbs. - Upgrade: Requires
cquarry>=1.6.
- Feature: Write-verb surface completed on cquarry ≥1.5's expanded module —
--set-authors ID "A; B"(semicolon-separated; author_sort recomputed),--set-rating ID STARS(0–5),--set-comments/--clear-comments,--set-column ID #LABEL VALUE/--clear-column(layout auto-detected, enumerations validated against configured values, non-editable columns refused), and--remove-book ID [--confirm-remove](dry-run by default printing title+formats; irreversible with the flag). All verbs queue OPF regeneration viametadata_dirtied. - Feature:
--format-statsprints per-format book counts and total bytes (get_format_stats()). - Upgrade: All write verbs funnel through a shared
_run_writedispatcher with uniform lock/validation error handling (exit 1) and argument validation (exit 2). - Upgrade: Requires
cquarry>=1.5.
- Feature:
--show-author-details— opt-in enrichment for--catalog,--all-wings,--export, and structured--searchoutput. Catalog lines gain a{sort; link}segment and JSON/CSV exports gainauthor_sorts/author_linksfields, sourced from cquarry ≥1.4's entity secondary columns (each author's true sort key and author-page URL). - Upgrade:
--searchnow resolves custom grouped-search terms (GroupName:query, from Calibre'sgrouped_search_termspreference, with upstream union/false-inversion semantics) and the newannotations:location (full-text over e-reader highlights) — both inherited automatically via cquarry 1.4's engine; no flags needed. - Upgrade: Requires
cquarry>=1.4.
- Feature: Page counts flow through exports —
--exportJSON gains apageskey per book, CSV apagescolumn, and the AI format a<N>psegment — sourced from Calibre's nativebooks_pages_linktable via cquarry ≥1.3. - Feature: Library provenance stamping — text catalogs (
--catalog,--all-wings) carry the library's identity UUID in their header line, and--audit's summary names it, so any output can be traced back to its source library after moves/restores (cquarryget_library_uuid()). - Upgrade: The audit's cover checks now resolve through cquarry's
get_cover_path()(canonicalcover.jpg/cover.pnglayout logic) instead of hand-built paths. - Upgrade: Requires
cquarry>=1.3.
- Feature: First write flow —
--set-title BOOK_ID TITLErenames a book through cquarry's separateWritableCalibreDBmodule (trigger-safe: registers Calibre'stitle_sort/uuid4SQL functions, refreshes the sort key, bumpslast_modified). Every mutation also records the book id in Calibre'smetadata_dirtiedqueue (requires cquarry 1.2), which is what upstream consumes to regenerate the book's sidecar.opfand re-push metadata to wireless readers on its next startup — external edits finally propagate. Requires closing Calibre first; lock contention fails cleanly with exit code 1. - Feature:
--auditoutput now includes a "Pending OPF sync" section listing the books queued inmetadata_dirtied(via cquarry's read-onlyget_dirtied_books()), so you can see exactly what Calibre will resync next time it starts. Absent when the queue is empty. - Upgrade: Requires
cquarry>=1.2.
- Feature: New
--export-annotationscommand dumps e-reader highlights, bookmarks, and notes as JSON (via cquarry'sannotationsreader), optionally scoped to one book with--id <BOOK_ID>. - Feature: New
--plugin-data NAMEflag (works with--catalog,--all-wings, and--search) appends third-party plugin values — e.g.goodreads_idorwordcountfrombooks_plugin_data— as<name: value>segments on each book line. - Upgrade: Migrated every mode off the retired string-typed
get_all_books()fields.authors,tags,formats, andlanguagesare consumed as the nativelist[str]arrays cquarry 1.1+ exposes; author names containing literal commas no longer risk splitting. - Upgrade: Requires
cquarry>=1.1. Search gains saved-search interpolation (--search 'search:"Name"'), multi-valued count operators (tags:#>2), language canonicalization (languages:English), slash date separators, tristate boolean keywords (checked/blank/...), strict errors on unknown virtual libraries, and newsize:/pages:locations — all available through the existing--searchflag with no CLI changes. - Note:
scripts/reconcile_file_metadata.pyalready fetches per-book records without a full-library scan, so the roadmap's single-record fast path was satisfied by design; no change was needed. - Feature: New
scripts/audit_conversion_overrides.pylists books carrying manual conversion recipes — cquarry’s extractor, which never unpickles the blobs.
- Fix: Replaced hardcoded custom columns logic in
librarything.pywith dynamic ID resolution from the database schema. - Fix: Restored functionality in
test_queries.shby targetingcquarry_cliinstead of the extractedcquarrymodule. - Fix: Migrated
spot_check.pyandaudit_isbns.pyto useconnect_ro()with WAL/SHM fallback for safe read-only locking against active Calibre DBs. - Fix: Corrected NULL title bug in
spot_check.pyby coalescing absent titles. - Fix: Fixed logical gap check in
display.pyseries output. - Fix: Reordered argument validation in
export.pyto prevent 0-byte file truncations on invalid format flags. - Fix: Removed
.title()enforcement during duplication checks inaudit.pyto preserve original casing.
- Refactor: Adapted to
vir-tuiv2.0.0 public API and decoupled menu fallbacks. - Fix: The non-curses fallback text menu now functions properly for CalibreQuarry by passing custom
letter_keysandaliasesduring initialization.
- TUI Extraction (
vir-tui): Extracted the generic CLI formatting (core.py) and curses menu primitives (tui.py) into thevir-tuishared repository. CalibreQuarry now depends onvir-tuifor all UI logic, ensuring perfect parity and centralized updates for all interactive prompts across the workspace.
- Shared Library Extraction (
cquarry): Extracted the core Calibre database reading layer and search expression grammar into a new, standalone Python library (cquarry). CalibreQuarry now depends on this shared library for all data access and search resolution, ensuring 100% parity across all tools in the workspace (like Hermitage and Wings). - Project Renamed to
calibrequarry: To prevent pip namespace collisions with the newly extractedcquarryshared library, the CLI project has been formally renamed tocalibrequarryinpyproject.toml, and its internal modules have been moved tocquarry_cli. The terminal command remainscquarry.
UI Upgrade: CLI scripts now feature rich output (ANSI formatting, tqdm progress bars, and a clear summary block). The project is no longer strictly stdlib-only and now depends on tqdm.
CI Configuration & Code Closures. The GitHub Actions pipeline (ruff check) failed because of several B023 late-binding closures inside tui.py which were unnoticed by the global configuration. Fixed those closures and added a test (test_version.py) to prevent version drift between pyproject.toml, config.py, and VERSION files.
LibraryThing Export Integration. Ported the standalone export_librarything.py script natively into cquarry. You can now use the --exportlt flag to generate LT-formatted CSVs directly. Crucially, this can be combined with --search to export only specific subsets of your library (e.g., --search "date:>2026-08-05" --exportlt), making targeted updates significantly easier. The export fully handles LibraryThing data quirks, such as folding ISBN-10 to ISBN-13, expanding translators into individual tags, clearing sentinel dates, and breaking output into manageable 500-book chunks split by "Read" vs "Unread" statuses.
A bug, maintenance and improvement sweep across the package and all seven companion scripts. Eight fixes, three additions, four cleanups, every one pinned by a regression test. The suite grows from 243 to 273 tests.
Every database opener built its file: URI by raw interpolation. ? and # are URI syntax, so a library at Books #2/metadata.db resolved to a different path entirely and failed with the thoroughly unhelpful "no such table: books". This affected db.py and six of the seven scripts; fetch_library_codes.py already percent-encoded its path, which is the form the rest have adopted via a shared db_uri_ro helper. The package's helper is in helpers.py; the standalone scripts each carry their own copy, as they carry everything else.
--tag scoping in audit_isbns.py and fetch_library_codes.py matched one arbitrary tag per book. Both pulled a single tag through a LIMIT 1 subquery with no ORDER BY, so a book carrying two tags was filtered on whichever one SQLite happened to return: a book tagged both Fic.Fantasy and NonFic.Tech was invisible to --tag NonFic.Tech roughly half the time, and fetch_library_codes.py's per-branch hit-rate report was attributing books to a branch at random. Scoping now considers every tag, and a book is reported under a tag the filter actually selected on. This is latent on the reference library, where all 7,439 books carry exactly one tag; it was found by reading the query, reproduced with a two-tag fixture, and fixed as insurance rather than to change any result observed so far.
compress_pdf.py --out-dir pointed at the PDF's own directory destroyed the original. The output lands at out_dir/<same name>, which in that case is the source path, so the mode documented as leaving the original untouched moved the compressed temp over it, left no .pre-compress.pdf rollback, and then printed "Original untouched at:" followed by the path of the file it had just replaced. It is now refused (exit 2), with the comparison resolving both sides so a symlinked directory cannot slip past.
write_all_wings silently overwrote colliding wing catalogs. Sanitizing a wing name to a filename is lossy: "Tabletop: RPG" and "Tabletop RPG" both reduce to Tabletop_RPG, and a punctuation-only name reduces to nothing, yielding _Library.txt. One wing's catalog replaced another's with no warning, leaving a file that claimed to be a wing it was not. Names are now de-duplicated with a numeric suffix, and a name that sanitizes away gets a positional fallback.
write_catalog could not write into a directory that did not exist yet. run_audit and run_export both create the parent first; write_catalog opened directly, so --catalog --output reports/catalog.txt exited 1 on a FileNotFoundError. It now matches its siblings.
spot_check.py --review crashed on an empty sample. With no books selected (--n 0, or a --limit that selects nothing) no chunk files were written, and the closing instructions indexed written[0]. It now reports an empty sample.
audit_drm.py and audit_epub.py had no locked-database fallback. The package, validate_metadata.py and reconcile_file_metadata.py all read from a temporary snapshot when Calibre holds the lock; these two let the OperationalError escape as a traceback, so a library-mode run while Calibre was open simply crashed. Both now degrade the same way, pinned by tests that take a real BEGIN EXCLUSIVE lock rather than faking the error.
Two writers used the locale's encoding instead of UTF-8. spot_check.py's report and bundle, and audit_drm.py's CSV, were the only file writes in the repository without an explicit encoding=; a title outside the locale's character set raises UnicodeEncodeError partway through a long run. Reachable only under a genuinely non-UTF-8 locale (en_US.ISO-8859-1 and the like) rather than the far commoner LANG=C, which auto-enables Python's UTF-8 mode: this is consistency with the rest of the repository more than a crash anyone was hitting.
--audit reports cover_file_missing. A book whose database row says has_cover but whose cover file is gone from disk was silently clean: the audit flagged a missing cover record and a low-resolution cover, but not the case where the two sources of truth disagree. Every cover consumer hits that, Calibre's own grid included.
The curses prompt accepts non-ASCII input. It read one byte at a time and discarded anything outside 32-126, so Brontë could not be typed into a search query or an output path. It now reads whole characters (get_wch), which for a library full of translated fiction is the difference between the TUI being usable for search and not.
Every external tool call is bounded by a timeout. spot_check.py (exiftool, djvused), reconcile_file_metadata.py (calibredb, exiftool, qpdf, djvused, pgrep), fetch_library_codes.py (pgrep) and compress_pdf.py's inspection probes previously had none, so a single wedged tool on a single damaged file hung a whole-library run with no indication of which book it stopped on. Each failure path is handled rather than merely raised: a hung reader flags that book and the run continues. Ghostscript itself is deliberately left unbounded in compress_pdf.py, because a legitimate 1 GB sourcebook takes many minutes and killing that mid-convert would be the wrong answer.
audit_epub.py extracted each book's rendered text twice under all: emptytext built it per spine document and ocr rebuilt the same string, which is the expensive half of a pass whose entire purpose is touching each EPUB once. It is now computed once per book and shared. Two stale documentation references to validate_library.py, a script that is not in this repository, now point at validate_metadata.py; audit_epub.py's usage line said "all three audits" when there have been four since v3.6.0. audit_drm.py imported sqlite3 inside a function while importing everything else at module scope, and stats.py carried a conditional whose two branches computed the same value.
audit_isbns.py counted any labelled ISBN as the book's own. Books quote other books' ISBNs constantly, and one citation is indistinguishable from a self-identification if you only count numbers, so a handful of famous false accusations followed: The Atrocity Archives names The New Hacker's Dictionary's ISBN in a glossary entry, Metamagical Themas lists one among Hofstadter's self-referential joke titles, and C++ Primer Plus advertises six other Sams books in its back matter.
The fix is to require corroboration rather than to enumerate the ways a citation can look. A copyright page never carries a bare number: it sits beside a copyright line, a rights reservation, a binding, a printing statement, or a CIP block. A citation carries none of that, so an ISBN now counts as the book's own only when such a marker appears within 260 characters.
Applying that symmetrically was itself a bug, caught by measuring before committing. The first version demanded self-identification in both directions and cost 639 confirmations across a real library while removing only 25 false findings. The two directions need different evidence: a book printing the same number the catalogue holds is conclusive whatever the surrounding prose says, because a citation coinciding with your own stored value does not happen. Only a different number needs to have been claimed. Restricting the test to the negative direction gives 46 fewer false findings with confirmations slightly up (2,974 to 2,979).
Both directions are now pinned by tests carrying the real passages. 243 tests.
A new companion script, audit_isbns.py, and the first new capability since the v3.8 sweep. It answers a question nothing else in the Calibre ecosystem asks: does the ISBN stored against a book actually identify that book? Calibre downloads metadata but never re-examines what it stored, so a wrong ISBN stays invisible, and an ISBN is what other systems key on when you hand them a catalogue.
The motivating evidence: a four-source sweep of a 6,786-ISBN library that was already validator-clean found 51 identifiers pointing at a different book. The dominant shape is a same-publisher sibling, which is why the defect survives every existing check: the number is well-formed, the checksum passes, and only the book it names is wrong. Programming Clojure carried tmux 2's ISBN, Spelunky carried Super Mario Bros. 3's, and A Book on C carried 9782147483649, the 2147483649 integer-overflow constant dressed as an ISBN.
This release ships the offline half: verification against the ISBN each book prints on its own copyright page. That is the best authority available for exactly the books no bibliographic database has heard of (small-press RPGs, indie ebooks, print-on-demand reprints), and it needs no network, no credentials, and no new dependencies.
It reads body text only, never embedded metadata. reconcile_file_metadata.py writes the database's values into those metadata blocks, so comparing against them would be comparing the database with itself and would cheerfully confirm every error the tool exists to find. A test pins this: an OPF carrying a matching identifier must not produce a confirmation.
Not crying wolf is most of the work. Three benign things resemble a mismatch and are classified apart. A bibliography prints other books' ISBNs (The Art of UNIX Programming prints 49), so above --max-printed distinct numbers a file is read as a citing work. A bundle or series volume legitimately prints several ISBNs, reported AMBIGUOUS for a human to resolve rather than guessed at. A format variant (print versus ebook) differs only in its final digits, so a printed number sharing the stored one's registrant prefix is reported VARIANT rather than MISMATCH; a genuinely wrong ISBN almost always comes from a different publisher block entirely.
There is deliberately no --apply, and there will not be one. Single-source verdicts proved wrong often enough during the motivating sweep that an auto-fixer would have "corrected" Curse of Strahd, Cold Mountain, Kitchen and The Master and Margarita, every one of which was already right.
Scoping reuses fetch_library_codes.py's anchored-hierarchical --tag rule rather than inventing a virtual-library flag: a comma-separated prefix list covers a multi-root wing without coupling a standalone script to the package's VL resolver. Text extraction is stdlib zipfile for EPUB and optional pdftotext/djvutxt for PDF and DJVU, whose absence is reported rather than fatal. MOBI/AZW3 are skipped.
Three things were corrected before merge, all found by running it against a whole 6,783-ISBN library rather than a single wing.
The severe verdict was renamed SUSPECT to MISMATCH, because the old name overclaimed. On a real library the commonest reason a book prints a different ISBN than the catalogue holds is that the file is a different edition, which is worth knowing but is not a wrong book. An ISBN alone cannot separate the two, and a report that says "suspect" 100 times about mostly-benign edition drift trains its reader to ignore it.
The publisher-prefix comparison was too long. Registrant lengths run 2 to 7 digits and are inversely proportional to publisher size, so a nine-digit prefix (right for a one-book press) split HarperCollins from itself: Sabriel's stored 978-0-06-447183-1 and printed 978-0-06-000548-1 are one publisher and were being reported as rivals. Six digits reclassified 28 findings from the severe bucket to VARIANT.
UNREADABLE conflated two opposite things. Half of those books were MOBI/AZW3, which this tool skips by design; the rest were supported formats that yielded nothing, which may be damaged files. They are now SKIPPED and UNREADABLE respectively.
Also documented: the printed ISBN can itself be wrong, which is the limit of the whole premise. The TSR Forgotten Realms Campaign Setting prints a number belonging to The Jungles of Chult (a permanent typo), and Night Witches prints Durance's because a small press reused its previous copyright page. Both surfaced as VARIANT and both resolved in favour of the database, which is exactly why nothing is ever written automatically.
Runs: 121 books in 15s on one wing and 468 in 63s on another, both with zero severe findings; 6,783 in about 15 minutes across the library. The suite grows from 208 to 237 tests, with the real-world classifications pinned so a future refactor that silently reclassifies them fails loudly.
A full-repository bug, maintenance, and documentation sweep: the package, all seven companion scripts, the tests, and every doc. The package core came out clean (one micro-refactor: _num_predicate in search.py returned a two-tuple whose second element nothing read; it now returns just the predicate). The scripts yielded nine real fixes, every one now pinned by a regression test. The suite grows from 195 to 208 tests, and fetch_library_codes.py and validate_metadata.py gain their first tests.
audit_drm.py reported an unparseable PDF as CLEAN. qpdf --is-encrypted exits 0 for encrypted, 2 for not encrypted, and 3 for a file it cannot parse; the code collapsed 2 and 3 into "clean". A corrupted or truncated PDF (exactly the kind of loose file the tool exists to vet) therefore skipped the trailer /Encrypt fallback that was built for the couldn't-classify case, and a Standard-encrypted broken file read as definitively DRM-free. Exit 3 now defers to the fallback and reports BENIGN encrypted-unclassified when an /Encrypt dictionary is present.
audit_epub.py content contradicted itself on an injected signature in a declared-foreign book. The expected-foreign flag (declared language / NonFic.Language.* tags) was applied to injection-signature hits too, so a piracy notice in a legitimately-French book printed under a green "(expected-foreign)" label, reported "0 file(s) need review", and still exited 1. A signature is a defect regardless of language; it is now always counted and listed as needing review.
compress_pdf.py skipped verification when it mattered most. page_count() returns None both when pdfinfo is missing and when pdfinfo cannot parse the file, and the page-count comparison only ran when both counts were known. A Ghostscript output so broken pdfinfo could not read it (the strongest possible bad-conversion signal) therefore passed "verification" and replaced the original. When the original's page count is known and the output's is not, the run now aborts with the original untouched.
reconcile_file_metadata.py split author names on commas. The file-side author string was split on &, ;, and ,, but ebook-meta only ever joins authors with &; a comma belongs to the name itself. "Martin Luther King, Jr." parsed as two bogus authors, never matched the database, and the book reported as drifted forever, surviving every re-embed. Same shape as the v3.8.0 identifier-space bug, one field over. The split now uses & and ; only.
spot_check.py --record silently truncated comma-delimited notes. A verdict line's note field was parts[4] of an unbounded split, so a comma-mode note of "wrong author, should be Jane Doe" recorded as "wrong author" with the rest dropped and no error. The split now caps at five fields, keeping free-text notes intact.
spot_check.py --record no longer runs without --against. The id reconciliation is documented as the load-bearing part of review mode ("nothing is written unless the ids reconcile"), but --against was optional, and omitting it skipped the check entirely, accepting exactly the short-but-plausible verdict lists it exists to refuse. Recording without an ids file is now a setup error (exit 2).
fetch_library_codes.py clobbered its own backup. The pre---apply backup was stamped with the date only, so a second run the same day overwrote the first run's restore point, the copy that actually holds the pre-change database. The stamp now carries seconds plus a collision counter; no backup is ever overwritten. Also fixed: a missing metadata.db exited 1 via sys.exit(str) while the documented contract (and every other setup-error path) says 2.
validate_metadata.py inflated FORMAT_FICTION_PDF counts. The check joined through the tags table before filtering, producing one warning per (book, tag) pair; a crossover book tagged Fic.Fantasy and Fic.Horror was reported twice. Now one warning per book, listing every matching tag.
audit_epub.py's latin-1 decode fallback was dead code. bytes.decode("utf-8", "replace") can never raise, so the documented fallback was unreachable and the except path re-read the same corrupt entry just to fail again. A corrupted spine entry now reads as empty text through one honest path, and the docstring says so.
test_queries.shexercisedvl:"The Tabletop", a wing that no longer exists (split into the two "Tabletop:" wings in the live library), so the VL-resolution smoke was silently testing the unknown-VL path. It now targetsFantasy Wing.- All seven scripts are executable now (five carried a shebang without the bit; the documented
python3 scripts/...invocation is unchanged). - Docs drift closed across the board:
spec.md§5 gains the missingfetch_library_codes.pyrow and the--reviewhalf ofspot_check.py(and now says three scripts write, not two);roadmap.mdgains Phase 7 recording the v3.7.0–v3.8.0 companion work; the README's test-suite section describes all seven test files instead of the original two;CLAUDE.md's architecture tree addsaudit_drm.py,spot_check.py,modes/tags.py, and the five test files it didn't list, and renames the long-goneaudit_epub_content.pytoaudit_epub.py.
New companion script fetch_library_codes.py: derive Library of Congress Classification codes from the LoC SRU catalogue and store them as identifiers. Written after the existing Calibre plugin for this job, "Library Codes - SRU", was diagnosed as unable to do it at all.
The plugin fails two independent ways. It refuses composite custom columns outright (library_codes_dialog.py validates cc_datatype != "text" and clears the active flag), which is fatal if you store the code as an identifier and project it with {identifiers:select(lcc)}. And its ISBN lookup queries the LCDB index dc.identifier, which the server has never supported: it answers with SRU diagnostic 1/16, 1007 (Bib-1 114 Unsupported Use attribute), "Unsupported index". The index that resolves an ISBN there is bath.isbn, confirmed live against lx2.loc.gov:210 alongside dc.title and dc.creator, which do work and are what the plugin's author/title fallback happens to use. So the plugin's primary path, the one its own description advertises, has been dead while the fallback quietly carried it.
Storing to identifiers rather than to a column value is the deliberate difference. Identifiers are the catalogue's canonical home for external keys, a composite column displays the value with no second copy to drift, and reconcile_file_metadata.py already carries identifiers into embedded file metadata. Books that already hold an lcc are skipped unless --refresh is passed, so the tool is naturally incremental against a growing library.
Two operational facts shape the design. The Library of Congress rate-limits harder than the plugin's 1.6s pacing suggests: at 0.6s between requests it began resetting connections after roughly twenty queries, and a clean 60-query run at 2.0s posted zero errors. Pacing therefore defaults to 2.0s with exponential backoff, the run aborts after eight consecutive failures rather than hammering a service that has stopped answering, and every result including a miss is cached to ~/.cache/cquarry/library_codes.json. That cache is what makes a multi-hour pass survivable: an interrupted run resumes for free, and a repeat of the same 12 books fell from 30s to 0.09s.
Coverage is partial and worth knowing before committing hours to it. Measured on a 7,362-book reference library, a clean 60-book random sample of ISBN-bearing books resolved 18, or 30%. The average is misleading, because the distribution is steep: NonFic.Tech hit 5/7, while Fic.Fantasy managed 3/16 and Fic.Contemporary, Fic.Horror, Fic.Translated and Gaming.TTRPG returned nothing at all across the sample. LoC catalogues academic, technical and canonical trade titles well and genre fiction, indie and small-press releases, translations and tabletop material poorly. The dry run is therefore the default and reports a per-branch hit rate, and --tag scopes a pass by anchored-hierarchical tag prefix so the dense half can be run without paying for the sparse half.
--apply backs metadata.db up to the sibling .backups directory before writing and refuses to run while Calibre is open. INSERT OR REPLACE is correct here specifically because identifiers is UNIQUE(book, type), unlike the enum custom-column link tables where the same statement would append a second row.
New: spot_check.py --review, a judgement mode for the checks no pattern can make. The mechanical lints in that script decide whether a field has the wrong shape. They cannot decide whether a title is the right title, whether an author field holds the person who wrote the book, or whether a description describes this book rather than another one. Review mode puts those three questions in front of a reader in a form that can be answered in bulk: full title, authors, context line and complete description, emitted in numbered chunks with a matching .ids file, answered with a verdict file, and accumulated in a ledger.
The ledger is what makes it worth doing more than once. Reviewed books drop out of later samples, so repeated runs converge on full coverage while the random sampling that makes a partial pass statistically meaningful is preserved. --worklist prints the BAD entries as a punch list; nothing is ever written back to metadata.db.
Recording refuses unless the verdict ids reconcile exactly with the ids emitted, in both directions, and refuses again on any verdict word that is not OK or BAD. This is the load-bearing part rather than a nicety: a reviewer working through hundreds of records drops some, and a short list looks identical to a complete one. On the reference library the same failure has now been recorded five times against agent output, once during the run that motivated this mode. Six tests cover the drop, invent, malformed, roundtrip, worklist, and chunking paths; the suite grows to 70.
Also fixed: reconcile_file_metadata.py truncated any identifier value containing a space, and reported the book as drifted forever. parse_identifiers split the ebook-meta output on [,\s]+, commas or whitespace, so lcc:BF637.S4 G63 2007 parsed as lcc = "bf637.s4" with the rest silently dropped. The parsed value never matched the database, so the book was reported as drifted no matter how many times it was successfully re-embedded. The split is now on commas alone, which is the format ebook-meta actually emits.
This is a latent bug rather than a new one, and it went unfound because nothing could reach it: every identifier type in use until now (isbn, goodreads, storygraph, google, amazon, oclc) is space-free. Library of Congress call numbers are the first that are not, so writing lcc identifiers is what exposed it. Measured on the reference library, a reconcile pass over 655 freshly embedded books reported 469 of them as still drifted before the fix and 1 after, that one being an AZW3, which genuinely cannot carry arbitrary identifier types. Tests grow to 69.
One deviation worth recording: the XML comes from a plain-HTTP endpoint and is parsed with xml.etree.ElementTree. defusedxml is the conventional hardening and is not an option in a stdlib-only project, so the response is instead capped at 8 MB before parsing, which bounds the entity-expansion exposure without a dependency. ElementTree does not resolve external entities, so there is no XXE path.
spot_check.py reported a complete book as EPUB_EMPTY_SPINE when the package used the legacy OEB 1.0 namespace. check_epub resolved the manifest and spine with a hardcoded {"o": "http://www.idpf.org/2007/opf"}, so any package declaring http://openebook.org/namespaces/oeb-package/1.0/ instead (OverDrive-era conversions) matched nothing at all: no manifest, no spine, and therefore a HARD failure and a nonzero exit code on a book that opens perfectly. Found by the first full-library pass, which flagged exactly one hard failure across 7,339 books, #8048 Dying Inside: 31 content documents, 446,083 characters of body text, a spine listing every one of them, and a checker that could not see any of it.
Manifest items and spine itemrefs are now matched by local element name through a _by_local_name helper, so the package's declared namespace stops mattering. This is the same root cause as the known audit_epub.py emptytext false positive on legacy OEB files; that analyzer is untouched here and still carries it.
Tests grow to 64 (an OEB 1.0 package resolves its manifest and spine).
Three spot_check.py correctness fixes and one new advisory flag, all found by running the checker against the 7,339-book reference library and then auditing what it did not catch.
_MOJIBAKE missed the commonest lead byte of all, â. The pattern enumerated individual accented characters (Ã[©¨¤¶¼£±]), which covers é/ã but not â, the double-encoded form of a curly apostrophe and by far the most frequent mojibake in scraped blurbs. One live case was missed on the reference library (#2658 Observability Engineering, youâ??re doing). Replaced with [ÃÂ] followed by anything in U+0080-U+00BF: that band is never valid text, because a real Portuguese à is followed by an ASCII vowel (Ãvila) and never by latin-1 supplement punctuation. Verified against ten cases including São Paulo, café society and Ãvila, which must stay clean.
The comment, title, and author lints read raw markup instead of text. lint_comment stripped tags but never decoded HTML entities, so a description of & repeated forty times measured 200 characters and cleared the 120-character stub gate on 40 characters of real content (one live case). The same blindness hid any mojibake stored in entity form. A new plain_text() helper strips tags, drops script/style bodies, decodes entities, and normalizes non-breaking spaces; the lints and the review bundle's blurb excerpt all route through it, so every check now sees what a reader sees.
OPF hrefs were resolved without URL-decoding. Manifest hrefs are percent-encoded per the EPUB spec, so a content document whose filename contains a space arrives as %20 and never matches the zip namelist, which reads as a missing spine item and therefore a HARD failure and a nonzero exit. No book in the reference library trips it (checked all 4,890 EPUBs: zero false positives cleared by decoding), so this is a latent correctness fix rather than an observed one, but the failure mode is silent and severe enough to close. Fragments are stripped before resolution as well.
Advisory COMMENT_TRUNCATED: a description that stops mid-word. The motivating defect, found 2026-07-30: a metadata download returned four of six Jacqueline Carey blurbs cut off mid-sentence (Blessed Elua founded Terre d'A, Raphael de Mereliot, her manipul). Nothing caught them, because every existing check passes: the text exists, is well-formed, and is long (977 to 2,265 characters). Only the final word gives it away. A whole-library sweep then found roughly fifteen more.
The check is deliberately gated and deliberately advisory. It needs a wordlist and reads /usr/share/dict/words on exactly the terms the PDF and DJVU checks already use for exiftool and djvused: used when the system has it, silently skipped when it does not, never a Python dependency. Descriptions legitimately end without punctuation all the time (blurb attributions, series lists, contents dumps), so the heuristic requires the final word to be lowercase, absent from the wordlist, ASCII, preceded by whitespace, in prose of at least fifteen tokens, and not part of a URL or a numbered contents line.
Measured honestly on the reference library: 22 flagged out of 7,339, of which 6 are confirmed truncations (27% precision, 43% recall against a hand-built set of 14). The residual false positives are systematic and worth knowing before trusting a flag: modern technical vocabulary absent from an old wordlist (microservices, autoscaling, lifecycle, asyncio), and truncations whose final fragment happens to be a real word (the co, on the st) are missed entirely. It is a lead generator for the judgment pass, not a verdict, which is why it is not in HARD and does not affect the exit code.
Tests grow to 63 in tests/test_scripts.py (mojibake lead-byte coverage with the Portuguese and French negatives, entity decoding and the stub-gate skew it caused, and truncation detection with its proper-noun, URL, and complete-prose negatives).
audit_epub.py gains a fourth analyzer: ocr, flagging OCR/conversion-damaged prose. The defect class none of the existing three catches: a bad conversion splits paragraphs mid-sentence at line-wrap or page-break positions, so the text reflows broken ("could just make out the shape" / "of another boat"). The motivating case was a damaged Jingo EPUB that measured 80 such splits where a clean edition of the same text measured 0. The primary signal is exactly that split shape (a paragraph ending without terminal punctuation, the next starting lowercase, paired only within one spine document), reported as a per-book rate normalized by paragraph count. all runs it inside the same single decompression pass as the other three.
The hard problem was separating damage from intentional style, and rate alone cannot do it: deliberately unpunctuated literary prose (Fosse's Septology, Evaristo's Girl, Woman, Other, Kingsnorth's The Wake, Faulkner) posts split rates far above genuinely damaged books. The discriminator that works is where the fragment ends. Style breaks at clause boundaries ("...she started out in theatre"); damage breaks at line-wrap positions, which land on function words ("sat the disembodied" / "heads who were..."). On the reference library every style book measured at most 11% of splits ending on a function word and every hand-confirmed damage case at least 26%, so the flag gate requires 25%, alongside minimum-splits, minimum-paragraphs, and rate floors. The function-word set is a small closed list, the same spirit as the content analyzer's stopword votes, not a dictionary.
Five false-positive idioms found during validation are guarded explicitly: paragraphs interrupted by a rendered figure (an inline formula or card-diagram image reads as a split otherwise; _Blocks now records when an image falls between two blocks' text), display math set as text (mostly non-alphabetic fragments), back-of-book indexes rendered as paragraph blocks (the "See also" signature), epistolary sign-offs (a dangling short unterminated fragment), and the block-quotation idiom of academic prose (pairs into or out of a <blockquote>, plus attribution fragments ending on "that"). Secondary signals are reported but never gate: en-dashes embedded inside words (bottom–feedin'), doubled opening quotes (' 'Course, with the first quote required to follow whitespace so British single-quote dialogue's close-then-open sequences stay invisible), and space-stripped proper nouns recurring alongside their hyphenated form (AnkhMorpork vs Ankh-Morpork).
Hand-validated against the full 4,605-EPUB reference library, every flagged book inspected: 105 flagged, 104 confirmed damage, one borderline residue (display quotes publisher-styled as plain paragraphs, indistinguishable without CSS). Documented out of scope: character-substitution errors ("sonic" for "some") need a wordlist the stdlib-only contract rules out, and damage whose signature is word truncation or whitespace corruption rather than paragraph splitting. Tests grow to 60 in tests/test_scripts.py (split detection, dialogue-fragment and scene-break non-splits, image-interrupted pairs, style-vs-damage discrimination, threshold boundaries, and an all run including the new analyzer).
The persistent-curses-screen rework, closing the last item in the roadmap's "Port from the Lattice TUI audit" section (Lattice T7, shipped there as v4.10.0). Purely a lifecycle change; no menu, prompt, or mode behavior differs.
One curses screen per session. Every menu, prompt, pause, and pager used to be its own curses.wrapper init/teardown, so multi-prompt flows visibly flashed to the shell between widgets. interactive_menu now opens the screen once and every widget draws into it (_with_screen); a widget invoked outside a session still gets its own one-shot wrapper, so nothing changes for direct callers. _reset_terminal becomes a no-op while the session screen lives (running stty sane under curses would undo cbreak/noecho beneath it). Measured under a pty: a full menu, prompt, Esc-cancel, mode, pager, quit session enters the terminal's alternate screen exactly once, where the same flow used to enter it once per widget.
One degradation path. A curses failure at session startup or mid-session (unknown terminal, capability lost) funnels through _degrade_to_text: the screen is suspended once and the rest of the session runs the text fallback. All the v3.3.2 guarantees hold (no stuck terminal, no silent exit 0); a pager that dies mid-display now also degrades and prints the mode's output as plain text instead of eating it. Ctrl-C at the menu (curses or text fallback alike) ends the session cleanly with exit code 130 and the screen restored, while EOF at the text menu stays a quiet Quit.
tests/test_tui.py grows to 29 cases; the session lifecycle was additionally verified end-to-end under a pty (alternate-screen count, degraded startup on an unknown TERM, and a full cancel-then-run flow against the live library).
The rest of the Lattice TUI audit ports (roadmap section "Port from the Lattice TUI audit": T2, T4, T6, and the fallback-menu generation). Lattice shipped all of these in its v4.9.0; this release keeps the two shared curses skeletons aligned.
Esc in a prompt now cancels back to the menu instead of accepting the default (behavior change, port of Lattice T2). In menus Esc has always meant back; in prompts it meant "accept the default", so mis-selecting a mode and mashing Esc launched it with all defaults. A cancelled prompt now unwinds the whole prompt chain back to the menu with nothing launched; bare Enter still accepts the default, and Ctrl-C or EOF at a prompt cancels the same way (the text-fallback prompts used to exit the whole program on Ctrl-C). The hint bar now reads "Enter Accept, Esc Cancel, Ctrl-U Clear". Cancelling the first-run prompt exits without persisting anything (exit code 1, the CLI's no-database signal, so unattended runs never read as success), and cancelling "Change database path" leaves the saved path untouched.
Output-file prompts expand ~ (port of Lattice T4). ~/reports/x.txt used to create a literal ./~/reports/ directory, since the TUI has no shell to expand it. Every output path the TUI collects now goes through a shared _prompt_out() that expands the tilde but does not absolutize, so relative paths keep meaning the current directory. The results pager also gains a footer naming the resolved absolute path, so "where did my report go" answers itself; the footer is suppressed when the mode errored or was cancelled, so it never claims a file that was not written.
The no-curses fallback menu is generated from the same sections the arrow-key menu renders (port of Lattice's v4.8.1 fix). The numbered listing and its key map were hand-maintained twins of _MAIN_SECTIONS, exactly the pattern that silently desynced in Lattice (fallback keys dispatching the wrong modes). _build_fallback now derives both from the sections, the word aliases ("catalog", "stats", ...) stay as an explicit supplemental dict, and tests pin that every entry is reachable with its number matching its label.
Widget paper cuts (port of Lattice T6). A bad answer at a number prompt re-asks with a "not a number, try again" note instead of silently using the default. On terminals shorter than the menu, the box shifts so the selected row stays visible instead of clipping blind below the bottom edge. The pager gains horizontal panning (arrow keys or h/l) with ellipsis markers on lines that continue off-screen, computes its width once instead of on every keypress, and keeps the scroll position valid across resizes. Ctrl-U clears a prompt field, so editing a long pre-filled path no longer means backspacing through all of it. And the export format prompt validates its answer against json/csv/ai, re-asking instead of accepting any string.
tests/test_tui.py grows from 7 to 24 cases, pinning all of the above.
TUI hardening ported from the Lattice TUI audit (2026-07-01). cquarry/tui.py shares its curses skeleton with Lattice's; the two carry-overs from that audit's high-severity findings land here (roadmap section "Port from the Lattice TUI audit", items H7 and H6's exception-boundary half). First: a curses init failure no longer reads as Quit. On capability-poor terminals (TERM=vt100, dumb terminals) the color setup or curs_set raised curses.error, which the menu loop treated as the user quitting, so the TUI silently exited 0 even though the text fallback menu works. Cosmetic capabilities are now non-fatal (a monochrome TUI beats a dead one), and a real curses.wrapper failure flips the session to the text fallback menu instead of exiting. Second: _run_with_capture now has an exception boundary. A mode error used to escape as a raw traceback and lose the captured output; it is now paged under an [Error] heading with the traceback plus whatever was captured, and Ctrl-C pages a [Cancelled] notice the same way. New tests/test_tui.py (7 cases) pins both behaviors. The remaining Lattice carry-overs (Esc-cancels-prompt, ~ expansion in output prompts, generated fallback menu) stay on the roadmap until their Lattice counterparts land.
audit_drm.py no longer flags a freed EPUB on a leftover marker file. v3.3.0 treated the mere presence of META-INF/rights.xml (Adobe ADEPT) or sinf.xml (Apple FairPlay) as DRM. But those are token/voucher files, not the lock itself: the actual lock is content encryption, which a DRM'd EPUB records in encryption.xml against its XHTML. When a book is freed, the content is decrypted but the marker can stay behind, so a bare marker with no content encryption is a residual artifact, not a locked book; it reads and embeds fine. The first whole-library sweep surfaced exactly one such case (Warhammer Helsreach: a sinf.xml, no encryption.xml, 37 plain-XHTML chapters), which is the same residual-artifact shape as the PDF that motivated the tool. EPUB classification now keys on actual content encryption (encryption.xml with non-font entries) and names the scheme from whichever marker is present; a standalone marker is reported BENIGN as a "residual DRM marker". PDFs are unchanged: a residual handler dictionary there still breaks metadata embedding, so it is still flagged. With this fix the library's real DRM count is 48 (all recoverable PDF ADEPT dictionaries), with the lone FairPlay EPUB correctly cleared. Tests extended to 21 cases (residual markers benign; markers plus encrypted content still DRM).
audit_drm.py: a cross-format DRM scanner (read-only). The metadata and structural audits never look at encryption, so a DRM-locked file can pass epubcheck, report its page count, and even import, yet silently refuse to let its embedded metadata be rewritten. The case that prompted this was a z-library PDF carrying a residual Adobe ADEPT EBX_HANDLER dictionary that qpdf --check and pdfinfo both reported as "not encrypted" while exiftool choked on it, failing the reconcile embed. The new script classifies EPUB, PDF, and Kindle (MOBI/AZW3) files; DJVU has no DRM scheme and is reported N/A. It runs in library mode (formats and paths from metadata.db, opened strictly mode=ro) or directory mode (a recursive scan of loose files before import, for the pre-import battery), with the usual exit codes (0 clean, 1 DRM found or scan error, 2 setup error).
The design priority was not detecting encryption; it was not crying wolf. Two benign things look like DRM to a crude check and are explicitly cleared. Font obfuscation: an EPUB META-INF/encryption.xml that scrambles only the embedded fonts (the IDPF or Adobe #RC algorithms) is not a content lock; because publishers often name obfuscated fonts fonts/00001.dat with no font extension, an entry is cleared when it uses a font-scrambling algorithm OR targets a font resource (extension or a fonts/ path), which an algorithm-only or extension-only check gets wrong. Permission flags: a PDF "encrypted" with the Standard handler and an empty user password opens with no password and is only flagged against printing or copying, so it is classed with qpdf as PERMISSIONS, not a lock. Real DRM is rights.xml (Adobe ADEPT) or sinf.xml (Apple FairPlay) in an EPUB, a content-encrypting encryption.xml, a non-Standard PDF security handler found by a streaming byte scan (so a residual or inactive dictionary is still caught, which is exactly the Irodov case), a password-locked PDF, or a non-zero Mobipocket encryption-type field.
The first whole-library sweep (6,651 files) found 50 DRM-locked files (49 Adobe ADEPT, 1 Apple FairPlay), with 135 benign font-obfuscation/permission cases correctly cleared and one early false positive fixed before release: five EPUBs whose Adobe #RC font obfuscation targets fonts/*.dat were initially misread as encrypted content, which is what drove the algorithm-or-target rule above. Ships with a unittest suite (tests/test_audit_drm.py, 19 cases) building zip, PDF-byte, and PalmDB fixtures for each format and verdict.
reconcile_file_metadata.py --repair-pdf now deletes the .~qpdf-orig backup qpdf leaves behind. qpdf --replace-input, used to rebuild a broken cross-reference table before re-embedding, writes the pre-repair original to <name>.~qpdf-orig beside the file and never removes it. Across many reconcile passes these full-size copies accumulated inside the library tree, which is the worst place for them: Calibre scans that tree, and each one is a complete duplicate PDF. A sweep of one library turned up 21 such files totalling 403 MB. embed_pdf now unlinks the backup as soon as qpdf reports success (return code 0 or 3), before retrying the embed; a missing backup is a no-op, so the change is safe whether or not qpdf wrote one. Regression tests mock the exiftool/qpdf boundary to assert the backup is removed after a successful repair and that an absent backup does not raise. Pre-existing strays from older runs are not cleaned by the tool; remove them once with fd -H '\.~qpdf-orig$' "<library>" -X rm.
audit_epub.py emptytext now flags partial / placeholder exports (new PARTIAL verdict). The whole-book character count missed a failure mode: a DRM-locked or sample export where most chapters are an identical "content unavailable" placeholder while one or two real chapters carry enough text to clear the THIN floor, so the book validates, repairs clean, and reads as full-length to the old total-char check. The canonical case was a BookShout export of Johannes Cabal: The Fear Institute, where 15 of 17 chapters were the same 138-char "something went wrong loading... bookshout.com" stub; even after structural repair to zero epubcheck fatals it stayed a 2-chapter sample. The analyzer now flags PARTIAL when a known DRM-placeholder signature appears anywhere in the spine, or when the same short stub (12 to 600 chars) repeats across at least 3 spine documents and at least 30% of the spine. PARTIAL is a real defect (counts as FOUND, exit 1, needs re-sourcing), distinct from the advisory THIN. The false-positive guard is the per-document distribution: a well-made book full of small but DISTINCT section dividers does not trip it (only repeated-identical stubs do), so the three full novels in the batch that surfaced this stayed OK. Library and directory modes both report it.
audit_epub.py now percent-decodes spine hrefs, fixing false EMPTY verdicts. OPF manifest hrefs are IRIs, so a content document whose archive filename contains a reserved character (commonly !, written %21; Sigil and calibre emit these routinely) was matched against the raw zip namelist undecoded, failed to resolve, and dropped out of the spine. A text-full book whose every chapter file had such a name resolved to zero readable spine documents and was reported EMPTY: the exact false positive hit on Martha Wells's The Serpent Sea (every split_NNN.html was named CR!RT...). The resolver now decodes the percent-encoding (UTF-8, with multi-byte runs decoded together) and strips any #fragment before matching the namelist. Stdlib-only via a small re-based decoder (_pct_decode); no urllib dependency added. The fix lands in the shared spine resolver, so all three analyzers (content, pagenumbers, emptytext) benefit. Regression tests cover the decoder (reserved char, multi-byte UTF-8, invalid escape) and an end-to-end encoded-spine EPUB.
The three EPUB-content audits are merged into one scripts/audit_epub.py. audit_epub_content.py, audit_epub_pagenumbers.py, and audit_epub_emptytext.py shared the same spine resolution, library/directory dual-mode, read-only contract, and exit codes, and differed only in the per-book verdict; they are now three analyzers behind one tool, selected by subcommand: audit_epub.py content|pagenumbers|emptytext|all [directory]. The detection logic of each is unchanged (same thresholds, same results). Two wins beyond removing the duplicated scaffolding: all opens each EPUB once and runs all three analyzers in a single decompression pass (the expensive part is decompression, so this is much faster than three separate full-library runs), and there is now one spine resolver to maintain instead of three slightly-diverging copies. The old script names are removed; update any caller to audit_epub.py <mode>. The --min-chars / --thin-chars knobs (emptytext) carry over.
scripts/audit_epub_emptytext.py: find empty / no-body-text EPUBs. Catches the failure mode every other audit misses: a content-less stub that still validates. The canonical case is the "Bookmate" export, where the archive holds only cover and promo images plus a tiny HTML placeholder, the OPF spine points only at that placeholder, and the book itself is absent. Such a file passes epubcheck, "repairs" clean in a structural repairer (its one referenced document is well-formed), and shows no foreign text to audit_epub_content.py because there is no text at all; the metadata looks perfect. The detector resolves the spine, drops <script>/<style>, strips tags, decodes entities, and counts the rendered characters: EMPTY (<=2000, --min-chars) is a real defect to re-source, THIN (<20000, --thin-chars) is advisory because a genuine short story or a publisher sample can also land there. A Bookmate origin is not itself a defect; most Bookmate exports carry their full text, so the flag is on empty content, not provenance. Library mode (DB-driven, mode=ro) and directory mode (vet downloads before import), mirroring the other two EPUB audits. First full-library run flagged four real defects (a Bujold stub, two image-only scans with no text layer, and a Draft2Digital sample of a full novel hiding in the THIN tier) against three genuinely short stories left alone.
spot_check.py no longer flags OCaml and NCurses as case garble. The intercaps allowlist (_CASE_OK) now includes OCaml and NCurses alongside SQLite, QBasic, and the rest, so legitimate library titles stop tripping the advisory case-garble heuristic.
scripts/audit_epub_pagenumbers.py: find print page numbers baked into EPUB body text. Bad PDF/OCR-to-EPUB conversions capture the print page number (and often the running header) as a literal paragraph in the flow instead of real EPUB pagination, so it reflows into the middle of a sentence ("where the hay cart 16 was taking him"). The detector reads each book's blocks in spine order and flags a number only when it genuinely interrupts prose: a lowercase continuation after it, a word split across it (the previous block ends in a hyphen), or it abuts a repeated running header/footer. Section and chapter numbers (which open the next block with a capital) and endnote/footnote numbers and chronology years are left alone. Library mode (DB-driven, mode=ro) and directory mode (vet downloads before import), mirroring audit_epub_content.py. Validated by hand against the full reference library: 21 flagged, every one a true positive; the false-positive tail (an experimental footnote-poem, a scraped web-serial's vote counts, placeholder section labels) all fell under the hit-count or book-span floors. It also surfaces piracy watermarks and bad OCR scans that ride along with the page-number cruft.
scripts/spot_check.py: randomized metadata + file-integrity audit. Samples N random books (reproducible with --seed) and checks what pattern-based sweeps miss: title corruption and mojibake, junk author entries, missing or stub descriptions, EPUB archive integrity (CRC, container/OPF sanity, spine completeness, text volume), PDF header/page count, and DJVU page count. Emits a machine-readable flag report plus a review bundle (title/author/tag/series/blurb per sampled book) for a human or LLM judgment pass. Read-only against metadata.db; validator-owned checks are not duplicated. First 600-book run against the live library caught a wrong-book description, a truncated description, three mojibake descriptions, and an EPUB with ten dangling spine references.
reconcile_file_metadata.py PDF embeds no longer fail silently. Two related defects: exiftool refuses to rewrite XMP packets containing duplicate properties (seen in the wild: a doubled prism:doi) unless -m is passed, and it exits 0 with "files unchanged", which the tool read as success; The embed now passes -m. (An interim attempt to also write a bare -Publisher tag was reverted: on PDFs exiftool maps it to the same XMP dc:publisher bag, so double assignment appended duplicate entries and broke the round-trip.)
write_catalog no longer corrupts the shared book cache. It sorted the list returned by get_all_books() in place, silently reordering the session cache for every later consumer. It now sorts a copy; a regression test pins the behavior (tests/test_modes.py).
compress_pdf.py cannot clobber a rollback original. If a .pre-compress rollback file already exists, an in-place run now aborts instead of overwriting the only copy of the true original with an already-compressed file.
audit_epub_content.py finds the library from either home. The library root resolves to wherever metadata.db sits: next to the script (the copy living inside the library) or the current working directory (running the repo copy from a library), in that order.
reconcile --id rejects malformed id lists cleanly via a dedicated parser instead of an unhandled ValueError.
Exception chaining (raise ... from) throughout search.py and helpers.py; unused loop variable removed in wing-overlap analytics; re.Scanner access satisfied for type checkers; new test coverage for the backup guard, library-root resolution, and cache isolation (suite: 87 tests).
Python 3.14+. The supported floor moved from 3.9 to 3.14 to match the development environment. The code does not depend on bleeding-edge syntax, but only 3.14+ is tested and supported.
Comprehensive search engine. The search expression parser was rebuilt as a dedicated, stdlib-only engine (src/cquarry/search.py) that ports Calibre's grammar and matching semantics. It now resolves field locations beyond tags and authors: series, publisher, rating, formats, languages, pubdate/date/last_modified, identifiers/isbn, comments, cover, id, uuid, and #custom columns, in addition to tags, authors, all, and vl:. It supports contains/=exact/~regex/^accent match kinds, numeric and date relational operators (rating:>=4, pubdate:>2015, date:30daysago), and field:true/field:false presence tests. Boolean grouping, implicit AND, quotes, and escapes follow Calibre's grammar, evaluated with its candidate-set semantics. The previous build only handled tags:, author(s):, and vl:; other prefixes silently matched nothing.
--search prints to stdout. With no --output, results stream to the terminal instead of forcing a file. --format json|csv|ai emits the matching books in that structured shape; otherwise a plain-text listing is produced. An empty query (--search '') returns the whole library, matching Calibre.
Deeper cover audit. Cover dimension reading no longer stops at the first 1 KB, so a JPEG whose SOF marker sits behind a large EXIF/ICC block is measured correctly; PNG covers are now read too.
Interactive TUI analytics no longer crashes. Selecting any item under the ANALYTICS menu (Author statistics, Reading pace, Tag tree, Wing overlap) raised a NameError because those functions were never imported into tui.py. They now work.
Half-star ratings are visibly distinct. A 2.5 rating rendered identically to 2.0 (both ★★☆☆☆); it now shows ★★½☆☆ using the universally available ½ glyph.
Series "complete" is computed correctly. Completeness now means "no missing integer volumes" rather than "book count equals the top index", so a series with novellas (0.5) or duplicate editions is no longer wrongly marked incomplete.
Portability of the series rollup. get_all_series was rewritten to aggregate in Python instead of using GROUP_CONCAT(... ORDER BY ...), which required SQLite 3.44+.
Normalized custom columns now load. Single-valued text and enumeration custom columns (e.g. a "Status" / reading-state column) are stored by Calibre in a value table plus a books_custom_column_N_link table, exactly like multi-valued columns. The loader keyed off is_multiple and tried to read a book column straight from the value table, which errored and silently returned nothing. It now detects the link table, so --show-custom and #column searches work for every custom-column type.
Honest parity claims. The README and spec no longer claim "100% parity"; they document exactly which locations and operators are supported and the deliberate, dependency-bound deviations (stdlib re instead of the regex module, unicodedata folding instead of ICU, no GPM templates or saved-search references, and cquarry's anchored hierarchical tags: match).
Companion scripts. scripts/compress_pdf.py (Ghostscript-based PDF shrinking with verify-or-rollback; writes files and metadata.db) and scripts/audit_epub_content.py (read-only EPUB content auditor) are now versioned alongside the toolkit and fully documented, explicitly outside the read-only cquarry package contract.
Portable test suite. tests/test_search.py and tests/test_helpers.py cover the parser grammar (adapted from Calibre's own tests), the matcher against an in-memory provider, a full-stack integration test on a temporary SQLite fixture, and the rating/series/image helpers, all without needing a live Calibre library.
Tag Dump (--tags). A flat, alphabetized list of every tag in the library with its book count, written to stdout. Drop-in replacement for the noisy calibredb list_categories -r tags shell pipeline — pipe it to a file with cquarry --tags > tags.txt. Also reachable as "Tag dump" under LISTS in the interactive TUI. Honors --quiet (suppresses header/footer, leaves the body intact for scripting). Distinct from --analytics tags, which renders the hierarchical tree.
Full Parity Search Engine. Refactored the --search expression parser to achieve 100% parity with Calibre's native syntax.
- Author Searching: Added full support for
author:andauthors:prefix tokens. - Fallback Text Search: Un-prefixed terms (e.g.
cquarry --search "author:Anne Rice") now correctly fall back to searching anywhere across book titles, authors, and tags. This accurately mimics Calibre's implicit booleanANDhandling of unquoted spaces. - Complex Grouping: Verified and documented support for nested boolean logic, parenthetical grouping, and negative lookaheads (e.g.,
NOT(tags:Fic.Romance OR tags:Fic.Contemporary)ortags:"Fic.Fantasy.Grimdark" AND author:"Phil Tucker"). - Test Suite. Added an automated test suite (
tests/test_search.py) mapped against Calibre's actualSearchQueryParserbehavior to guarantee ongoing expression fidelity.
Completed Software Status. Phase 4 has been concluded, and CalibreQuarry is now considered feature-complete and stable. It has undergone rigorous end-to-end testing against real-world Calibre databases.
Bug Fixes:
- Fixed a
NameErrorcrash in--auditmode where the newly introducedcolorhelper was not imported, preventing the final summary from printing when issues were found.
Custom Column Support. Added a --show-custom "Column Name" flag that extracts data from user-defined custom Calibre columns. The values are automatically appended to text catalogs and are natively included in JSON, CSV, and AI exports.
Color CLI Output. Introduced simple, lightweight ANSI color formatting for headers, warnings, and error highlights across CLI modes to improve readability when bypassing the interactive pager.
Extended Audit Checks. The --audit mode has been significantly expanded to include three new checks:
- Duplicate Detection: Identifies books with identical titles and primary authors across the library.
- Cover Quality Audit: Scans the actual JPEG cover files on disk (without external library dependencies) and flags covers with low resolution (below 500px on their longest edge).
- Format Migration Report: Flags books that are only available in deprecated legacy formats (MOBI, LIT, LRF, DJVU, PDB, AZW).
Extended Analytics. Added a new --analytics argument with four detailed reporting modes: author (per-author breakdowns of formats, ratings, and series), pace (books added per month/year trend), tags (hierarchical taxonomy tree visualization), and overlap (virtual library wing overlap analysis). These are also accessible via a new ANALYTICS section in the interactive TUI.
Search Query Export. You can now pass arbitrary Calibre search expressions directly to the CLI via --search "query" to export matching books. The results are written to a plain text file. This feature is also accessible via the interactive TUI under the OUTPUT menu.
AI-Readable Export. Added a new ai format to the --export option. This format outputs the library data as a highly token-efficient, flat text list designed specifically for LLM ingestion and recommendation prompts (e.g., Title by Author [Tags] - Rating/5).
Top Authors and Top Tags in --stats. The statistics output now includes a "Top authors" section (10 most prolific by book count) and a "Top tags" section (15 most-used tags), inserted between the ratings distribution and the tag taxonomy breakdown.
CalibreQuarry has been refactored from a single ~1450-line monolithic script (cquarry.py) into a proper Python package architecture, with TUI improvements modeled after the Lattice project.
Layer-Based Package Design. The codebase now lives in src/cquarry/ and is split by logical functionality: config.py, db.py, helpers.py, cli.py, tui.py, and a modes/ directory for individual feature operations (catalog.py, stats.py, audit.py, display.py, export.py). The monolithic script is gone.
Modern Build System (Hatch). CalibreQuarry now uses pyproject.toml managed by Hatch. Install via pip install . or pipx install . and the cquarry command is available globally. Also runnable via python -m cquarry.
Persistent Database Configuration. Both CLI and TUI now share a unified database resolution chain: explicit --db flag, saved config (~/.config/cquarry/config.json), default search paths, then an interactive prompt if running in a TTY. The path is saved on first successful resolution, eliminating the need to pass --db in future sessions. A "Change database path" option under the SETTINGS section in the TUI main menu allows updating the stored path.
Calibre Lock Handling. When metadata.db is locked by a running Calibre instance, CalibreQuarry now automatically copies the database (including WAL/SHM journal files) to a temporary snapshot and reads from that instead. A notice is printed to stderr, and the temp files are cleaned up on exit. Previously, a locked database would produce an unhandled sqlite3.OperationalError.
Fully Immersive TUI. All operations now run through a _run_with_capture() wrapper that intercepts stdout and stderr via io.StringIO buffers. Output is displayed within the scrollable curses pager rather than dropping the user back to raw terminal output. This matches the immersive TUI pattern established in Lattice v4.1.2.
Styled Curses Pause. The post-operation "Press Enter to continue" prompt now renders inside a styled Unicode box within the curses session (_tui_pause), instead of falling through to a raw input() call. Accepts Enter, q, or Esc to dismiss.
Null Byte Sanitization. The scrollable pager now strips null bytes from captured output before rendering, preventing ValueError: embedded null character crashes on corrupted data.
Curses TUI. Running the script with no arguments now launches a full-screen, arrow-key navigable terminal UI utilizing the curses library (matching the interface of getMusic). Non-TTY environments or systems without curses will gracefully fall back to a styled text-based box menu. The TUI features a custom scrollable pager that intercepts standard output, allowing you to comfortably read and navigate command outputs directly within the interface.
Implicit AND in VL expressions ignored subsequent tags. Calibre's parser evaluates adjacent tags like tags:Fic.Fantasy tags:Magic as an implicit AND. The _parse_and method previously discarded tags after the first one unless the AND keyword was explicitly written out. It now correctly intersects all implicit constraints.
Exact tag matching (=) was case-sensitive. Calibre's tag matches are always case-insensitive. The exact match SQL query (WHERE t.name = ?) was missing COLLATE NOCASE, causing tags:"=Fic.Fantasy" to fail if capitalization varied.
Duplicate author headers in --primary-only mode. Generating a catalog with --primary-only caused highly fragmented author groups. The script relied solely on the SQL ORDER BY b.author_sort (which sorts by the full multi-author string). Books are now presorted in Python natively by their derived primary-only display key.
Non-deterministic GROUP_CONCAT output. The metadata fields built via GROUP_CONCAT(DISTINCT ...) (like authors and tags) returned unpredictably ordered results depending on SQLite's internal row execution. This occasionally resulted in the wrong primary author being selected. The SQL query has been rewritten to use correlated subqueries with explicit ORDER BY clauses for deterministic structure.
Double NOT cascading crashes. Expressions combining consecutive exclusions (e.g., NOT NOT vl:Name) failed because _parse_not routed directly to _parse_atom for the inner operand. This has been updated to recursively call _parse_not to handle complex nested negations gracefully.
Fractional series indices ignored in recent display. show_recent dropped series identifiers completely if the index contained a decimal (e.g., 1.5) due to a missing fallback formatting block.
_read_value crashed on unmatched quotes in VL expressions.
expr.index('"') raised ValueError if a virtual library search expression
had an opening quote with no closing quote. A single malformed VL definition
took down the entire tool. Now consumes the rest of the string as the value.
Series index 0.0 silently dropped. if idx and idx == int(idx) treated
0.0 as falsy — any book at series position zero lost its series display. Hit
in both write_catalog and show_recent. Changed to explicit is not None
checks.
Division by zero in show_stats. Empty library crashed on three separate
lines: format bar chart (count * 40 // total), rating bar chart
(max(rating_counts.values())), and unrated percentage
(unrated * 100 / total). All three guarded.
max_index None dereference in show_series.
s['max_index'] == int(s['max_index']) threw TypeError when max_index was
None. Same fix propagated into detect_series_gaps.
format_stars produced garbage on corrupt ratings. A DB value of 12 (6.0
stars) yielded a negative empty count. Python silently returns "" for
"☆" * -1, so no crash — but the display was meaningless. Rating now clamped
to 0–5.
CSV export blanked zero-valued fields. stars or '' and
series_index or '' used the or pattern, which treats 0.0 as falsy.
Changed to explicit is not None checks.
JSON export had leading whitespace in split fields.
b['authors'].split(',') on GROUP_CONCAT output produced
["Author1", " Author2"]. All split fields now strip.
Tokenizer keyword boundary missed underscores. isalnum() doesn't match
_, so a hypothetical tag starting with or_ or not_ would misparse as a
boolean operator. Boundary check now includes underscore.
_prompt_str displayed [None] in interactive prompts. When called with
default=None, the user saw the literal text [None]. Now shows empty
brackets.
get_all_books() results cached. The 8-JOIN metadata query was called once
per wing in --all-wings mode — 18 times against a 3,800-book library. Now
fires once and returns the cached list.
_get_all_book_ids() cached for NOT operations. The VL parser queried
SELECT id FROM books on every NOT clause. Multiple NOT expressions in a
single VL definition hammered the DB. Cached on first call.
count_books() uses warm caches. If the books or IDs cache is already
populated, returns len() instead of hitting SQLite.
Parser built once in main(). Was constructing build_parser() to parse
args, then building it again on the help-output fallthrough path. Stored the
reference.
--version flag. Uses action="version" so argparse handles it during
parse_args() — works even when no database is present. Version also shown in
the interactive menu banner.
Unused imports removed. defaultdict and Path — imported, never
referenced.
f-string with no placeholders. f" Languages:" → " Languages:".
show_wings caught bare Exception. Narrowed to ValueError, which is
what resolve_vl actually raises.
quiet parameter wired up everywhere. show_recent, show_series, and
show_stats all accepted quiet but ignored it. Now suppresses headers and
decorative output when passed.
main() catches PermissionError. A read-only DB with wrong filesystem
permissions previously produced an unhandled traceback.
_match_tags docstring corrected. Said "containing" for the non-exact case
but the SQL does prefix match, not substring. Added a note that regex patterns
(tags:~regex) are unsupported.
prog="cquarry.py" added to ArgumentParser. Version and help output now
show the script name consistently regardless of invocation path.