The v1.23.1 hotfix shipped with four stale version carriers (src/cquarry/__init__.py, src/cquarry/config.py, spec.md, API.md all still reading 1.23.0), so the published wheel reported cquarry.__version__ == "1.23.0" and every consumer reading the constant saw the wrong release. The guard that should have caught it, tests/test_version_sync.py, was the one file in the suite written as plain pytest-style functions -- and CI runs unittest discover, which silently collected zero tests from it. The file is now a unittest.TestCase (469 -> 471 collected tests), every carrier agrees at 1.23.2, and the discover run that gates each push executes the guard.
Found by the post-blitz verification day's deep pass on the foundation repos.
- Bugfix:
set_series(book_id, None)andremove_entity_everywhere("series", ...)wroteseries_index = NULL, which the real Calibre schema rejects.books.series_indexis declaredREAL NOT NULL DEFAULT 1.0in every live library, so both verbs raisedIntegrityErrorinside the caller's batch and rolled the whole pass back (field find on book 9136 during the 2026-09-16 math/classics phase 3; the test fixtures' nullable column hid it until now). - Clear semantics are reset-to-1.0, matching Calibre's own no-series state. Both verbs
now reset
series_indexto 1.0 when the link goes: the DDL default, what upstream's own series-removal path writes, and what series-less books hold in practice (5438 of the 5442 series-less books in the reference library; zero NULLs anywhere). - Pinned against the real DDL. New
TestSeriesClearAgainstNotNullSchemafixture carries the NOT NULL column; the two older tests asserting the NULL clear were corrected to 1.0.
THE FINAL AUDIT's feature spine (L4 ranks 1-4) plus the approved bare setter, the LOW bug tail, the docs-truth debt, and the storefront polish, in one release.
integrity.find_missing_format_files(db): catalogued format rows whose file is absent on disk; the integrity family's last disk hole. Ridesget_format_path(verify=True)'s check exactly the way the cover checks rideget_cover_path; emptybooks.pathbooks are skipped (nowhere to look, same rule asfind_missing_cover_files).CalibreDB.external_changes_detected():PRAGMA data_versionas the cheap staleness token long-lived holders (Hermitage, Carrel) have had recorded twice; poll it and callrefresh()only on True. The answer stays True untilrefresh()re-primes the baseline, so a poll loop cannot miss a change; on a locked-database snapshot connection it can never fire, which is that boundary's reminder to reopen.- Consistent snapshots. The locked-database lock-escape now copies
through sqlite3's backup API: one consistent page image with the WAL
folded in, replacing three racing
copy2calls (the same torn-snapshot shape Wave 14 flagged in CalibreQuarry's own backup). Python's backup retries a busy source forever, so the copy runs on a leash:CalibreDB.SNAPSHOT_LEASH(10 s) bounds the wait, and a writer still holding the lock past it trips a documented fallback to the old raw file copy rather than hanging the reader.backup_to(dest)gives consumers the same consistency for their own backups; CalibreQuarry'scopy2site can retire onto it at that repo's release. - Annotations. The
annotations:search location bulk-loads its text map in one query (the per-book probe was an N+1), andget_annotations_decoded()projects Calibre's raw rows onto{book, format, kind, annot_id, timestamp, text, notes, title}so a renderer never re-learnsannot_data's shape (wiring it into Hermitage's Codex is that repo's recorded lane). write.set_series_index(book_id, index)(the approved write API): the bare index correction completing the 1.19 passthrough family. Updatesbooks.series_indexin place, the link row untouched (no delete-and-reinsert likeset_series); the book must already belong to a series,Noneraises, and an equal value is an honest no-op.- The LOW bug tail (L2, all eight):
pubdate:Ndaysagowith a gigantic count converts toParseExceptioninstead of a raw OverflowError;set_formatstores integer sizes likeadd_format; the boolean vocabulary is one shared constant pair (BOOL_TRUE_WORDS/BOOL_FALSE_WORDS) so search and write agree on_checked,_blank,_emptyand friends;custom_columns.idis int()-cast before every f-string table name on both sides (the Wave-13 defense's siblings);_custom_column_metaraises the house ValueError on schemas predating editable/display;genre_distribution's node emission is iterative (a hostile deep tag no longer risks RecursionError, output order identical);uuid4registers without deterministic=True (SQLite may reuse a deterministic result within a statement);empty_trash/expire_trashno longer materialize a.caltrashtree in libraries that never trashed;list_books' dead end-recompute andadd_book's always-true batch guard are gone; composite custom columns read as a documented empty instead of a stderr warning. - Docs truth (the release-sync debt from the 2026-09-13 blitz):
API.md's tristate note now allocates correctly (numeric/rating/date
locations take exactly true/false; boolean locations take the tristate
set; re-dated 1.18), the false "identifier keys sweep as text" claim
is gone, the refresh() locked-snapshot boundary clause reached
README/spec/API, the README glance table gained the metadata-quality
trio, spec section 4 lists formats+languages in the bare-term sweep,
and
find_db's config auto-persist is documented. Prose batch: README's write mega-bullet split per verb family, the bindery-cli rename and link, the spec identifier-cleaning garble and dossier seam fixed, the three live em-dashes recast, the roadmap opener rewritten, and the comment contracts (the module-docstring "EVERY mutation" overclaim and the last retired ATTACH claim) corrected. - Storefront: pyproject carries the canonical description (it
discloses the opt-in write path; the library-graduation check-in
closes as acknowledged), Bug Tracker/Changelog/Documentation project
URLs, an absolute logo URL (PyPI rendered a broken image), a README
badge row, pytest
pythonpath = ["src"], and the version-sync guard now covers the patchnotes head and__init__.py(eight of eight carriers guarded). - Housekeeping: .gitignore trimmed to what a stdlib library produces;
REPORT-12-Sept.md retired per the :988 precedent (the declined-lines
payload extracted into the roadmap, archive copy in audit-final); the
repo-local
testing_facility/(28 MB of byte-identical bootstrap duplicates) reclaimed; ci.yml SHA-pinned like publish.yml; the no-force-push/no-delete ruleset applied on main; wiki and Projects disabled (repo settings, outside the file tree). - Both import skills swept per the skill-sync rule: phase-3-import gains the set_series_index line; the spine touches nothing the phase-1 import loop teaches.
- Suite: 439 -> 464 tests.
The CalibreQuarry blitz lane's two cquarry touches under cross-repo grant #117 (decision 2026-09-15), shipped as one additive release.
-
CalibreDB.precedent_tags(authors, limit=12): the tag-by-precedent suggestion read (distinct tags across the named authors' books, NOCASE author match, capped and alphabetized for stability). Promoted from CalibreQuarry's run.py phase-3 prompt, whose four-table JOIN was the only unrecorded raw-SQL read in the frontend tier (THE FINAL AUDIT L2.6). Two deliberate deltas from the promoted form: results are ORDER BY name (the original relied on SQLite's arbitrary DISTINCT order) and the limit is a parameter (was a hardcoded 12). Documented in API.md; five fixtures in test_db.py. -
The publish path hardened (the audit's publish-workflow box): pypa/gh-action-pypi-publish is SHA-pinned (release/v1 was a moving branch holding id-token: write), the workflow carries a top-level contents: read permissions block with the publish job keeping only id-token: write and a scoped contents: write release job, the concurrency group refuses cancellation mid-publish, the test job runs the CI ruff gates first, the build gets a strict twine check and a wheel smoke-install, and a create-release job mints the GitHub Release from the tag's verbatim message. actions-only dependabot keeps the pins current. The no-force-push ruleset on main remains recorded for cquarry's own lane.
Erratum (2026-09-15, post-release): this entry as tagged also said the pypi environment's deployment policies now admit only the v*.. pattern. That part was reverted the same day: the REST API only creates branch-type policies, and a tag deployment is rejected outright whenever custom branch policies exist, so the policy broke this very publish (the v1.22.0 tag shipped only after the policy was reverted and the failed job re-ran). Tag policies are UI-only today; the workflow's v*.. tag trigger remains the effective gate, and a UI-applied tag policy is a recorded reopen item.
-
Suite: 434 → 438 tests.
The six-lens audit's riders decision, executed: 1.20.1 shipped green and pushed, so the two recorded candidates that have a real consumer rode as this minor release. Additive API only; no floor bumps owed anywhere.
list_books(sort="ids", ids=...): the caller's-order mode. The listing's contract said its order comes fromsort, never from the id order; the special keyids(requiresids, stands alone) now replaces the sort with the caller's id sequence verbatim. A duplicated id keeps its first slot, ids absent from the library are skipped,descendingreverses the sequence, andoffset/limitslice after the ordering. This is the mode Carrel-calibre-web'spreserve_orderre-sort shim (added page-by-page afterlist_bookscame back title-sorted) has been waiting for; the fork can retire the shim at its own release.- The metadata-quality trio in
integrity.py. The routed bindery OPF-085 item (51 warnings counted in the audit), promoted to the shared predicate family so every consumer answers identically:find_invalid_uuids(db)(books whoseuuidis empty or does not parse as a UUID; the pre-uuid-column degrade spelling""reports honestly),find_sentinel_pubdates(db)(the0101-01-01undefined date sentinel and its0100-01-01ancestor, the same pair the search engine already treats as dateless), andfind_bad_language_codes(db)(a linked language code that is not exactly three lowercase ASCII letters: bare names likeEnglish, two-letter codes, empty strings; shape check only, so a valid but rare code never false-positives). CalibreQuarry renders the audit rows in its own lane.
Both import skills swept; neither rider touches the import loops, so no skill edits. Consumer floors stay where they are (adoption notes in the roadmap's cascade block).
A bug-fix release from the 2026-09-12 six-lens audit (Wave 13): five write-path defects plus a LOW hardening batch, no API changes. Two of the fixes restore contract claims the docs already made; the release closes the gap instead of rewording the promise.
save_original_formatattaches the FTS sidecar before its batch (HIGH). The save opened its transaction without_ensure_fts_attached(), so the nestedadd_format/set_formatattached inside it; on SQLite builds that forbid in-transaction ATTACH (pre-3.21.0) the caught failure cached_fts_state=Falsefor the connection lifetime, silently killingdirtied_formatsqueueing andneeds_scanfor every later format verb. The one-line fix mirrorsrestore_original_format. On modern SQLite the poison cannot fire (in-transaction ATTACH has been legal since 3.21.0, 2017), so this is attach-first discipline plus armor for old builds, and the docstring no longer claims modern SQLite forbids it.- A failed batch flush can no longer strand committed books'
directories.
batch()'s finally ran the three post-commit flushes BEFORE resetting_batch_dirs/_batch_poisoned; a flush OSError left the caller with a COMMITTED transaction plus registered directories that a LATER failed exit would rmtree (rows pointing at deleted directories). The state now resets in a nested finally before the error propagates, and_remove_book_diris idempotent (a retried flush skips directories that are already gone instead of raising on the missing source). A failed flush's committed removals stay queued and a later flush completes them. remove_bookclears the FTS queue for the book's formats. It cleanedmetadata_dirtiedandannotations_dirtiedbut never the sidecar'sdirtied_formats, while its docstring and API.md claimed the queues are cleaned; Calibre would have re-extracted vanished files. The sidecar attaches before the transaction, the book's formats are captured before the cascade, and each queue entry is cleared; the docstring and API.md now name the FTS queue explicitly.add_bookseeds the FTS/pages queue for its formats. Seeded formats were inserted directly, bypassingadd_format, so new books never entereddirtied_formatsor gotneeds_scan; upstream's own add path queues everydataINSERT.add_booknow attaches before its batch and queues each seeded format; the queue writes join the batch and roll back with it.- A failed commit no longer diverges rows from files. Bare
(non-batched)
update_title/set_authorsapplied the filesystem re-lay BEFORE the commit, so a commit failure left rolled-back rows under moved/renamed files (exactly what_relayout_book_path's docstring promised could not happen). The re-lay now queues in every path and lands after the rows commit: the outermost batch commit inside a batch, the setter's own commit otherwise, and a failed commit drops the queued op instead of leaking it into a later batch. - LOW hardening.
expire_trashcounts only actual removals (a failed rmtree used to be counted as removed);refresh()'s docstring no longer overpromises on the locked-DB snapshot path (the snapshot is not retaken; reopen the CalibreDB for current data);load_configcatches onlyOSError/JSONDecodeErrorand notes a broken config on stderr instead of failing silently into "no config";get_tag_browser_countsnarrows its label-map suppress tosqlite3.Errorand double-quotes the view identifiers it reads (defense-in-depth: the names come from sqlite_master).
Both import skills swept: they teach "since 1.18 format writes queue FTS", which these fixes restore rather than change; no skill edits needed. Patch release: consumer floors stay where they are.
Brandon un-deferred C.5 and C.7 ("get that done as 1.20.0, along with anything upstream"). The upstream delta since the 09-10/11 research turned out to be two trivial commits (a notes-import birthtime fix -- notes stay declined -- and a typing nit); no schema, search, or write surface changed, and the real-library schema version still matches the add_book census.
- Trash lifecycle:
list_trash()/empty_trash()/expire_trash(older_than=). The management half ofremove_book(delete_files="trash"), over upstream's.caltrash/b/.caltrash/flayout.empty_trashis upstream'sclear_trash_dir(remove whole, recreate empty);expire_trashisexpire_old_trash's mtime rule with upstream's 14-day default (timedeltaaccepted;<= 0expires everything);list_trashis the added reviewable inventory ({category, book_id, mtime, files}) so the destructive half can be looked at before it runs. Pure filesystem verbs: trashed rows are already gone, so nothing touches the database. create_custom_column(label, name, datatype, *, is_multiple, editable, display)/delete_custom_column(label). The library bootstrap. Creation mirrors upstream's DDL statement-for-statement: thecustom_columnsrow, value/link storage per the normalized vs direct split, the fkc guard triggers, the seriesextraindex column, and thetag_browserviews (including thefiltered_one -- views are lazy, so its Calibre-only function is safe to reference). Calibre'supdate_all_last_mod_dates_on_startpref is set like upstream so the next start refreshes every book. Two deliberate deviations, documented: column numbers allocate past the highest row id AND storage-table number (a barelastrowidcan collide with tables still awaiting Calibre's purge of a flag-deleted column), and the link-table update guard fires onUPDATE OF valuewhere upstream'sOF authornames a column the table does not have (its guard was dead code). Deletion drops NOTHING, exactly like upstream: it setsmark_for_delete=1and the physical purge is Calibre's own next-startup job; the flagged column stays listed and functional until then. Created columns work end to end immediately --set_custom_columnwrites through them, and a fresh reader loads and searches them, including the series#label_indexlocation.
Riders: none this wave -- no consumer consumes these verbs yet (CalibreQuarry's management surfaces and the acquisition importer come later), so the ecosystem floors stay at >=1.19.0 per the four-program rule's no-purpose-no-bump clause.
Suite: 412 passed (was 400).
Brandon approved four of the six researched write-API candidates from REPORT-12-Sept.md (Phase 13 section C); the other two stay deferred with their consumers. Every verb ships with tests against the trigger-hazard schema, keeps its file work behind the commit boundary, and rides the 1.18 machinery (path re-laying, FTS dirtying, the batch queues).
rename_entity(kind, old, new)/remove_entity_everywhere(kind, name). The fix-the-misspelled-name verb for authors, series, publishers, and tags. A rename whose target already exists merges: colliding links drop against the survivor's UNIQUE(book, fk) and the old row is deleted. Author renames recomputebooks.author_sortand re-lay every affected book's on-disk path; series merges renumber incoming books to max+1 over the survivor's other books; series removal nullsseries_indexlikeset_series(None). Both return the affected book count and queue every affected book for OPF resync.set_cover(book_id, data)/remove_cover(book_id). The audit-to-fix loop closes:find_low_res_coversnames the bad cover,set_coverreplaces it. JPEG/PNG sniff-or-raise (theadd_bookrule; no stdlib transcoding), written tocover.jpg/cover.pngwith any stale cover under the other extension removed;remove_coversweeps both and clearshas_cover.set_author_sort/set_title_sort/set_timestamp. Verbatim passthrough corrections for mangled sorts (set_timestampnormalizes likeset_pubdate). A laterset_authors/update_titlerecomputes over an override by design; that interplay is documented on both sides.save_original_format(book_id, fmt)/restore_original_format(book_id, "ORIGINAL_<FMT>"). The undo-able repair primitive bindery's lane asked for. The copy is a realORIGINAL_<FMT>data row plus a<stem>.original_<ext>file -- upstream's own convention, so Calibre sees it natively. Restore swaps the bytes back, keeps the target's filename stem, removes the original row/file/queue entry, and always queues the restored format for FTS re-extraction (the bytes changed even when the row was already correct).
Deferred: the trash lifecycle (nothing uses
remove_book(delete_files="trash") in anger yet) and
create_custom_column/delete_custom_column (the acquisition importer
is a future project; schema DDL waits for a real consumer).
Suite: 400 passed (was 381). Consumers: CalibreQuarry and bindery-cli bump their floors to >=1.19.0 (the consuming riders); Hermitage consumes none of the four and stays at the 1.18.0 pin.
Everything below ships against REPORT-12-Sept.md, three read-only research passes that compared this repo line-for-line against the upstream Calibre clone (schema + upgrade map, search stack, write paths). The promotion candidates the report surfaced stay shut pending approval; every ungated box in Phase 13 (roadmap sections A-D) is closed here, each with its pinning or regression test.
- The
full-text-search.dbsidecar is now read (the last unread Calibre data file).CalibreDB.get_book_text(book_id, fmt)returns one format's fullbooks_textrow includingsearchable_text;get_text_extractions(book_id=None)returns the bulk status rows WITHOUT the megabyte texts (err_msgaudit, format-hash change detection);search_book_text(query, *, fmt=None, ids=None)is a case- and accent-folded Python-side content search returning{book_id: {FORMAT, ...}}. Same lock-escape snapshot handling as metadata.db; a missing sidecar degrades every read to empty;refresh()drops the sidecar connection with the rest. No FTS5 machinery: the index tables tokenize through Calibre's custom tokenizer and are unqueryable outside Calibre, so the plain table is the read surface. integrity.find_failed_text_extraction(db)returns{book_id: {FORMAT: err_msg}}for scans, DRM, and corrupt files; the third sanctioned non-cached read in the module after the two cover-file checks.CalibreDB.get_page_metadata(book_id=None)exposes thebooks_pages_linkauxiliary columns (algorithm,format,format_size,timestamp,needs_scanas bool): provenance for displayed page counts, withneeds_scan=Truemeaning Calibre has queued a recount.#label_indexis real for custom series columns. The engine registered the location but it could never resolve (#myseries_index:>3silently matched nothing);load_custom_columnnow reads the link table'sextrafloat and the engine serves it. An exact label literally ending in_indexkeeps the token; ancient schemas degrade. Also riding the same SQL:CalibreDB.custom_column_links(col_name)surfaces the normalized value tables'linkURL column (no Calibre UI populates it).- Doc line (spec 3.2, CLAUDE.md, API.md): a stored annotation's notes follow its highlighted text joined by LF + unit separator + LF; the research line dropped the leading LF and is corrected in the roadmap.
- Fixed:
identifiers:KEY:TRUE/FALSEinverted on uppercase (the one real bug the research found). The presence gate folded the value but the selection compared the raw one, soidentifiers:isbn:TRUEmatched the COMPLEMENT of the right answer. Upstream lowercases once and uses that value for both; so does cquarry now. - Bare
true/falseare the sweep-wide presence test (upstream parity), not substring matches:truematches books where any swept field holds a non-blank value, and the identifiers store and cover flag count as present. Identifier keys never text-sweep (the old spec claim was wrong; the code now matches upstream instead of the claim). Bare numeric probes lostid(upstream excludes it) and cover-as-0/1 (cover joins via presence), leaving series_index, rating, pages, size. - Super-quotes
"""..."""ported (upstream's documented escape hatch for quote/paren/regex-heavy queries):title:"""a "b" (c)"""parses and matches; it used to mis-tokenize. template:raises a clear ParseException naming the missing template engine (model: upstream TemplatesNotAllowed) instead of silently matching nothing. The optional template implementation stays out of scope.- Strictness fixes (the honesty pass, each pinned): date locations
take exactly
true/falseas presence words and no match-kind prefixes --pubdate:blankorpubdate:~2020now raise, like upstream's date-conversion error, instead of matching dateless books. Numeric locations take exactlytrue/false--rating:checked/rating:blankraise like upstream's non-numeric error (the tristate vocabulary is bool-only, where it still works). Text fields' presence words narrowed to exacttrue/false(title:yesis substring text again). New dated spec-deviation entries: tristate bool fidelity (9), lexer strictness (10), the benign extensions consolidated (11), the entity case-change policy (12).
- Path re-laying on
update_title/set_authors. Curation renames no longer leave theAuthor/Title (id)directory andTitle - Author.extformat files under the old names: the layout moves with the rows (directory rename, per-format file renames, emptied-parent removal, stale-target replacement, case-only spelling fix, db-only correction for path-less legacy rows). The fs half lands only after the commit, deferred to the outermostbatch()commit via_pending_relayoutsand dropped on rollback, so a failed pass never leaves rows pointing at renamed directories. - Fixed in the same machinery: a failed batch never cleared
_pending_removals, so a later successful batch could flush stale removals against rows the rollback had resurrected. The rollback path now clears both deferred queues; regression test included. - FTS + pages dirtying alongside format writes.
add_formatandset_formatqueue the (book, format) pair into the sidecar'sdirtied_formatsand setbooks_pages_link.needs_scan, so a repaired file no longer leaves Calibre's content index and page counts stale forever (Calibre never re-reads a file on its own; it processes its queues).remove_formatclears the queue entry so Calibre never re-extracts a vanished file. The sidecar is attached before the transaction opens, the queue writes roll back with the batch, and a missing sidecar degrades to a no-op. The stalebooks_textrow of a removed format stays Calibre's own to clean (the sidecar's delete triggers need Calibre's custom tokenizer; deleting here would corrupt the index). clean_identifierparity.set_identifier(and the mirroredclear_identifiernormalization) cleans both halves like upstream: the type is stripped, lowercased, and stripped of:/,; the value is stripped with,mapped to|-- Calibre's comma-free stored shape.
Suite: 381 passed (was 327). Consumers bring up per roadmap Phase 13 section D (the four-program wave).
set_format(book_id, fmt, name, size). The sanctioned remove+add of onedatarow in a single transaction -- the row half of swapping a repaired file into place. Honest no-op (False) when the identical row already exists;Truewhen written. Retires bindery's hand-rolledremove_format+add_formatcomposition.remove_book(book_id, delete_files=None). The file story is decided:"trash"moves the book's directory into the library-local.caltrash/b/<id>/(upstream's own trash layout, exact stdlib parity viashutil.move),"permanent"deletes it outright, and the defaultNonekeeps rows-only semantics for callers that sweep themselves. File removal lands only after the rows COMMIT; inside abatch()it defers to the outermost commit, so a rollback never leaves resurrected rows fileless.CalibreDB.find_candidate_duplicates(title, authors, isbn=None). The library half of duplicate screening: the ISBN rule first (separator- and case-insensitive against the book'sisbnidentifier), then normalized title (folded, subtitle and leading article scrubbed) plus folded first author, both exact. Returns{"id", "matched_by": "isbn" | "title_author"}entries sorted by id; the file half stays in the frontend where the embedded-metadata readers live.CalibreDB.export_rows(*, ids=None, include_custom=True). Flat one-dict-per-book rows: the hydrated row with every non-composite custom column flattened in as a#labelkey (Nonewhen the book has no value, so the key set is uniform).ratingstays the raw internal 0-10 andpubdatethe raw TEXT; conversion is the renderer's job. Researched against CalibreQuarry's inline exporter SQL, which this retires.integrity.find_identifierless(db). Books with an empty identifiers store; promoted from Hermitage's inline Insights predicate and joining thefind_*family.
- Q2 (notes and dist):
database_report.mdandresearch.mdare deleted; the three facts that existed nowhere else moved into CLAUDE.md (asymmetric link-table uniqueness, the FTS5 sibling table names, the identifiers EAV pseudo-type note). The gitignoreddist/1.9.0 build artifacts are gone. - 3.2 satisfied: batch-scoped filesystem compensation needs no further
exposure -- 1.15.0 made it automatic inside any consumer's
batch(). - 3.4 deferred: the banned-labels policy stays unbuilt -- no upstream anchor, no consumer has specified labels or semantics; revisit when CalibreQuarry's lane brings a concrete policy.
- The
on_duplicate=policy parameter onadd_bookwas not built: callers consultfind_candidate_duplicatesfirst, the same confirmation-before-act patternremove_bookuses.
- Test inflation deflated: 415 collected items down to 326 distinct.
The phase-6 expansion suite re-ran in full inside every subclass
(
TestBatchContext,TestSetPubdate,TestSetWriteConveniences,TestAddBookeach inherited all ofTestWriteSideExpansion's tests). The fixture plumbing is now_WriteSideFixtureand the expansion tests live in_WriteSideTests; plain mixins carry noTestCasebase, so each suite runs exactly once, and the one variant whose schema genuinely differs (TestAddBook's trigger schema) keeps its own re-run variant. Also: the triple-pasted simple DDL lifted into_make_simple_db, thetransaction()success twin replaced with an identity assert (the failure twin still exercises the alias end to end), the near- tautological search-integration count folded into the first real assertion, the duplicated get-entities unknown-kind assertion deduped, and the_now()timestamp shape pinned by regex. - Docs repaired. The
tag_rollupexample in the roadmap now teaches the shipped subtree-totals rule (its own ship note had declared the old mixed rule dead); the annotations no-FTS deviation is spec section 5 item 9, so the canonical list matches API.md's; "8-JOIN" corrected to the real 6 joins plus Python hydration in spec.md and CLAUDE.md; the 1.13.0 patchnotes baseline repaired to 258 so same-day entries agree; both README IndentationError snippets fixed (theget_dirtied_booksone rode 1.15.0, the quickstart's mid-block dedent this release). - Version-sync guard.
tests/test_version_sync.pyasserts all six version carriers agree (VERSION, pyproject.toml,__init__.py,config.py, spec.md, API.md); nothing guarded them before, and carriers had drifted once (1.12.0 shipped withconfig.pyand API.md stale). - Recorded, not decided (Brandon's call): the committed working notes
(
database_report.md,research.md) and the gitignoreddist/1.9.0 build artifacts stay until he decides; the state and options are in the roadmap box.
- Recursion never escapes as
RecursionError. Grammar-valid adversarial queries (400-deep nesting, 5000-term chains, deepvl:/search:chains) now surface asParseExceptionfromsearch()and from the virtual-library and saved-search matchers, with upstream's own two guard sites as the model. - An empty query after any location matches nothing.
title:used to match presence (all books) while upstream, and cquarry's owntags:, returned nothing; every location now agrees with upstream, includingvl:,search:, and date fields. - Invalid boolean keywords raise.
cover:mayberaisesParseExceptionlike upstream'sBooleanSearchinstead of silently matching nothing. - Two-letter language codes canonicalize.
languages:jamatches books stored withjpn(upstream'scanonicalize_langstep, now mirrored with a full ISO 639-1 to 639-2 map for cquarry's languages). - Component exact matching strips parts.
authors:=..CherryhmatchesC.J. Cherryh's last component (upstream matches stripped components; cquarry compared raw split parts), the leading-dot literal comparison mirrors upstream's one-dot form, and the module docstring's flagship example is replaced with one that is actually true. - The
allsweep is upstream's. Bare terms now coverformatsandlanguages(canonicalized), sweep identifier KEYS as text, and probe numeric fields (id,series_index,rating,pages,size,coveras 0/1) by exact equality; date fields take no part, all exactly like upstream's sweep. - Smaller papercuts, fixed or dated. Fixed:
search:=Namestrips the exact-match prefix instead of reporting a known search as unknown; a bare#Nwith no relop text-searches the literal string instead of raising; identifier presence accepts exactlytrue/falsesoyes/checkedno longer leak into value matching; whitespace-only stored values count as absent. Documented as dated spec section 5 items instead: date parsing/comparison leniency (fromisoformatvsdateutil, calendar dates vs local instants) and composite custom columns matching nothing (the template engine stays out of scope per section 7).
- Native lists end to end for multi-valued custom columns.
load_custom_column()andfield()returnlist[str]for multi-valued columns, andset_custom_columntreats a bare string as ONE value instead of comma-splitting it, so a storedDoe, Johnsurvives as a single value; the old join-on-load, re-split-on-use round-trip turned it into phantom values and made count and exact searches lie. Consumers that.split(",")custom-column output must stop. CalibreDB.refresh(). The cache-invalidation boundary for long-lived holders: one call clears every cache (rows, ids, search view and engine, preferences, custom columns, path index) so a connection held open across an external write stops contradicting itself.- Corrupt preferences and ancient schemas degrade instead of
crashing. A corrupt or non-dict
virtual_libraries/saved_searchespayload degrades to{}with the same treatment the generic preferences accessor already used;get_identifiers(), the search view's identifier sweep, andfield()'s comments read degrade on schemas missing those tables, likeget_all_books()always did. resolve_vl()andresolve_saved_search()canonicalize before the lookup."My VL"and" My VL "resolve to a known library instead of raising; the guard and the lookup can never disagree again.- Papercuts. The
0101-01-01sentinel sorts as dateless inlist_books(undated books no longer come first on descending pubdate); duplicatebooks_ratings_linkrows no longer fan a book out into several rows;title_sortstrips a wrapping quote pair before the article move (upstream'squote_pairs, also making the trigger UDF more faithful);strip_htmldrops unterminatedscript/stylebodies; the format-path docstrings state the POSIX truth aboutnormcase.
set_rating(book_id, 0)clears likeNone(Calibre maps 0 to unrated; a 0-rating link used to land as a spurious Tag Browser entry).set_languageswritesitem_orderwhen the schema carries it instead of leaving every row at 0._write_pattern_a's no-op detection orders old rows deterministically instead of relying on scan order.add_formatrejects negative sizes.- Dispositioned as documented: new authors default their sort to the display name (upstream's surname-flip heuristic runs on GUI edits, not row creation); dated note in the spec.
- Both import skills synced with the behavior changes: phase-1 gained the byte-identity floor note in its duplicate screen; phase-3 gained the Ctrl-C-safe batch note, the one-value-per-bare-string rule, and the add_book double-import clause.
- The gated items are recorded, not decided:
remove_book's filesystem story carries both options and a recommendation in the roadmap box; the sweep's papercut list is fully dispositioned (implemented or documented, each with its commit); the promotion candidates remain open for Brandon.
- No more torn writes on Ctrl-C.
WritableCalibreDB.__exit__used to commit unconditionally, and every setter's rollback caught onlyException, so aKeyboardInterruptbetween a setter's SQL statements escaped the rollback and was committed on context exit: a link row without itsmetadata_dirtiedrow, a publisher change queued but unmarked. Rollback paths now catchBaseException, the context-manager exit rolls back whenever an exception is in flight, and its commit failures propagate instead of silently discarding the transaction (matchingbatch()'s exit). The class docstring's "nothing is written unless a method returns normally" is now honored rather than aspirational. - A failed batch no longer strands orphan book directories.
add_book's directory removal was per-call only: when a later book in a sharedbatch()failed, the SQL rolled back but the earlier books' directories stayed behind, looking like real books no row pointed at. Directories created inside the outermost batch are now tracked and a failed exit removes them all; adds made outside any batch never register, so a later failed batch cannot touch committed books. - Nested-batch failure sticks. An inner
batch()exception caught by the outer block used to be swallowed into a commit (the local success flag only saw the outer's clean yield). A failure now poisons the pass: the outermost exit rolls back, and the flag clears on exit so the next batch on the same handle commits normally. - add_book refuses byte-identical re-imports. The "never imported twice" invariant is enforced with a data-table check: a format seed whose bytes match an already-catalogued file (same format and size pre-filter, then content compare) raises before anything is written, dry runs included, under any title or author. Metadata-based duplicate screening (ISBN/title/author fuzz-matching) stays a frontend concern and remains the gated promotion candidate.
- Unknown datatypes raise instead of being stringified. The writers
now dispatch against Calibre's own datatype set:
text,enumeration,series, andratingwrite through link-table storage,int,float,bool,datetime, andcommentswrite direct; anything unknown, or a known datatype on the wrong layout, raisesValueError. - Rating-typed custom columns take 0-5 stars and store x2 on the
same internal 0-10 scale as
set_rating(verified against upstream: its writer takes internal values and the x2 lives only in the GUI; cquarry's API takes stars so one library never carries two conventions). 0 stars means unrated and clears, matching upstream's purge of 0-rating rows. On the read side,field()surfaces rating custom columns as stars, so#myrat:4means 4 stars exactly likerating:4; the sweep's two-scales-in-one-library search gap closes with it. - Datetime-typed custom columns normalize like
set_pubdate: ISO text in UTC, never a naivestr()without the offset. - Empty enumerations accept nothing. Validation used to be skipped
when
display.enum_valueswas empty (upstream silently drops such writes instead); cquarry raises with a message that says so. Noneentries in value lists are skipped, never stringified into the literal string'None', in bothset_custom_columnandadd_custom_column_values.
remove_bookhad zero tests despite dynamic-table DELETEs, orphan pruning, and irreversibility: a dedicated fixture now carries the realbooks_delete_trgcascade and pins the upstream-faithful residue (an unreferenced Pattern-A value row survives removal, since no trigger purges it and upstream leaves it too).- The locked-database snapshot fallback, a README headline feature, has
its first witness: an
EXCLUSIVElock forces the snapshot path, the stderr notice fires, and the copy plus its-wal/-shmsidecars are cleaned up onclose(). - Six documented read APIs get their first tests
(
get_format_stats,get_identifiers,count_booksraw-then- cached,get_virtual_libraries,get_all_tags,get_tag_countswith zero-count tags), the~regex match kind gains its first exercise including the malformed-pattern-to-ParseExceptionpath, andformat_path_indexpins the honest POSIX case semantics. - Suite: 306 collected items at 1.14.0 to 360 here.
- Calibre-parity book creation, one batch, copy-only.
WritableCalibreDB.add_book(title, authors, ...)inserts(title, series_index, author_sort)FIRST and letsbooks_insert_trgfillsort/uuid(the caller never passes either), takes the id fromlastrowid, then writes theAuthor/Title (id)path with upstream's component budgets (PATH_LIMIT100 POSIX) and fallbacks. The ONE documented deviation: the ASCII fold is a plainencode("ascii", "replace")where Calibre runs an ICU user-codec first. Seeds: format FILES (copied atomically and placed asTitle - Author.ext,datarows carrying the real bytes on disk), a cover (path or bytes, JPEG/PNG sniff-or-raise so an unparseable cover is never catalogued withhas_cover=1), identifiers (types normalized likeset_identifier), one language (canonicalized likeset_languages), pubdate, publisher. No tags/series/ratings/ comments/custom columns at creation: phase 2 clears and sets those itself, phase 3 curates. An empty title becomesUnknown(Calibre parity); an empty author list is legal and leaves the book exactly whatfind_authorlessexpects. Everything runs in ONEbatch()with the setters reused inside it: any failure rolls the SQL back and the tracked directory is removed, so a failed add leaves zero rows, zero links, and no directory._touch_book()queues the id; Calibre generates the sidecar.opfon next startup. Copy only, by decision: sources are never moved or deleted; queue hygiene is the runner's policy. dry_run=Truewrites nothing and returns the computed plan: the id predicted fromsqlite_sequence(labeledpredicted_id), the predicted path, resolved authors with their sort keys and whether each row is new, format filenames and sizes, the cover filename, and thebooksrow that would be inserted.- New helper
helpers.sniff_image_format(data): the bytes-level sibling ofget_image_size(); the cover guard builds on it. - Manual pass against a copy of the testing-facility library
(user_version 27): predicted id matched the real id, triggers
filled sort/uuid, the format file landed with a truthful size,
and the read side (
get_book,get_format_path) resolves the book. The Calibre-GUI half of the pass (open the book, watch the.opfregenerate on next start) awaits Brandon. - Upstream sync: CalibreQuarry Phase 17 (
run phase2) is the driving consumer and awaits this release; Bindery (--install-to-calibrerepairs existing books only), Carrel-calibre-web (read-only), and Hermitage (read-mostly, no Flatpak pin bump) are unaffected per the Phase 10 roadmap box. - Suite: the
test_write.pyfixture gains the real INSERT-path hazards (books_insert_trg,books_pages_link_create_trigger,series_insert_trg, thefkc_insert_*guards) and 13 new add_book tests cover the row/link/path contract, atomic file placement, the sniff-or-raise, dry-run honesty, and both failure-compensation shapes.
clear_tags(book_id) -> intdetaches every tag from a book in one call: link rows deleted, now-orphaned tag rows pruned after (thefkc_delete_on_tagsorder),last_modifiedbumped and the book queued for OPF resync only when a link actually went away. The count of removed links comes back and an already-untagged book is an honest zero; until now a caller had to know the tags to clear them through per-nameremove_tag.add_custom_column_values(book_id, label, values) -> intgives multi-valued custom columns append semantics.set_custom_columnis replace-only, so puttingBrandonbeside an existingRinin#audiencetook a read-modify-replace dance; the new call dedupes against the book's existing values (the link table isUNIQUE(book, value)), collapses duplicates within the input, inserts only genuinely new links, and returns the honest count. Scoped tois_multiplePattern-A columns: single-valued and direct-storage columns raise with a pointer toset_custom_column, and a bare string is aTypeErrorrather than a comma-split guess.clear_rating(book_id) -> boolis a self-documenting alias ofset_rating(book_id, None)so an audit trail names the operation; the orphaned rating row prunes with it.- Consumers: CalibreQuarry's Phase 16 set-mode verbs and Phase 17's
#audiencestep sit directly on these three calls and need nothing new downstream; Bindery, Hermitage, and Carrel-calibre-web are unaffected in their recorded postures and no Flatpak pin moves. The phase-3-import skill'sUNIQUE(book, value)gotcha now teaches the setters instead of raw delete-then-insert, and the phase-1 skill was swept clean. - The test fixture's link tables now carry the real
UNIQUE(book, value)shape plus an is_multiple#audiencecolumn; suite 258 → 280 (the 255 baseline here was a typo; the same-day 1.12.0 entry below says 255 → 258 -- repaired 2026-09-09 so the two agree).
- New
analytics.genre_distribution(db)answers "what fraction of the library is each genre" over Calibre's genre-as-hierarchical-tags convention. Every dot-path node rolls up its subtree (Fic.Fantasy.Epiccontributes toFic,Fic.Fantasy, and itself), shares are fractions of every book in the library, and a book counts once per node even when several of its tags share an ancestor. Multi-genre books therefore land in several roots and the shares can legitimately sum over 1.0; the docstring says so where renderers will read it. Nodes come back depth-first (parents before children, siblings share-descending then name) so both a top-level headline slice and a full tree render are plain dict walks; books with no tags count under"untagged", last. - Deliberately not a duplicate of
get_tag_counts(flat per-tag link counts): the rollup, the denominator, and the untagged bucket are the delta the analytics module's scope rule demands. CalibreQuarry's--analytics genres(3.27.0) is the first renderer. - Suite 255 → 258.
list_books(sort="sort")never sorted. The hydrated rows store Calibre's title-sort undertitle_sort, so the publicsortkey resolved to a missing field, every comparator result was "equal", and the listing came back in cache order — which happens to BE title-sort order (the SELECT'sORDER BY b.author_sort, b.sort), so nothing looked wrong and the 1.10.0 tests passed by that coincidence. Found by the first real consumer: Carrel's descending search sort returned ascending results. The public key now maps to the row field (_ROW_KEY_ALIASES), and a regression test inserts books whose insertion order differs from their title order so a no-op sort can never pass again.- Suite 254 → 255.
sortaccepts a key sequence (primary first, one direction for all) alongside a single key, andauthor_sort/seriesjoin the key set — together covering the eight-button search-page sort header (authaz= author_sort, series, series_index;authza= the same descending) that motivated the API. None-valued keys still sink last regardless of direction; unknown keys raise. Implemented as an explicit comparator (functools.cmp_to_key) so the None handling and multi-key ordering are one readable rule.- Tests: suite grows 251 → 254 (multi-key ordering both directions with a NULL tie-breaker, unknown-key-in-sequence, empty-sequence).
- Consumer postures unchanged from 1.10.0: the fork adopts immediately; CalibreQuarry/Hermitage/Bindery waivers stand.
CalibreDB.list_books(*, ids, sort, descending, offset, limit)pages the cached rows for frontends that resolve a book-id set through the search engine and then paginate it — the read-side answer that lets Carrel unwind its hybrid (cquarry id sets paged through the fork's stock ORMfill_indexpage). Pure over the cache, no SQL of its own; the sort keys aresort(Calibre's title-sort),title,timestamp,pubdate,rating,series_index,id; None-valued keys sort last regardless of direction; unknown keys raise ValueError.- Consumer postures (the §2.3 wave): Carrel-calibre-web adopts in 0.6.31 (the wings/saved_searches/categories hybrid unwinds onto it). CalibreQuarry, Hermitage, and Bindery waived: CalibreQuarry's CLI has no pager surface (its verbs cover the need), Hermitage's grid is client-side over the full load, and Bindery has no listing surface at all.
- Tests: suite grows 245 → 251 (ids filtering, every sort key both directions, None-last semantics, offset/limit slices, error paths).
- Columns are addressable three ways.
load_custom_column()accepted only the exact Display Name (e.g.Translator(s)), while Calibre's native search,WritableCalibreDB.set_custom_column, and the#search grammar all speak the internal#label(#translators) — a UX asymmetry the 2026-09-02 import batch kept paying for. Newfind_custom_column(key)resolves one record by#label, bare label, or display name: a leading#matches the label only (labels are unique, never ambiguous); otherwise an exact display-name match wins (the historical key, so every existing caller keeps working) and a bare label is the graceful fallback; the label side is case-insensitive, mirroring the write module's_custom_column_meta.load_custom_column()routes through it, and its not-found error now lists both the name and the#labelof every column.get_custom_columns()stays keyed by display name (consumers iterate it; rekeying would break them). - Consumer note: Hermitage passes display names into
load_custom_columnand keeps working unchanged (display names still resolve); CalibreQuarry's--show-customlikewise. Bindery and Carrel-calibre-web do not read custom columns through this path.
WritableCalibreDB.clear_identifier(book_id, id_type)deletes one pair from the EAVidentifierstable and queues the book for OPF regeneration (_touch_book()), matching every other mutation. The type is normalized exactly likeset_identifier(stripped, lowercased; empty raises); a pair that is already absent is an honest no-op returning False; unknown books raise like every other setter. Motivated by the 2026-09-02 phase-3 import: cleaning invalidmobi-asinUUIDs and migratingamazonISBN-10s required dropping to raw SQL because the explicit helper did not exist (set_identifier(b, t, None)always worked but was not discoverable).- Tests: suite grows 241 → 245 (dual-resolution matrix with a display-name/ label ambiguity case, clear round-trip, normalization, no-op, dirtied-queue, unknown-book raise).
get_book_dossier(book_id, *, include_comments=False). The composed deep fetch detail views hand-assembled from ~10 read calls: the standard row,cover_path, per-format detail,custom_columnskeyed#labelas{name, datatype, value}(values exactly asfield()yields), annotations, reading positions, plugin data, conversion overrides, and ; only when flagged;commentsas{html, plain}.Nonefor unknown books. The frontend keeps rendering; cquarry owns the assembly.format_path_index()+find_book_by_path(). Every catalogued format path → book id in onedata ⋈ booksquery, built exactly likeget_format_path()and keyednormcase(normpath()), cached. Bindery'sCalibreIdResolverwas the seed consumer.cquarry.integrity. The mechanical "incomplete" predicates promoted from CalibreQuarry's--auditfrontend: untagged, unrated, authorless, formatless, coverless, missing cover files, deprecated formats (caller supplies the set), low-res covers ({id: (w, h)}), duplicates ((title, primary author)groups), series gaps. Pure over the cached rows; every id list sorted; the two cover-file checks are the only functions that touch the disk.cquarry.analytics.addition_timeline(month/year),author_stats(count-desc then name, star-scale averages, unrated excluded),rating_distribution(half-step stars,"unrated"last),vl_overlap(multi-wing combos only, unknown wings raise throughresolve_vl).- helpers, ISBN family.
isbn_normalize,isbn_check_digit_is_valid(ISBN-10 mod-11 / ISBN-13 EAN), andto_isbn13(978-prefix conversion; 13-digit inputs pass through; deliberately NO source check-digit validation, matching the LibraryThing exporter's contract this replaces). - helpers,
tag_rollup(counts). Subtree totals for dot-path counts: every node carries its own count plus everything below it; the rule Hermitage's_total_countand Carrel's category union already render, so adopting it is output-identical. The Phase 9 roadmap's example showed the keyed node keeping its bare count (Fic.Fantasy: 3where the subtree rule gives 5); the example was inconsistent with the render parity it was designed for and is corrected in the roadmap tick. - Docs split:
API.md+ README unbusy. The full per-method reference moved from README's Public API section intoAPI.md(with every new API above); the README keeps the hero, quick-starts (a dossier example joined the batch one), install, a one-line-per-module API-at-a-glance linking toAPI.md, the full search-grammar section, and the back matter. Spec gained §3.7 (integrity) and §3.8 (analytics); §6's consumer table refreshed for the sync releases below.
- Test suite 209 → 241:
test_integrity.py,test_analytics.py, plusTestDossierAndPathIndexand the ISBN/rollup batteries intest_helpers.py.
WritableCalibreDB.transaction()restored as an exact alias ofbatch(). The 2026-08-29 phase-3 import calledwith db.transaction():and hitAttributeError; 1.7.0 had shipped the deferred-commit context asbatch()only, and the session fell back to rawsqlite3. The alias keeps the pre-1.7 call shape working (sameBEGIN IMMEDIATEat entry, one commit at clean exit, full rollback on failure); the phase-3-import skill moved tobatch()in the same pass.
- Test suite 207 → 209: alias commit-at-exit and mid-block rollback cases mirroring the batch pair.
set_pubdate(book_id, value). Acceptsstr(YYYY-MM-DDor a full ISO datetime),date,datetime, orNone; naive datetimes are taken as UTC and the value is stored asdatetime.isoformat(' ')in UTC, which reproduces Calibre's TEXT rows byte-for-byte ('1991-10-01 07:00:00+00:00').Nonewrites the0101-01-01 00:00:00+00:00undefined-date sentinel, which the search engine already treats as absent. No-op honest: an equal instant returnsFalsewithout bumpinglast_modifiedor queuing OPF resync. This retires the raw-SQL pubdate workaround that put unix integers in the TEXT column and cost 8 linter errors on 2026-08-27.with wdb.batch():moves the commit boundary to the end of the block:BEGIN IMMEDIATEat outermost entry (the write lock held across the pass), nested batches join the one transaction, and a fault-injected mid-batch failure rolls back everything, including themetadata_dirtiedqueue. Every setter keeps its signature and per-call return semantics; only the commit boundary moves (_begin/_commit/_rollbackguard on_batch_depth).- Comments read surface.
get_book(book_id, include_comments=True)adds the raw stored HTML under acommentskey; rows otherwise keep omitting comment text (documented in docstrings at last). New bulkget_comments(book_id=None) -> {book: html}is the sanctioned bulk read, replacing consumers' reach-ins todb.conn(Hermitage's next sync adopts it). - Docs: spec §3.6 documents the setter and batch semantics; CLAUDE.md gains the
batch commit-boundary, pubdate-TEXT, and comments-omission contract notes;
README's write example shows
batch().
- Test suite 163 → 207 pytest-green runs: 17 new tests (batch atomicity/nesting/fault injection; pubdate round-trips str/date/datetime/aware-tz/sentinel/no-op/unknown-book; comments access including absent-table degradation) plus the parent cases the new fixture subclasses (TestBatchContext, TestSetPubdate, TestCommentsAccess) re-run by inheritance.
- Search fix (upstream parity). An empty numeric query;
rating:,size:,pages:,series_index:,id:with no value; now matches nothing (upstreamNumericSearch'sif not query: return matches) instead of raisingParseException. Regression-tested. - Refactor: shared
_BOOK_SELECT.get_all_books()andget_book()now build their rows from one SQL constant. They drifted once (v1.6.0'ssizefinding was the second instance of that class); the duplicate is what made it possible. - Transaction control pinned deliberately.
WritableCalibreDBnow setsautocommit=sqlite3.LEGACY_TRANSACTION_CONTROLexplicitly with the rationale in code: the write path'sBEGIN IMMEDIATE(take-the-write-lock-upfront,busy_timeouton acquisition) is impossible under PEP 249autocommit=False, which holds a transaction open from the first statement; verified empirically before choosing. - Lint hardening. Adopted
contextlib.suppress(5 sites), comprehensions over append-loops (3 sites), removed an unnecessary lambda wrapper, renamed an ambiguousl, fixed an ambiguous×in a docstring. Swept with strict rule families (F/E7/B/SIM/PERF/C4/RET/PLW/RUF); zero functional findings beyond the fixes above. - Deliberate non-change:
os.pathis kept overpathlib;Path.resolve()resolves symlinks whereos.path.abspathdoesn't, which would silently changeget_format_path()/ snapshot paths for symlinked libraries.
Hermitage, CalibreQuarry and Bindery swept with the same strict rule families: no functional
findings (global-statement caches, Pillow with-rebinds and script-level subprocess.run
calls are deliberate patterns). The .split(",") contract on native list fields: zero
violations across all three.
Driven by a full audit against the 7,631-book testing-facility library, cross-checked
against upstream Calibre source (calibre/db/search.py):
- User-category search (parity fix).
@Name:querynow works exactly as upstream'sget_user_category_matches: books holding any member value (exact match on the member's location),@Name:.queryincludes subcategories,falseinverts, other query text is ignored as upstream, <2-char queries match nothing. Groups and real fields win over same-named categories; unknown@Namesmatch nothing instead of silently degrading to anall:text sweep (the previous behavior). Providers opt in via an optionaluser_categories()hook (CalibreDBsuppliespreferences.user_categories). - Row-shape parity fix.
get_book()now returns the exactget_all_books()row shape: it was missingsize. Both rows additionally carryuuidandidentifiers(previously only the internal search view had them); enrichment degrades gracefully on ancient schemas. - Language ordering (parity fix). A book's
languagesfollowbooks_languages_link.item_order(link-id tiebreaker), matching Calibre; schemas predating the column keep link-id order. - New read APIs.
get_feeds()(thefeedsnews-recipe table),get_annotations_dirtied_books()(the annotations sibling of the OPF dirtied queue), andget_tag_browser_counts(); Calibre's owntag_browser_*sidebar rollups includingavg_rating, with custom columns rekeyed to#label. View quirks worked around without touching the database: the ratings view'sratingcolumn aliased toname, and the series view'stitle_sort()UDF (moved to stdlibhelpers; registered on the connection for the duration of the read only, then removed). Thetag_browser_filtered_*variants are deliberately skipped; they call Calibre's GUI-statebooks_list_filter()function, which only exists inside a running Calibre (as does themetaview'ssortconcat()aggregate;metastays unread by design). get_entities("ratings")completes entity coverage:{id, name, sort, link, count}with the half-star integer surfaced as text (resolves the v1.4.0 deferral ofratings.link).- Docs.
database_report.md§6 records the in-process-function landmines discovered during the audit (meta,tag_browser_filtered_*,title_sortin views).
- Test suite grew from 141 to 160 tests: user-category search battery (precedence,
inversion, subcategories, unknown names, hookless providers), feeds / annotations-dirtied
/ tag-browser reads with old-schema degradation, ratings entities, row-shape parity, and
language ordering on both current and pre-
item_orderschemas. - Verified against the real testing-facility library: all 9 tag-browser categories read,
recursive 31-VL "Unsorted" resolution unchanged,
get_all_books()still ~0.12 s.
- Entity setters:
set_authors(relink +author_sortrecomputation from per-author sort keys joined " & ", orphan pruning),set_series(+series_index: defaults 1.0 fresh assignment, preserved on reassignment; clearing nulls both),set_publisher,set_rating(0–5 stars stored ×2 with UNIQUE(rating) find-or-create dedup),set_languages(canonicalized to ISO codes via a new publicsearch.canonical_language). All NOCASE-matched, transactional, no-op-honest, orphan-pruning. set_comments: 1:1 upsert/clear on the UNIQUE(book) comments row; raw HTML stored verbatim (readers sanitize).- Custom-column writers:
set_custom_column(book_id, label, value)auto-detects storage pattern by link-table existence (Pattern A value+link vs Pattern B direct), validates enumerations againstdisplay.enum_values, accepts tristate bools, refuses non-editable columns and composite columns (no storage). - Format management:
add_format/remove_formatregister/dropdatarows (duplicate formats rejected case-insensitively);set_has_covertoggles the flag; files remain the caller's responsibility by design. - Book lifecycle:
remove_bookcleans custom columns in both storage patterns (PRAGMA-detected) and both dirtied queues before firing the cascade trigger, then prunes orphaned entities; verified against a real user_version-27 library including normalizedcustom_column_Nlayouts lacking abookcolumn. - New read API:
get_format_stats();{fmt: {count, bytes}}aggregates in one query (unblocks CalibreQuarry's deferred per-format disk-usage report).
- Test suite grew from 128 to 141 tests covering every setter's change/no-op/error paths, dedup semantics, validation failures, format round-trips, and removal cascades/pruning.
- Entity secondary columns: book rows gain
author_sortsandauthor_links; arrays parallel toauthorscarrying each author's true sort key and link URL (empty strings on ancient schemas). Newget_entities(kind)returns{id, name, sort, link, count}for authors / series / publishers / tags / languages, name-sorted;PRAGMAguards keep pre-column schemas working. (ratings.linkremains unread; no consumer need; revisit on demand.) - Custom-column display config:
get_custom_columns()values now includeeditable,normalized, and the decodeddisplayJSON (enum_values/enum_colors/composite_template…), with documented defaults on schemas predating those columns. - Generic preferences accessor:
get_preference(key, default)reads anything in thepreferencestable (JSON decoded where it parses, cached); typed helpers cover the high-traffic keys:get_field_metadata(),get_grouped_search_terms(),get_user_categories(),get_tag_browser_state()(order + hidden). - Grouped search terms (search parity): the engine resolves
GroupName:queryas a union over the group's member locations,GroupName:falseinverted, real fields winning over same-named groups, nesting raisingParseException; upstream semantics frompreferences.grouped_search_terms. Providers opt in via an optionalgrouped_search_terms()hook. - Annotation search: new
annotations:location matches each book's concatenated annotationsearchable_textwith full text-match kinds (substring/exact/regex) plustrue/falsepresence. Bare terms never sweep annotations, mirroring upstream. Documented deviation: ordinary matching instead of FTS stemming/ranking.
- Test suite grew from 118 to 128 tests: parallel-array shape, entity shapes/counts/kind errors, display-config decoding, preference typing, grouped expansion/inversion, annotation matching and all-exclusion.
- Native page counts: the
pages:search location now reads Calibre's ownbooks_pages_linktable first (upstream-managed since the CountPages integration; guarded by an OperationalError catch for older schemas), keeping an int custom column labelledpagesas fallback. Page counts also surface as apageskey on everyget_all_books()/get_book()row. This resolves former documented deviation §5 item 7. - New API
get_formats(book_id)returns per-format detail{fmt: {path, size_bytes, name}}(unverified path from the original DB location; catalogued uncompressed size; filename stem) so consumers can pick/report formats without rawdataqueries. - New API
get_cover_path(book_id, verify=True)resolves<library>/<books.path>/cover.jpgwith acover.pngfallback and disk verification (None when absent);verify=Falsereturns the catalogued path unconditionally. Raises ValueError for unknown books. - New API
get_library_uuid()exposes the library's identity UUID fromlibrary_id; stable across moves/restores, unlike per-book uuids; intended as the cache key for per-library state in web/GUI consumers. None on schemas without the table.
- Test suite grew from 111 to 118 tests: native-vs-fallback page precedence, end-to-end
pages:searches, UUID round-trips, format-map shape, and cover-path variants (jpg/png/missing/unverified).
- OPF sync fix: every
WritableCalibreDBmutation now records the book id in Calibre'smetadata_dirtiedqueue (INSERT OR IGNORE; guarded by a cachedsqlite_masterexistence check for pre-existing schemas). Previously writes only bumpedbooks.last_modified, but upstream regenerates a book's sidecar.opf; and re-pushes metadata to wireless readers; only for ids present inmetadata_dirtied(backend.pydirty_books()/dirtied_books()), so external edits never reached OPF/wireless sync. No-op mutations still queue nothing. - New read API:
CalibreDB.get_dirtied_books()returns the sorted, deduplicated ids awaiting resync so consumers can show what Calibre will pick up at its next startup ([]when the table is absent). Strictly observational; clearing the queue remains Calibre's job.
- Test suite grew from 104 to 111 tests: per-mutation dirtied-queue assertions (including no-op and duplicate-insert semantics against the real
UNIQUE(book)schema) and reader coverage for missing-table tolerance.
- Fix:
get_last_read_positions()now matches Calibre’s real schema; the table has nouser_typecolumn and its time field isepoch, notepoch_time. Against a live library the old SELECT silently returned an empty list on OperationalError; rows now surfaceid/book/format/user/device/cfi/epoch/pos_fracexactly as stored. Caught by Carrel-calibre-web’s CI fixture, which uses a schema dumped from a real library.
- New built-in locations:
size(total bytes across formats, honorsk/m/gsuffixes),pages(sourced from an int custom column labelledpageswhen present),title_sort, andseries_sort("Series [index]"). All are exposed inget_all_books()/search views. - Saved-search interpolation:
search:"Name"resolves through the newCalibreDB.get_saved_searches()/saved_search(). Nesting works, cycles raiseParseException, unknown names raiseParseException, and lookups are case-insensitive. - Multi-valued count operator:
tags:#>3,authors:#=2,formats:#<5,identifiers:#>=2. - Language canonicalization:
languages:Englishmatches books stored asengvia a 55-entry ISO 639-2 map; unknown tokens pass through untouched. - Slash date separators:
pubdate:=1965/08/01andtimestamp:2006/07work alongside hyphens. - Expanded boolean keywords:
checked,unchecked,blank,emptyplus_-prefixed variants, accepted in both boolean (cover:) and numeric/tristate (rating:blank) positions. - Strict virtual-library errors:
vl:UnknownraisesParseExceptioninstead of silently returning no matches; VL name resolution is case-insensitive end-to-end (resolve_vl,vl_expression). - Component matching everywhere: Calibre's leading-dot modifiers under
=;.foo(subtree) and..foo(component); now apply to all text fields, not just hierarchical tags. allsweeps custom text columns: bare terms also search custom text/enumeration/tags-like columns.
get_book(book_id)fetches one hydrated record without scanning the library.search_books(query)returns hydrated books for a search expression directly.get_format_path(book_id, fmt, verify=True)resolves<library>/<books.path>/<name>.<fmt>from the original DB location (snapshot-safe).tags_to_tree(tags)builds nested dicts from dot-delimited hierarchies for TreeView rendering.normalize_rating(int)is the canonical name for the 1–10 → 0–5 star conversion (calibre_rating_to_starskept as alias).get_vl_ui_state()exposes Calibre's storedvirt_libs_hidden/virt_libs_orderso frontends can mirror the GUI sidebar exactly.resolve_saved_search(name)resolves a saved search to book IDs with case-insensitive matching.
get_annotations(book_id=None)extracts highlights/bookmarks/notes from theannotationstable with JSONannot_datadecoding.get_last_read_positions(book_id=None)maps per-device reading progress (pos_frac, CFI, epoch time).get_plugin_data(book_id=None, name=None)reads third-party payloads (Goodreads IDs, word counts, page counts) frombooks_plugin_data.get_conversion_profiles(book_id=None)lists manual conversion overrides; the pickled blob stays raw bytes (never unpickled).strip_html(html)reduces comments HTML payloads to safe plain text (tag stripping, entity unescaping, whitespace collapse).
cquarry.write.WritableCalibreDB: a separate class that is unreachable from read-onlyCalibreDB. Registers Calibre's trigger dependencies (title_sort(),uuid4(),PYNOCASE) before any statement, usesBEGIN IMMEDIATEtransactions, bumpsbooks.last_modifiedon every mutation, and cleans link tables before tag deletion to satisfyfkc_delete_on_tags.- APIs:
update_title(),add_tag()/remove_tag()(returns whether state changed),set_identifier()/set_identifiers()batch upserts honoringUNIQUE(book, type).
- Dropped
re.Scanner: the tokenizer is a plainre.finditerscanner over a documented pattern; pure documented stdlib. - Test suite grew from 47 to 104 tests covering every feature above.
- Fix: Prevented infinite loops in
get_jpeg_sizeby asserting frame payload lengths are valid. - Fix: Re-wrote AST quoted-colon parsing block in
_base_tokento successfully preserve strings likeidentifiers:isbn:"value". - Fix: Fixed logical rating searches (
#rating:false) by properly declaring#ratingasDT_RATING. - Fix: Added dynamic series index generation for custom
#seriescolumns in location routing. - Fix: Eliminated Calibre format-splitting bugs for author/tag strings containing commas by omitting
GROUP_CONCATin favor of dictionary mapping and native python lists. - Fix: Shielded
resolve_vlfrom virtual library recursion explosions. - Fix: Extracted date-time components accurately in ISO-8601 targets, preventing exact match failures on ISO strings containing
T.
- Build: Configured pyproject.toml to ignore strict ruff lints blocking the CI pipeline.
- Lazy-Loaded Comments & Custom Columns:
db.pyno longer eagerly loads thecommentsHTML payloads or large custom column tables into memory when building the search view. These are now fetched from SQLite strictly on-demand per book ID during search expression evaluation. This reduces memory footprint and snapshot copy time for libraries with extensive HTML comments.
- Initial Extraction: Graduated
cquarryinto a standalone shared library. - Database Engine (
cquarry.db): Features theCalibreDBwrapper, which intelligently managesmetadata.dbaccess, falling back to a WAL-consistent snapshot if the Calibre desktop application holds an exclusive write-lock. Exposesget_all_books(), tags, series, and identifiers with performant SQLite JOINs and internal memory caching. - Search Grammar Engine (
cquarry.search): A full recursive descent parser implementing Calibre's search expression logic. Provides boolean logic (AND,OR,NOT), exact matching (=value), hierarchical tag prefix matching (tags:FicmatchesFic.Fantasy), date math (date:>14daysago), and nested Virtual Library resolution (vl:"My Books"). - Helpers: Inherits standard Calibre domain formatters from CalibreQuarry (star rating converters, deterministic missing series gap detection, and binary image dimension sniffing).