Skip to content

fix(scan): scheduled RescanFolders should reconsider unmatched files - #187

Open
4o66 wants to merge 1 commit into
pennydreadful:developfrom
4o66:fix/rescan-scheduled-filter-matched
Open

4o66 wants to merge 1 commit into
pennydreadful:developfrom
4o66:fix/rescan-scheduled-filter-matched

Conversation

@4o66

@4o66 4o66 commented Sep 22, 2026

Copy link
Copy Markdown

Database Migration

NO

Description

The scheduled task builds RescanFoldersCommand() with Filter = FilterFilesType.Known.
Known skips any file whose size and mtime are unchanged, regardless of whether it is
still linked to an edition
.

When a book is deleted — including by a metadata refresh — BookDeletedEvent reaches
MediaFileService, which calls UnlinkFilesByBook; the BookFiles rows survive with
EditionId = 0. Under Known those rows are never looked at again, because the file on
disk has not changed. The row's own existence is what prevents its repair, so the library
cannot recover on its own — the only ways back are a manual unmapped-files import or
deleting rows by hand. 3,172 files were stranded this way on my library.

FilterFilesType.Matched already implements the right rule: skip only if unchanged
and (matched or already remote-searched, per #183). RefreshAuthorService already
pushes its rescan with Matched; the scheduled task is the last caller still on Known.
This changes that default to match.

Evidence. A full-library scan on the stock default logs Filtering 9812 files for unchanged files -> Analyzing 0/9812 files while 1,382 BookFiles rows sit at
EditionId = 0. That is Known working as designed: nothing on disk changed, so every
file is excluded no matter how many are unlinked, and the unlinked ones can never come
back. A scan that does consider unlinked rows analyzed 252 of the same 9,812 and imported
113 files.

Cost, and a caveat on the marker this leans on. Matched skips an unchanged,
unmatched file only once LastRemoteSearchTime is set, and
IdentificationService.StampRemoteSearchResult sets it only when
Edition == null && RemoteSearchSucceeded. Two classes of file therefore never get
stamped and are re-analyzed on every scheduled scan:

  • Files matched to an edition but then rejected at import — Has the same filesize as existing file, Not an upgrade for existing book file. Edition is non-null so no
    stamp is written, and EditionId stays 0 so they are never skipped. This is the larger
    class on my library.
  • Files whose remote search returned before any TryRemoteSearch call — the Goodreads ID
    path in CandidateService is deliberately not routed through it, and there is an early
    yield break when a file carries no usable author or title tag.

On a 9,812-file library, 1,110 of 1,382 unlinked files are stamped and correctly skipped;
the remaining 272 are re-analyzed nightly, costing 5m24s per scan and importing nothing.
Bounded here, and it scales with the number of permanently unmatchable files rather than
with library size — but on a library with many of those it would be felt.

That marker is yours from #183 and I did not want to second-guess it inside a one-line
default change, so this PR leaves it alone. If you would rather stamp on any completed
identification attempt, or add a backoff, I am happy to do that here instead of — or as
well as — flipping the default.

Todos

  • Tests — RescanFoldersCommandFixture. FilterFixture already covers the Matched
    semantics themselves (filter_unmatched_should_return_existing_file_if_unmatched,
    added in don't keep searching if we didn't get a match #183)
  • Translation Keys — not needed
  • Wiki Updates — arguably; the scheduled task's behavior changes for unmatched files

Notes for reviewers

  • Users with many permanently unmatchable files will see the nightly scan get slower —
    that is the point of the change, but it is a visible difference and may deserve a line
    in release notes.

Issues Fixed or Closed by this PR

None. I looked for an existing report of this and did not find one.

The scheduled task used FilterFilesType.Known, which skips any file whose
size and mtime are unchanged regardless of whether it is still linked to an
edition. When a book is deleted, BookDeletedEvent reaches MediaFileService,
which calls UnlinkFilesByBook and leaves the BookFiles rows with
EditionId = 0. Under Known those rows are never looked at again, because the
file on disk has not changed, so the row's own existence is what prevents
its repair.

FilterFilesType.Matched already implements the right rule: skip only if
unchanged and (matched or already remote-searched). RefreshAuthorService
already scans with Matched; the scheduled task was the last caller on Known.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant