Skip to content

Feature: Add the ability to process more than 1 page of AA search results - #1383

Open
RoninTech wants to merge 5 commits into
calibrain:mainfrom
RoninTech:feature/handle-all-aa-result-pages
Open

RoninTech wants to merge 5 commits into
calibrain:mainfrom
RoninTech:feature/handle-all-aa-result-pages

Conversation

@RoninTech

Copy link
Copy Markdown
Contributor

Now that we are showing how many AA results are actually being displayed it is front and centre that the default is to ignore AA results beyond 50. ABB searches already have a config item to allow for processing more than 1 page of search results. This PR adds a new config item for Max Anna's Archive Results Pages that defaults to 1 to match the original AA result handling, and allows for more than 50 results when it is set to > 1. The maximum allowed pages to process is 10 which matches AA's search results, which only tell us there are 500 or more results.

The Max Anna's Archive Results Pages default of 1 will perform the same as current AA direct searches. Increasing it will obviously increase the time for a search to return but can provide a larger set of the original AA paginated results.

Here are a couple of images to show the change:

Direct-dl-settings-max-AA-pages direct-max-3-pages-AA-downloads

Coded with llama.cpp, opencode and 🤖

@RoninTech

RoninTech commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor Author

There is a slight misalignment of the release number column, not perfectly centred between book cover and title/author, which is beyond my limited html/css abilities. Hopefully that's an easy fix @calibrain. 🙂

@calibrain calibrain left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @RoninTech
I just reviewd it, and got a few comments :

Please let's avoid doing one request per book to get the download numbers.
It's really heavy on the system (in some cases it created 100+ extra request, that all have to go through the slow browser to bypass Cloudflare)

Could you instead make a query using display=list (instead of our display=table) since that one contains download stats ?

Also, If a later page fails or the search time budget runs out partway, the whole search errors out. Pages already fetched are lost

@RoninTech

Copy link
Copy Markdown
Contributor Author

Sounds good. I'll take a look and get back to you.

@RoninTech

RoninTech commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor Author

The AA search display=list output does show the downloads info but it is behind javascript calls so a simple GET request will not be able to access them. In shelfmark it seems that the mechanism to use in this situation is the CDP_GET which is a lot slower but can see the JS populated download values.

Please let's avoid doing one request per book to get the download numbers.
It's really heavy on the system (in some cases it created 100+ extra request, that all have to go through the slow browser to bypass Cloudflare)

The way I was doing it was by fetching the AA search result pages (1 to 10 based on config setting) with simple requests.get() calls, then using the <AA_URL>/dyn/md5/inline_info/{book_id} API for each book, also via requests.get(), to get a tiny JSON object that includes the "downloads_total". Here is an example of that JSON:

{"reports_count":0,"comments_count":0,"lists_count":1,"downloads_total":130,"great_quality_count":0}

Using this AA inline_info API has no browser overhead, no JS execution, and no Cloudflare challenge so it is very fast, and we sped it up further by doing batches of 5 in parallel.

Today I changed the code to use CDP_GET for the search results pages so we can see the JS populated downloads without extra requests and it is quite a bit slower in comparison. Here is the difference doing an AA search for "The Silmarillion Tolkien" which has 144 matches on AA and so returns 3 pages of results:

Fetch Type Page 1 Page 2 Page 3 AA API Total
Simple GET 3.5s 3.6s 3.5s 36.7s 47.4s
CDP_GET 24.4s 24.9s 22.9s N/A 72.2s

I'm running this on a little mini-PC with a good net connection and don't notice any impact to performance while doing either of these style of searches. If we keep the "Max Anna's Archive Results Pages" set to it's default of 1 then that should mean a max of 1 (for the list page) + 50 (for each books inline_info) requests.

ASIDE: While doing this I noticed that I have access to "great_quality_count":0 via the inline_info AA API. That is based on AA users assigning stars to a book release which appears as, for example: · ⭐2 ·. It's a great input into the decision on which AA book release to grab.

shelfmark_direct_dl_stars

Let me know how you'd like me to proceed with this. And thanks for this great tool, I'm really enjoying using shelfmark and it's sped up my process a lot. 👍

@RoninTech
RoninTech force-pushed the feature/handle-all-aa-result-pages branch 2 times, most recently from 3f68b49 to 093fb14 Compare September 28, 2026 08:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants