Watches your Goodreads to-read shelf, downloads each book, then places it into Kavita, BookLore, Grimmory and Audiobookshelf and adds it to Open Notebook — ready to read, listen to, and take notes on. One container, and it never makes a second copy of a file.
Self-hosted · single-operator · no build step
Goodreads to-read
-> acquire Shelfmark: search + download (ebook, audiobook)
-> classify one category, from the book's genres
-> place ebooks: rename into <Category>/
audiobooks: rename into Author/Title/
-> index rescan Kavita, BookLore, Grimmory, Audiobookshelf
-> notebook add to the matching Open Notebook notebook
-> verify confirm each service really sees the file
-> shelve move off to-read onto collected-pdf / -audiobook / both
Two periodic jobs run alongside that chain: discover reads the shelf, and
reconcile checks what was recorded against what Goodreads actually shows.
- Reads the to-read shelf through a captured browser session, keyed on the Goodreads id — a re-run refreshes metadata instead of creating duplicates
- Paginates a shelf that has outgrown one page
- Reconciles recorded shelves against Goodreads every 6 hours and resets anything that drifted
- Moves finished books off the exclusive to-read shelf using the shelve / confirm / destroy / re-shelve / confirm sequence Goodreads requires, creating the destination shelf first if it does not exist
- Searches, ranks, hands the best release to Shelfmark, and watches the task
- Rejects releases on the indexer's Newznab category, on size, and on video-release name patterns before ranking — so a film cannot be queued as an audiobook
- Records every release id tried, so a dead release falls through to the next-best candidate (up to 4 per book) rather than failing the book
- Re-queues a task missing from Shelfmark's in-memory queue for 10 minutes
- Refuses to start an audiobook download below a configurable free-space floor
- Walks four genre providers — Goodreads AppSync, OpenLibrary, Google Books,
embedded epub
dc:subject— stopping at the first that answers - Each provider is isolated, so one timing out cannot stop the next; the embedded source works with every API down
- Falls back to matching the title against the same rules, stored as
genre_source = title, so an inferred category is always distinguishable - Flags anything matching nothing as needs review rather than misfiling it
- Ebooks →
<Category>/Author - Title (Year)/, audiobooks →Author/Title/ - Every placement is
os.rename: a move that would cross a filesystem raises rather than falling back to a copy, so a duplicate is impossible - An existing destination is never clobbered — the incoming file is dropped only when it is provably the same inode; everything else gets a suffix
- Re-running placement is a no-op
- Rescans Kavita, BookLore and Grimmory for books, Audiobookshelf for audiobooks — nothing is imported or uploaded
- Skips an app whose tree did not change rather than failing it
verifysearches each service before anything irreversible happens on Goodreads, matching on a normalised word-boundary form with the leading article ignored, treating a differing series number as disqualifying- A miss within 10 minutes of a rescan blocks rather than fails, because rescans are asynchronous
- Each ebook is attached to the notebook its category maps to, handed over as a
file_pathand never as an upload, so no second copy is stored - Attachment is verified by reading the source back; a source attached more than once is collapsed
- A per-service circuit breaker holds every book waiting on a service after 3 consecutive transient failures — an upstream outage becomes one row naming the service instead of hundreds of failed books
- A half-open probe lets one book through to discover when it is back, and a stuck hold can be released by hand
- Separately, each book gets a 24-hour grace on transient faults, measured from when the outage started — not from how many times it was retried
- Each sweep gets a 120-second budget and advances up to 6 books concurrently, serving least-recently-touched first, so a bounded pass still rotates through the whole list and publishes its result
- Every service is probed on a 5-minute timer with an authenticated call — never a bare health endpoint that a wrong API key would pass
- Every failed stage carries a
failure_kind(auth,network,server,busy,data), and the UI groups failures by cause rather than by book, with a suggested next step in plain language and a Retry all where retrying helps - A broken credential raises a banner on every page with a link to the fix
| Layer | What it uses |
|---|---|
| API + UI | FastAPI, Uvicorn — UI served straight from app/static/, no framework and no build step |
| Storage | SQLite (WAL), one file, no ORM |
| Goodreads session | Playwright + Chromium on a virtual display, streamed to the browser over noVNC |
| HTTP | httpx, BeautifulSoup |
| Crypto | cryptography (Fernet) for credentials, scrypt for passwords |
| Packaging | Docker, on Playwright's own image so Chromium is version-matched |
git clone https://github.com/ohmzi/goodreads-pipeline.git goodreads
cd goodreads
cp .env.example .env
python3 -c "import secrets; print(secrets.token_urlsafe(48))" # -> GOODREADS_SECRET_KEY
chmod 600 .env
docker compose up -d --buildLeft unset, 8091 is published on every interface, and a published port is
not covered by ufw — docker's iptables rules are evaluated before the
firewall's. It is worth narrowing with PUBLISH_HOST in .env, and easy to
get wrong in a way that takes the app offline: a binding to one address
answers only on that address, so a value your front end does not dial means
every request is refused with nothing in the app's logs at all.
Confirm where your front end reaches you from first:
docker compose logs goodreads | grep 'GET /login'and note that a proxy or tunnel running in a container dials you as its
bridge gateway (172.x.0.1), never as 127.0.0.1 — so 127.0.0.1 suits only a
proxy running directly on the host. 0.0.0.0 (the default) is always safe.
This decides which of the host's addresses the port answers on. It does not decide whether the app is reachable from the internet: a tunnel in front reaches it whichever value is set. See SECURITY.md.
Create a login, then open http://<host>:8091:
docker compose exec goodreads python -m app.cli set-password you --generateThere is deliberately no default account. Library paths and any extra networks
belong in a gitignored docker-compose.override.yml; the committed compose
uses placeholders so a clone runs unedited.
📘 Full instructions — mounts, networks, first run in the UI, and running without the container — are in SETUP.md.
python -m app.cli report # what completed, what failed, and why
python -m app.cli repair # fix notebook links and stuck stages (dry run)
python -m app.cli audit # duplicate audiobooks and titles in two places
python -m app.cli rename # tidy loose audiobooks into Author/Title
python -m app.cli reclassify # re-run classification over the library
python -m app.cli backfill-genres # re-resolve genres and re-categorise
python -m app.cli forget-session # drop the stored Goodreads session + browser profilereport and audit are read-only. repair, rename, reconcile and
reclassify change nothing without --apply. Two write immediately:
backfill-genres, and forget-session, which deletes the stored Goodreads
session and the browser profile and refuses while a login browser is running.
This app is the most sensitive service you will run. It holds the key that decrypts every stored credential, it can drive a browser signed into a real personal account, and it can write to a real library tree.
Passwords are hashed with scrypt and compared in constant time. Sessions are
stateless HMAC-SHA256 tokens carrying a password epoch, so changing a password
evicts every outstanding session — and so does signing out, on every device
at once. Failed logins are throttled by username and client address with a
progressive delay rather than a lockout, so nobody can shut the owner out of
their own UI. Every credential is Fernet-encrypted at rest behind a
GOODREADS_SECRET_KEY that must be at least 32 characters, and the data volume
is kept owner-only by a umask 077 plus a one-off tighten at startup. Requests
from another site are refused on the server, not only by SameSite, and the
calls the app makes to your other services refuse redirects, cap response
bodies, and withhold error bodies from anything that carried a credential.
The noVNC desktop has no published port: x11vnc binds to loopback inside the
container and the app bridges to it over a WebSocket that checks both the
origin and the session. That leaves the one published port as the only way in,
so the sign-in page is what stands in front of everything. PUBLISH_HOST
decides which of the host's addresses that port answers on — it does not decide
whether something in front can reach it.
⚠️ The threat model, credential storage design, reverse-proxy deployment, and an explicit list of known limits — including that the session cookie is a bearer token and that traffic is plain HTTP without a TLS terminator in front — are in SECURITY.md. Read it before exposing this beyond a host you control.
Secrets are split by lifetime:
.envholds what must exist before the database can be read; service credentials are typed into the Settings page, encrypted with Fernet, and stored in the database.
app/categories.yml is the rule table that decides where a book lands — a
category key per genre needle, mapped to a folder and an Open Notebook
notebook. Matching is substring, case-insensitive, and longest-needle-first, so
science fiction beats fiction regardless of file order. The file ships in
the image and /data/categories.yml overrides it, so tuning survives a
rebuild, and it is re-read per classification, so an edit needs no restart.
📗 Every environment variable, every credential, the genre provider chain, and the library-root mapping each service sees are in CONFIGURATION.md.
Goodreads removed its public API in 2020, then removed public shelf pages, and
/user/sign_in is now a stub that hands off to Amazon's sign-in flow. There is
no form to POST a password to — every maintained Goodreads tool today reuses a
persistent browser session instead.
So this runs Chromium on a virtual display inside its own container and streams it to the login page over noVNC. You sign in once, the cookies are saved, and everything afterwards runs over plain HTTP with them. When the session eventually expires the UI says so and you do it again.
| Document | What is in it |
|---|---|
| 🚀 SETUP.md | Install, mounts and networks, first login, first run, bare-metal |
| ⚙️ CONFIGURATION.md | Every variable, credential, and path mapping; categories.yml; genre resolution |
| 🔐 SECURITY.md | What is enforced and where (the whole posture on one page), threat model, credential storage, sessions, reverse proxy, known limits |
| 🔧 OPERATIONS.md | Upgrading, backup and restore, the maintenance CLI, health, troubleshooting |
| 🏗️ ARCHITECTURE.md | Components, module map, the database, concurrency, the path namespace problem |
| 🔄 PIPELINE.md | Every stage in detail, the scheduler, retries, the service breaker |
| 🔌 INTEGRATIONS.md | The surrounding services, the settings each needs, and what went wrong |
| 🌐 API.md | The HTTP API |
| 📝 VERSION.md | Release notes |
Apache-2.0 © 2026 Omar