Skip to content

Latest commit

 

History

History
77 lines (60 loc) · 3.83 KB

File metadata and controls

77 lines (60 loc) · 3.83 KB

pagespring

PyPI

Acquire and normalize online documentation into clean, convertible source files — the acquisition front-end to pagespeak.

Point it at a manual's URL. A pattern recognizes the source type, acquires the raw pages (stdlib urllib), and normalizes them into ONE clean HTML/markdown file with absolute asset URLs under incoming/<slug>/. That clean file is the deliverable; converting it into the finished RAG corpus is a separate step (pagespeak) that consumes incoming/ on its own — pagespring never runs it.

Lean by design: pf-core[cli] (PyPI) + beautifulsoup4, pyyaml, and pypdfium2 for PDF page counts and for cutting PDFs printed as 2-up spreads into single pages. Stdlib fetch, no ML stack.

Intended use

pagespring is for publicly available documentation — vendor manuals, help centers, open textbooks, API specs. It fetches only what the source serves to any reader: there is no login/session handling, no paywall traversal, and no bot-detection evasion. Where a help site's pages load their articles with a form POST to the site's own API, it sends the same guest request a browser does. It is a polite client: it identifies itself with a pagespring/<version> User-Agent (PAGESPRING_UA overrides it), honors 429 Retry-After, backs off on server errors, paces crawl requests, and caps crawl sizes.

It is a user-invoked archiver — closer to "Save Page As" than to an autonomous crawler. Every source is a URL you supply (one per ingest, or one per line of an ingest --batch file), and it never discovers sources on its own, so it does not consult robots.txt (which governs bots that find URLs themselves), except to warn when a Fluid Topics portal's robots.txt disallows the API an ingest reads. Before mirroring a site, check its terms of use. What you may do with the acquired copy (personal RAG corpus, internal search, redistribution) is governed by the source's license — the deliverable under incoming/ stays on your machine, and nothing is re-published by this tool.

Install

pip install pagespring

Quick start

pagespring ingest https://docs.tableplus.com   # acquire + normalize → incoming/tableplus/
pagespring ingest --batch manuals.txt           # one ingest per URL line, with a summary
pagespring renormalize <slug>                   # replay normalize from kept raw/ — no re-crawl
pagespring refresh --all                        # re-check every manual against its source
pagespring audit --all                          # $0 sanity checks on everything staged
pagespring localize <slug>                      # pull a deliverable's images later (resumable; --all)
pagespring patterns                             # list the source patterns
pagespring classify <url>                       # which pattern handles a URL (no fetch; --probe names docs_probe's route)
pagespring status                               # what's been acquired

Deliverables land in ./incoming/<slug>/ under the directory you run from.

Dev

bin/setup   # clone → venv + editable install with dev extras
bin/test    # pytest
bin/lint    # ruff check + ruff format --check + mypy (strict) + structural gate + framework-first

See docs/usage.md for the full command set and docs/architecture.md for the acquire → normalize flow and how to add a new source pattern.

License

Apache-2.0 — see LICENSE. Releases through 0.11.0 were published under the MIT license and stay MIT.