Asynchronous crawler for extracting text content from web pages.
- Async with
aiohttp+asyncio - Depth-limited recursion
- Optional URL filtering (regex)
- Optional “since” filtering via HTTP
Last-Modified - Pluggable callbacks (
data_handler,stop_handler,link_extractor)
- Python: 3.11+ (tested on 3.13)
- Dependencies (capped below the next major per PEP 440 intent):
aiohttp>=3.9,<4beautifulsoup4>=4.12,<5
Install:
pip install -r requirements.txt--
- Create a Python 3.13 virtual environment in the project root
python3.13 -m venv .venv- Activate the virtual environment
.venv\Scripts\Activate.ps1- Install dependencies
pip install -r requirements.txt- Run the sample script (aiohttp route)
python scrape_test.py- Create a Python 3.13 virtual environment in the project root
python3.13 -m venv .venv- Activate the virtual environment
source .venv/bin/activate- Install dependencies
pip install -r requirements.txt- Run the sample script (aiohttp route)
python scrape_test.pyYou can customize how links are extracted from each page by providing link_extractor.
If omitted, a built-in default extractor is used (collects all <a href> links, resolves to absolute URLs, and keeps only http/https).
Signature:
def link_extractor(page: ScrapedPage) -> list[str]: ...Example:
from urllib.parse import urlparse, urljoin
from web_crawler import scrape_website, ScrapedPage
def only_article_links(page: ScrapedPage) -> list[str]:
links: list[str] = []
# Example: keep only URLs containing "/articles/"
for a in page.soup.select("a[href]"):
href = a["href"]
if "/articles/" not in href:
continue
abs_url = urljoin(page.url, href)
if urlparse(abs_url).scheme in ("http", "https"):
links.append(abs_url)
return links
# Use it:
# await scrape_website(url, data_handler=..., link_extractor=only_article_links)You can optionally fetch “rendered” HTML using Playwright, which helps with SPA/JS-heavy pages. This is an additive feature: the default aiohttp-based crawler continues to work as-is. Playwright is only needed when you explicitly use it.
- Playwright and a browser runtime (e.g., Chromium) are optional dependencies.
pip install -r requirements-playwright.txt
playwright install chromium
# On Linux (recommended to include OS deps)
# playwright install --with-deps chromiumNotes:
- You may also install
firefoxorwebkitinstead ofchromium. - If Playwright is not installed, everything still works as long as you don’t call the Playwright adapter.
python scrape_test_playwright.pyThis script will:
- Navigate to the Yahoo News top page with Playwright
- Extract the rendered HTML, page title, links, and visible text (same extraction logic as the default path)
- Apply the same filters as
scrape_test.py(since within last 2 days and Yahoo article URL pattern)
If you want to switch to Playwright only for certain pages, enable the flag:
await scrape_website(
url=...,
data_handler=...,
use_playwright=True,
playwright_options={
"wait_until": "networkidle",
"timeout_ms": 30000,
"headless": True,
# "wait_for_selector": "article",
},
)Design notes:
- Delayed import: Playwright is imported inside functions, so environments without Playwright/Chromium are unaffected unless you enable it.
- Headers: Response headers are accessed in a case-insensitive way (
last-modified/Last-Modified). - Concurrency: Playwright is heavier than aiohttp; limit parallelism when using it.