Skip to content

Repository files navigation

BloomsReader

English | 简体中文

A Jina Reader (r.jina.ai)–compatible Cloudflare Worker that turns any web page or document into clean, LLM-friendly Markdown — from a URL or a direct file upload, with one command.

Web pages Direct fetch → simple render → full CDP render (optional), then readability extraction
Documents PDF / DOCX / PPTX / XLSX / images & legacy Office (.doc .ppt .xls) via MinerU; MarkItDown fallback
Cache D1-backed, 12-hour TTL, hourly cron cleanup
Rate limit Per-IP fixed window, default 1 request / 15 s (configurable)

Highlights

  • Familiar API — Drop-in-style routes and headers modeled after Jina Reader; works with existing tools that target r.jina.ai.
  • Great for LLMs — Readability extraction + MarkItDown produce clean, structured Markdown; no boilerplate or tracking scripts.
  • Documents made easy — PDFs, Office files, images and more parsed into text, tables and structure via MinerU; graceful MarkItDown fallback if MinerU is unavailable.
  • Costs almost nothing — Serverless, pay-as-you-go. Browser rendering only when needed; document parsing can run on the token-free MinerU lightweight API.
  • Private by design — Per-IP rate limiting; optional AUTH_TOKEN; SSRF blacklist for private / loopback / metadata addresses.
  • Built for the edge — Runs on Cloudflare Workers with a D1 cache; Tiny, typed, tested.

Quick Start

# Web page → Markdown (Jina style: URL directly appended)
curl https://<your-worker>/https://example.com

# JSON mode (same shape as Jina: {code, status, data:{title, url, content, usage...}})
curl -H "Accept: application/json" https://<your-worker>/https://example.com

# Document → Markdown (token-free MinerU lightweight API)
curl https://<your-worker>/https://cdn-mineru.openxlab.org.cn/demo/example.pdf

# Direct file upload (multipart)
curl -F "file=@report.docx" https://<your-worker>/

# POST with options
curl -X POST -H "Content-Type: application/json" \
     -d '{"url":"https://example.com","timeout":90}' https://<your-worker>/

Live demo: https://<your-worker>/https://example.com


How It Works

flowchart TD
    A[Request] --> B{Auth<br/><i>optional</i>}
    B -- "ok" --> C[Rate limit<br/><i>per IP</i>]
    C -- "allowed" --> D[Classify content]
    D --> E[Web page]
    D --> F[Document<br/>PDF / Office / image]

    E --> G{Direct fetch<br/>content valid?}
    G -- "yes" --> H[Readability extraction]
    G -- "no" --> I[Render]
    I --> J[Simple render]
    I --> K[Full CDP render]
    J --> H
    K --> H

    F --> L[MinerU agent<br/><i>≤ 10 MB</i>]
    L -- "success" --> M[Clean Markdown<br/>title · url · content]
    L -- "fail" --> P{Precision API<br/>token set?}
    P -- "yes" --> M
    P -- "no" --> Q[MarkItDown fallback]
    Q --> M

    H --> M
    M --> O["12 h D1 cache<br/><i>x-no-cache to bypass</i>"]
Loading

Usage

Endpoints

Endpoint Method Description
GET /{url} GET Web page or document → Markdown (Jina style). Encode the target as part of the path.
GET /?url={url} GET Same, with the target as a query parameter.
POST / POST {"url": "..."} (JSON) or multipart form with a file field.
GET /llms.txt GET Machine-readable API documentation (served to LLMs).
GET / GET Human-friendly landing page (never rate-limited).

Response format

Default (text/markdown):

Title: Example Domain

URL Source: https://example.com/

Markdown Content:
...
x-respond-with Output
(default) Markdown + title block; cache hits append Warning: This is a cached snapshot...
markdown Full-page Markdown (skips readability extraction)
html Page HTML
text Plain text
frontmatter / markdown+frontmatter YAML frontmatter + Markdown
json Equivalent to Accept: application/json

Request headers (Jina-compatible)

Header Description
Accept: application/json JSON response mode
x-respond-with Output mode (see above)
x-timeout Total timeout in seconds, 5–360 (default 360)
x-engine auto (default) / curl (direct fetch only) / browser (render first)
x-target-selector CSS selector; return only matching content
x-wait-for-selector CSS selector to wait for when rendering
x-remove-selector Comma-separated CSS selectors to strip from the result
x-set-cookie Cookie forwarded to the target (such requests are not cached)
x-no-cache: true Bypass cache reads (still refreshes the cache)
x-cache-tolerance Max acceptable cache age in seconds (default 43200)
x-retain-links all (default) / none / text
x-retain-images all (default) / none / alt
x-max-tokens Truncate output to N tokens

Document parsing headers (MinerU)

Header Description
x-mineru auto (default: agent → precision → MarkItDown) / agent / precision / off
x-mineru-token Per-request MinerU token (overrides MINERU_TOKEN)
x-mineru-model Precision model: vlm (default) / pipeline
x-mineru-language OCR language, e.g. ch / en
x-mineru-ocr: true Force OCR
x-page-ranges Page ranges, e.g. 1-10

Document routing

Format MinerU enabled (default) MinerU off / excluded
pdf / docx / pptx / xlsx / image MinerU agent (≤10 MB) → precision if a token is set → MarkItDown fallback MarkItDown
doc / ppt / xls (legacy Office) MinerU precision (token required) Unsupported (415)
csv / json / xml / ipynb / zip / txt / html MarkItDown MarkItDown
Never routed to MinerU — —
  • Which formats route to MinerU is controlled by the MINERU_FORMATS env var (comma-separated; default = the agent + legacy set above). Formats removed from that list fall back to MarkItDown.
  • The token-free MinerU agent API is limited to files ≤ 10 MB and ≤ 20 pages; larger files are auto-upgraded to the precision API when a token is configured.
  • Legacy .doc / .ppt / .xls require the precision API (needs a token); without one they return 415.

Rendering ladder (web pages)

  1. Direct fetch — real browser UA + content validity check (visible text after stripping script/style).
  2. Simple render — when fetched content is invalid and a renderer is configured: waitUntil: load + a short settle wait.
  3. Full CDP render (optional fallback) — networkidle wait, x-wait-for-selector, auto-scroll to trigger lazy loading. Off by default; enable via x-cdp-fallback: true or env CDP_FALLBACK=true.

Choosing a renderer (external CDP)

Default: Cloudflare Browser Rendering. You can also connect your own remote Chrome:

# Chrome side:
chrome --headless --remote-debugging-port=9222

# Worker side:
npx wrangler secret put CHROME_CDP_WS_URL   # e.g. wss://chrome.example.com:9222 or https://host:9222

In wrangler.jsonc set "CDP_PROVIDER": "external" to prefer external over Cloudflare. The external CDP client is a lean built-in implementation: Target.createTarget → attachToTarget → Page.navigate → wait (load / selector / network idle) → Runtime.evaluate (HTML) → closeTarget; http endpoints auto-discover webSocketDebuggerUrl and rewrite the host.


Configuration

Environment variables & bindings

Variable Required Description
BROWSER (binding) Recommended Cloudflare Browser Rendering
CACHE_DB (D1 binding) Recommended Result cache
MINERU_TOKEN Optional MinerU precision-API token
MINERU_FORMATS Optional Comma-separated formats routed to MinerU, e.g. pdf,docx,image; none/off disables MinerU entirely
CHROME_CDP_WS_URL Optional External Chrome CDP endpoint
CDP_PROVIDER Optional cloudflare (default) / external
CDP_FALLBACK Optional Allow full-CDP fallback by default; default false
AUTH_TOKEN Optional When set, requires Authorization: Bearer <token>
DEFAULT_TIMEOUT Optional Default timeout in seconds; default 360
MAX_UPLOAD_MB Optional Direct-upload size limit in MB; default 20
RATE_LIMIT_MAX Optional Requests allowed per IP per window; default 1; 0 disables rate limiting
RATE_LIMIT_WINDOW_SEC Optional Window in seconds; default 15

Secrets (MINERU_TOKEN, AUTH_TOKEN, CHROME_CDP_WS_URL) are set with npx wrangler secret put ... — never put them in wrangler.jsonc.

Rate limiting

Per-client-IP fixed window (CF-Connecting-IP, falling back to X-Forwarded-For), stored in D1. By default: 1 request per IP per 15 seconds. Exceeding it returns 429 with a Retry-After header.

  • / and /llms.txt are exempt.
  • Tune via wrangler.jsonc → vars (then redeploy) or the dashboard:
    • RATE_LIMIT_MAX=3 → 3 requests per 15 s
    • RATE_LIMIT_WINDOW_SEC=60 → 1 request per 60 s
    • RATE_LIMIT_MAX=0 → disabled
  • Counters live in the D1 rate_limit table (migration 0002_rate_limit.sql); the cron purges stale rows hourly.

Deployment

npm install

# 1. Create the D1 database, then paste the printed database_id into wrangler.jsonc
npx wrangler d1 create reader-cache
npx wrangler d1 migrations apply reader-cache --remote

# 2. (Optional) secrets
npx wrangler secret put MINERU_TOKEN      # MinerU precision API token
npx wrangler secret put AUTH_TOKEN        # Bearer token protecting this API
npx wrangler secret put CHROME_CDP_WS_URL # external Chrome CDP endpoint

# 3. Deploy (wrangler's build.command rebuilds the landing stylesheet first)
npm run deploy

The browser.binding = BROWSER in wrangler.jsonc provides Cloudflare Browser Rendering (remote Chrome over CDP). The free tier includes 10 minutes of browser time per day.

The landing page stylesheet is compiled from Tailwind CSS v4 (no component library). After editing src/landing/input.css or src/landing.ts:

npm run build:css   # outputs to public/landing.css, uploaded on deploy

Project Structure

reader/
├── src/
│   ├── index.ts            # Router, auth, rate limiting, pipeline orchestration
│   ├── options.ts          # Header / body option parsing & validation
│   ├── classify.ts         # Content-type & extension → kind routing
│   ├── fetcher.ts          # Direct fetch (browser UA, timeout, SSRF guard)
│   ├── webpage.ts          # Readability-style HTML → Markdown extraction
│   ├── render.ts           # CDP render abstraction (simple / full)
│   ├── cdp.ts              # Lean external-Chrome CDP client
│   ├── documents.ts        # Document parsing pipeline (MinerU / MarkItDown)
│   ├── mineru.ts           # MinerU agent / precision / signed-upload API clients
│   ├── markitdown.ts       # MarkItDown (TS) converter wrapper
│   ├── cache.ts            # D1 result cache (12h TTL) + hourly purge
│   ├── ratelimit.ts        # D1 fixed-window rate limiter
│   ├── url-input.ts        # Target URL parsing & normalization
│   ├── format.ts           # Markdown / JSON / error / llms.txt responses
│   ├── landing.ts          # Anthropic-style landing page
│   ├── landing/input.css   # Tailwind v4 design tokens & components
│   └── types.ts            # Shared types
├── migrations/             # D1 schema migrations
├── test/                   # Vitest unit tests
└── wrangler.jsonc          # Worker configuration

Development

npm test          # vitest unit tests
npm run typecheck # tsc --noEmit
npm run dev       # wrangler dev (no browser rendering locally)
npx wrangler d1 migrations apply reader-cache --local

Known Limitations

  • The MinerU lightweight API is rate-limited per exit IP; high concurrency may hit 429. It can also time out fetching sites such as GitHub/AWS — the service then retries with "Worker downloads + signed upload".
  • The Worker free tier has a 10 ms CPU limit: large zip decompression / big document conversions are better on a paid plan (raise limits.cpu_ms there too).
  • SSRF protection is a blacklist of hostname/IP literals (private / loopback / metadata). DNS-rebinding protection is not performed.
  • A single cache entry is capped at 8 MB; larger results are not cached.

License

MIT © 2026 BloomsReader Contributors

About

Jina Reader 风格的 Markdown 转换 Cloudflare Worker

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages