A Jina Reader (r.jina.ai)–compatible Cloudflare Worker that turns any web page or document into clean, LLM-friendly Markdown — from a URL or a direct file upload, with one command.
| Web pages | Direct fetch → simple render → full CDP render (optional), then readability extraction |
| Documents | PDF / DOCX / PPTX / XLSX / images & legacy Office (.doc .ppt .xls) via MinerU; MarkItDown fallback |
| Cache | D1-backed, 12-hour TTL, hourly cron cleanup |
| Rate limit | Per-IP fixed window, default 1 request / 15 s (configurable) |
- Familiar API — Drop-in-style routes and headers modeled after Jina Reader; works with existing tools that target
r.jina.ai. - Great for LLMs — Readability extraction + MarkItDown produce clean, structured Markdown; no boilerplate or tracking scripts.
- Documents made easy — PDFs, Office files, images and more parsed into text, tables and structure via MinerU; graceful MarkItDown fallback if MinerU is unavailable.
- Costs almost nothing — Serverless, pay-as-you-go. Browser rendering only when needed; document parsing can run on the token-free MinerU lightweight API.
- Private by design — Per-IP rate limiting; optional
AUTH_TOKEN; SSRF blacklist for private / loopback / metadata addresses. - Built for the edge — Runs on Cloudflare Workers with a D1 cache; Tiny, typed, tested.
# Web page → Markdown (Jina style: URL directly appended)
curl https://<your-worker>/https://example.com
# JSON mode (same shape as Jina: {code, status, data:{title, url, content, usage...}})
curl -H "Accept: application/json" https://<your-worker>/https://example.com
# Document → Markdown (token-free MinerU lightweight API)
curl https://<your-worker>/https://cdn-mineru.openxlab.org.cn/demo/example.pdf
# Direct file upload (multipart)
curl -F "file=@report.docx" https://<your-worker>/
# POST with options
curl -X POST -H "Content-Type: application/json" \
-d '{"url":"https://example.com","timeout":90}' https://<your-worker>/Live demo: https://<your-worker>/https://example.com
flowchart TD
A[Request] --> B{Auth<br/><i>optional</i>}
B -- "ok" --> C[Rate limit<br/><i>per IP</i>]
C -- "allowed" --> D[Classify content]
D --> E[Web page]
D --> F[Document<br/>PDF / Office / image]
E --> G{Direct fetch<br/>content valid?}
G -- "yes" --> H[Readability extraction]
G -- "no" --> I[Render]
I --> J[Simple render]
I --> K[Full CDP render]
J --> H
K --> H
F --> L[MinerU agent<br/><i>≤ 10 MB</i>]
L -- "success" --> M[Clean Markdown<br/>title · url · content]
L -- "fail" --> P{Precision API<br/>token set?}
P -- "yes" --> M
P -- "no" --> Q[MarkItDown fallback]
Q --> M
H --> M
M --> O["12 h D1 cache<br/><i>x-no-cache to bypass</i>"]
| Endpoint | Method | Description |
|---|---|---|
GET /{url} |
GET | Web page or document → Markdown (Jina style). Encode the target as part of the path. |
GET /?url={url} |
GET | Same, with the target as a query parameter. |
POST / |
POST | {"url": "..."} (JSON) or multipart form with a file field. |
GET /llms.txt |
GET | Machine-readable API documentation (served to LLMs). |
GET / |
GET | Human-friendly landing page (never rate-limited). |
Default (text/markdown):
Title: Example Domain
URL Source: https://example.com/
Markdown Content:
...
x-respond-with |
Output |
|---|---|
| (default) | Markdown + title block; cache hits append Warning: This is a cached snapshot... |
markdown |
Full-page Markdown (skips readability extraction) |
html |
Page HTML |
text |
Plain text |
frontmatter / markdown+frontmatter |
YAML frontmatter + Markdown |
json |
Equivalent to Accept: application/json |
| Header | Description |
|---|---|
Accept: application/json |
JSON response mode |
x-respond-with |
Output mode (see above) |
x-timeout |
Total timeout in seconds, 5–360 (default 360) |
x-engine |
auto (default) / curl (direct fetch only) / browser (render first) |
x-target-selector |
CSS selector; return only matching content |
x-wait-for-selector |
CSS selector to wait for when rendering |
x-remove-selector |
Comma-separated CSS selectors to strip from the result |
x-set-cookie |
Cookie forwarded to the target (such requests are not cached) |
x-no-cache: true |
Bypass cache reads (still refreshes the cache) |
x-cache-tolerance |
Max acceptable cache age in seconds (default 43200) |
x-retain-links |
all (default) / none / text |
x-retain-images |
all (default) / none / alt |
x-max-tokens |
Truncate output to N tokens |
| Header | Description |
|---|---|
x-mineru |
auto (default: agent → precision → MarkItDown) / agent / precision / off |
x-mineru-token |
Per-request MinerU token (overrides MINERU_TOKEN) |
x-mineru-model |
Precision model: vlm (default) / pipeline |
x-mineru-language |
OCR language, e.g. ch / en |
x-mineru-ocr: true |
Force OCR |
x-page-ranges |
Page ranges, e.g. 1-10 |
| Format | MinerU enabled (default) | MinerU off / excluded |
|---|---|---|
| pdf / docx / pptx / xlsx / image | MinerU agent (≤10 MB) → precision if a token is set → MarkItDown fallback | MarkItDown |
| doc / ppt / xls (legacy Office) | MinerU precision (token required) | Unsupported (415) |
| csv / json / xml / ipynb / zip / txt / html | MarkItDown | MarkItDown |
| Never routed to MinerU | — | — |
- Which formats route to MinerU is controlled by the
MINERU_FORMATSenv var (comma-separated; default = the agent + legacy set above). Formats removed from that list fall back to MarkItDown. - The token-free MinerU agent API is limited to files ≤ 10 MB and ≤ 20 pages; larger files are auto-upgraded to the precision API when a token is configured.
- Legacy
.doc/.ppt/.xlsrequire the precision API (needs a token); without one they return415.
- Direct fetch — real browser UA + content validity check (visible text after stripping script/style).
- Simple render — when fetched content is invalid and a renderer is configured:
waitUntil: load+ a short settle wait. - Full CDP render (optional fallback) —
networkidlewait,x-wait-for-selector, auto-scroll to trigger lazy loading. Off by default; enable viax-cdp-fallback: trueor envCDP_FALLBACK=true.
Default: Cloudflare Browser Rendering. You can also connect your own remote Chrome:
# Chrome side:
chrome --headless --remote-debugging-port=9222
# Worker side:
npx wrangler secret put CHROME_CDP_WS_URL # e.g. wss://chrome.example.com:9222 or https://host:9222In wrangler.jsonc set "CDP_PROVIDER": "external" to prefer external over Cloudflare. The external CDP client is a lean built-in implementation: Target.createTarget → attachToTarget → Page.navigate → wait (load / selector / network idle) → Runtime.evaluate (HTML) → closeTarget; http endpoints auto-discover webSocketDebuggerUrl and rewrite the host.
| Variable | Required | Description |
|---|---|---|
BROWSER (binding) |
Recommended | Cloudflare Browser Rendering |
CACHE_DB (D1 binding) |
Recommended | Result cache |
MINERU_TOKEN |
Optional | MinerU precision-API token |
MINERU_FORMATS |
Optional | Comma-separated formats routed to MinerU, e.g. pdf,docx,image; none/off disables MinerU entirely |
CHROME_CDP_WS_URL |
Optional | External Chrome CDP endpoint |
CDP_PROVIDER |
Optional | cloudflare (default) / external |
CDP_FALLBACK |
Optional | Allow full-CDP fallback by default; default false |
AUTH_TOKEN |
Optional | When set, requires Authorization: Bearer <token> |
DEFAULT_TIMEOUT |
Optional | Default timeout in seconds; default 360 |
MAX_UPLOAD_MB |
Optional | Direct-upload size limit in MB; default 20 |
RATE_LIMIT_MAX |
Optional | Requests allowed per IP per window; default 1; 0 disables rate limiting |
RATE_LIMIT_WINDOW_SEC |
Optional | Window in seconds; default 15 |
Secrets (
MINERU_TOKEN,AUTH_TOKEN,CHROME_CDP_WS_URL) are set withnpx wrangler secret put ...— never put them inwrangler.jsonc.
Per-client-IP fixed window (CF-Connecting-IP, falling back to X-Forwarded-For), stored in D1. By default: 1 request per IP per 15 seconds. Exceeding it returns 429 with a Retry-After header.
/and/llms.txtare exempt.- Tune via
wrangler.jsonc→vars(then redeploy) or the dashboard:RATE_LIMIT_MAX=3→ 3 requests per 15 sRATE_LIMIT_WINDOW_SEC=60→ 1 request per 60 sRATE_LIMIT_MAX=0→ disabled
- Counters live in the D1
rate_limittable (migration0002_rate_limit.sql); the cron purges stale rows hourly.
npm install
# 1. Create the D1 database, then paste the printed database_id into wrangler.jsonc
npx wrangler d1 create reader-cache
npx wrangler d1 migrations apply reader-cache --remote
# 2. (Optional) secrets
npx wrangler secret put MINERU_TOKEN # MinerU precision API token
npx wrangler secret put AUTH_TOKEN # Bearer token protecting this API
npx wrangler secret put CHROME_CDP_WS_URL # external Chrome CDP endpoint
# 3. Deploy (wrangler's build.command rebuilds the landing stylesheet first)
npm run deployThe browser.binding = BROWSER in wrangler.jsonc provides Cloudflare Browser Rendering (remote Chrome over CDP). The free tier includes 10 minutes of browser time per day.
The landing page stylesheet is compiled from Tailwind CSS v4 (no component library). After editing src/landing/input.css or src/landing.ts:
npm run build:css # outputs to public/landing.css, uploaded on deployreader/
├── src/
│ ├── index.ts # Router, auth, rate limiting, pipeline orchestration
│ ├── options.ts # Header / body option parsing & validation
│ ├── classify.ts # Content-type & extension → kind routing
│ ├── fetcher.ts # Direct fetch (browser UA, timeout, SSRF guard)
│ ├── webpage.ts # Readability-style HTML → Markdown extraction
│ ├── render.ts # CDP render abstraction (simple / full)
│ ├── cdp.ts # Lean external-Chrome CDP client
│ ├── documents.ts # Document parsing pipeline (MinerU / MarkItDown)
│ ├── mineru.ts # MinerU agent / precision / signed-upload API clients
│ ├── markitdown.ts # MarkItDown (TS) converter wrapper
│ ├── cache.ts # D1 result cache (12h TTL) + hourly purge
│ ├── ratelimit.ts # D1 fixed-window rate limiter
│ ├── url-input.ts # Target URL parsing & normalization
│ ├── format.ts # Markdown / JSON / error / llms.txt responses
│ ├── landing.ts # Anthropic-style landing page
│ ├── landing/input.css # Tailwind v4 design tokens & components
│ └── types.ts # Shared types
├── migrations/ # D1 schema migrations
├── test/ # Vitest unit tests
└── wrangler.jsonc # Worker configuration
npm test # vitest unit tests
npm run typecheck # tsc --noEmit
npm run dev # wrangler dev (no browser rendering locally)
npx wrangler d1 migrations apply reader-cache --local- The MinerU lightweight API is rate-limited per exit IP; high concurrency may hit
429. It can also time out fetching sites such as GitHub/AWS — the service then retries with "Worker downloads + signed upload". - The Worker free tier has a 10 ms CPU limit: large zip decompression / big document conversions are better on a paid plan (raise
limits.cpu_msthere too). - SSRF protection is a blacklist of hostname/IP literals (private / loopback / metadata). DNS-rebinding protection is not performed.
- A single cache entry is capped at 8 MB; larger results are not cached.
MIT © 2026 BloomsReader Contributors