A small PDF → LLM-ready text pipeline with a token-savings audit report.
Quantagen extracts text from research PDFs, filters out layout noise (headers, footers, page numbers, boilerplate), truncates reference/appendix sections, and reports how many tokens you saved — with a per-stage breakdown of where the savings come from. It ships as both a CLI batch runner and a Streamlit web UI, and both run the same pipeline.
Status: early-stage / experimental. This is a working proof-of-concept (~250 lines of application code), not a production tool. The numbers it reports are honest and reproducible, but there is no automated fidelity evaluation yet — see Known limitations.
PDF (PyMuPDF)
│ raw text, per page
▼
trim_structural_edges() ← config-driven % trim of document head/tail (main.py)
│
▼
filter_noise() ← Unicode-aware line-level filter (main.py)
│ drops page numbers, symbol-dense lines; keeps non-Latin text
▼
truncate_at_stop_markers() ← cuts at standalone REFERENCES / APPENDIX / INDEX headings
│ (rules/parser.py; prose like "index r." can never fire it)
▼
strip_recurring_noise() ← config-driven regex cleanup from config.json
│
├──► cleaned .txt written to outputs/
├──► token counts (tiktoken, gpt-4o encoding) + per-stage audit report
└──► abstract/conclusion extraction (printed in the CLI audit)
Both main.py (CLI) and app.py (UI) call the same clean_pipeline(), so their outputs are byte-identical for the same input.
- Recursive batch CLI — walks
inputs/for PDFs, processes each, prints a per-document audit report. - Streamlit UI (
app.py) — multi-file upload, one tab per document with raw/cleaned token metrics, a 600px preview pane, per-file.txtdownload, and a batch summary table. - Token accounting with stage breakdown — raw vs. cleaned token counts via
tiktoken, the reduction percentage, and which stage removed how many tokens (structural trim / line filter / truncation / patterns). Savings are attributable, not just asserted. - Unicode-aware filtering — letters from any script (Japanese, Cyrillic, Arabic, …) count as content; only genuinely symbol-dense lines are dropped.
- Heading-safe truncation — stop markers must start their own line and be heading-shaped, so a sentence containing the word "index" cannot truncate a document mid-prose.
- Config-driven rules (
config.json) — noise regexes, stop markers, and abstract/conclusion section markers live outside the code.config.jsonis located relative to the package, so you can run from any directory.
quantagen/
├── main.py # CLI entrypoint, clean_pipeline(), filter_noise(), token/audit helpers
├── app.py # Streamlit UI (uses the same clean_pipeline())
├── rules/parser.py # config-driven truncation, noise stripping, section extraction
├── config.json # structural rules, noise patterns, stop markers, section markers
├── inputs/ # sample PDFs (a Springer textbook + a research paper)
├── outputs/ # cleaned .txt files (gitignored; generated locally)
├── requirements.txt # dependencies
└── LICENSE # Apache-2.0
pip install -r requirements.txt
tiktokendownloads its tokenizer on first use, so the first run needs network access.
python main.pyRecursively scans inputs/ for *.pdf, writes cleaned files to outputs/, and prints an audit report per document. Real numbers from the bundled samples:
--- Audit Report: MLBasicsBook.pdf ---
Raw Tokens: 155055
Cleaned Tokens: 130111
Efficiency: 16.0872% (24944 tokens removed)
- structural trim 17588 tokens
- line filter 7356 tokens
Sections: abstract (1421 chars), conclusion (3352 chars)
--- Audit Report: researchpaperOnSustainableDev.pdf ---
Raw Tokens: 8540
Cleaned Tokens: 3758
Efficiency: 55.9953% (4782 tokens removed)
- structural trim 864 tokens
- line filter 268 tokens
- stop-marker truncation 3650 tokens
Sections: conclusion (168 chars)
Note how different document types behave honestly differently: the textbook keeps most of its body (its savings come from front matter and layout noise), while the research paper legitimately drops ~50% to its references section. An earlier version of this tool reported ~53% for both documents — that number was wrong (see Fixed issues).
streamlit run app.pyUpload one or more PDFs, inspect the cleaned text per tab, and download the results. The UI runs the identical pipeline as the CLI.
All tunables live in config.json:
| Key | Purpose |
|---|---|
structural_rules |
Fraction of the top/bottom of the document to strip (0–1). Set to 0 to disable. |
global_noise_patterns |
Regexes for recurring boilerplate (ISSN lines, © lines, Page N of M, stamp lines like DRAFT / CONFIDENTIAL) |
stop_markers |
Truncate at the first standalone heading line matching one of these (REFERENCES, APPENDIX, INDEX…). Must start the line; only a colon or section letter/number may follow. |
section_markers |
Headings used to extract abstract and conclusion |
- No automated fidelity evaluation. The pipeline removes what its heuristics consider noise, but nothing yet measures whether answers computed from the cleaned text match answers from the raw text. Until a QA-retention or precision/recall harness exists, treat the output as unvalidated.
structural_rulesis a blunt instrument. Trimming 5% off each end is right for documents with cover pages/colophons and wrong for documents without them. It is off-by-default-able (set to0) but not content-aware.- Section extraction is heuristic. In books with per-chapter "Summary" headings, the extracted "conclusion" is the first such section's text, not the book's conclusion. Fine for papers, imperfect for books.
- Detection is English-centric. Non-Latin text now survives cleaning (fixed — see below), but the stop markers, section markers, and noise patterns themselves target English/journal conventions.
- Paragraph structure is flattened — blank lines are removed during line filtering.
- No tests or CI yet. The behaviors described here are verified manually against the bundled samples; nothing guards regressions automatically.
process_document()still catches broad exceptions (it now prints the traceback).
For the record — these shipped broken and were fixed after being reproduced against the bundled samples (see git history for the individual commits):
- Noise patterns never ran:
strip_recurring_noise()looked up the wrong config key (noise_patternsvsglobal_noise_patterns), lackedre.MULTILINEfor^…$-anchored patterns, and loaded the config without UTF-8 (breaking the©pattern on Windows). Thedraft/confidentialpattern also matched the word "drafts" in prose; it is now anchored to stamp-shaped lines. - Stop-marker truncation cut documents mid-prose: matching was case-insensitive with no line anchoring, so the word "index" inside a sentence truncated the sample book at 51%, silently discarding its final chapters. Truncation now requires a standalone heading-shaped line, and the audit shows truncation's contribution separately.
- Non-Latin text was deleted wholesale: the density filter counted only
[a-zA-Z\d\s]as content, so Japanese/Cyrillic/Arabic lines scored as ~100% symbols. The filter is now Unicode-aware while still dropping symbol-dense layout lines. - CLI and UI ran different pipelines: the UI skipped
filter_noise(). Both now share oneclean_pipeline(). structural_ruleswas dead config and extracted sections were computed but discarded: both are wired up and surfaced in the CLI audit.
- Fidelity evaluation harness — hand-labeled noise lines for precision/recall of the filter, and/or QA retention (answer questions against raw vs. cleaned text) to report "token reduction @ ≥95% answer accuracy".
- Loss ledger — log every removed span with a reason and token count (
report.json) instead of destructive deletion, making the cleaning auditable and reversible. - Corpus-adaptive boilerplate detection — learn recurring header/footer lines from document frequency across a corpus instead of hand-written regexes.
- Markdown re-emission — structure headings/lists/tables before counting, cutting tokens further and improving downstream RAG retrieval.
- Packaging —
pyproject.toml,pip install quantagen, a properquantagenCLI entry point, and a pytest suite with fixture PDFs.