Skip to content

Process PDFs in page chunks to bound Marker memory - #43

Open
SkippySteve wants to merge 1 commit into
overcuriousity:mainfrom
SkippySteve:fix/chunk-pages-memory
Open

SkippySteve wants to merge 1 commit into
overcuriousity:mainfrom
SkippySteve:fix/chunk-pages-memory

Conversation

@SkippySteve

Copy link
Copy Markdown

Summary

Marker converts the whole selected page range in one pass, so large or image-heavy PDFs can be killed by the OOM killer during text recognition. This change converts in 10-page chunks by default, reuses the loaded Marker models, writes one joined Markdown file plus metadata, and frees each chunk's page images and CUDA cache before the next chunk.

--start-page now works without --max-pages by reading the PDF page count first. Marker layout/detection/OCR/recognition/table/equation batch sizes default to 1 (override with PDF2EPUB_GPU_BATCH_SIZE). --chunk-pages lowers the page window further, or 0 disables chunking.

The README documents the new flags, GPU memory knobs, and a common latex2mathml ModuleNotFoundError when the interpreter that runs main.py is not the project venv.

Fixes #36

Test plan

  • python -m compileall -q main.py modules/
  • ruff check --select E9,F .
  • Unit-tested _chunks, _selected_pages, _merge_metadata, and _low_memory_config

Marker holds the selected page range in memory, which OOMs on large or
image-heavy PDFs. Convert in 10-page chunks by default, reuse the loaded
models, join Markdown and metadata, and free each chunk before the next.

--start-page now works without --max-pages by reading the PDF page count.
Marker batch sizes default to 1 (PDF2EPUB_GPU_BATCH_SIZE). --chunk-pages
lowers the window further, or 0 disables chunking.

Document the new flags, GPU memory knobs, and the latex2mathml import
error that appears when main.py is run outside the project venv.

Fixes overcuriousity#36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Process killed at text recognition step

1 participant