Repository navigation
Process PDFs in page chunks to bound Marker memory - #43
Open
SkippySteve wants to merge 1 commit into
Open
SkippySteve wants to merge 1 commit into
SkippySteve wants to merge 1 commit into
Conversation
Marker holds the selected page range in memory, which OOMs on large or image-heavy PDFs. Convert in 10-page chunks by default, reuse the loaded models, join Markdown and metadata, and free each chunk before the next. --start-page now works without --max-pages by reading the PDF page count. Marker batch sizes default to 1 (PDF2EPUB_GPU_BATCH_SIZE). --chunk-pages lowers the window further, or 0 disables chunking. Document the new flags, GPU memory knobs, and the latex2mathml import error that appears when main.py is run outside the project venv. Fixes overcuriousity#36
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Marker converts the whole selected page range in one pass, so large or image-heavy PDFs can be killed by the OOM killer during text recognition. This change converts in 10-page chunks by default, reuses the loaded Marker models, writes one joined Markdown file plus metadata, and frees each chunk's page images and CUDA cache before the next chunk.
--start-pagenow works without--max-pagesby reading the PDF page count first. Marker layout/detection/OCR/recognition/table/equation batch sizes default to 1 (override withPDF2EPUB_GPU_BATCH_SIZE).--chunk-pageslowers the page window further, or0disables chunking.The README documents the new flags, GPU memory knobs, and a common
latex2mathmlModuleNotFoundErrorwhen the interpreter that runsmain.pyis not the project venv.Fixes #36
Test plan
python -m compileall -q main.py modules/ruff check --select E9,F ._chunks,_selected_pages,_merge_metadata, and_low_memory_config