Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
37 changes: 35 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,9 +58,21 @@ py -3.13 -m venv .venv

2. Install Python dependencies (this installs PyTorch as well):
```bash
pip install -r requirements.txt
python -m pip install -r requirements.txt
```

Keep the virtual environment active when running `python main.py`. If you see
`ModuleNotFoundError: No module named 'latex2mathml'`, your Python may be using
another environment. On Linux/Mac, run with the project interpreter directly:

```bash
.venv/bin/python main.py ~/scan-pages/Book1.pdf
```

If that interpreter also reports a missing dependency, install the full
requirements into `.venv`. For environments created with uv, use
`uv pip install --python .venv/bin/python -r requirements.txt`.

3. GPU acceleration (optional):

PyTorch is installed as a dependency in step 2. On Apple Silicon that wheel
Expand Down Expand Up @@ -132,7 +144,8 @@ python main.py [input_path] [output_path] [options]

Options:
--max-pages INT Maximum number of pages to process
--start-page INT Page number to start from
--start-page INT Zero-based page index to start from
--chunk-pages INT Pages processed at once (default: 10; 0 disables chunking)
--skip-epub Skip EPUB generation, only create markdown
--skip-md Skip markdown generation, use existing markdown files
```
Expand All @@ -151,6 +164,26 @@ Convert to markdown only:
python main.py thesis.pdf --skip-epub
```

### GPU memory usage

PDFs are processed in 10-page chunks by default, then joined into one Markdown
file and one EPUB. This bounds the page-image memory used by Marker; users do
not need to split or reassemble the PDF themselves. On an 8 GB GPU, the default
also sets all of Marker's layout, detection, OCR, recognition, table, and
equation batch sizes to 1.

If a particularly image-heavy PDF still runs out of memory, use a smaller
chunk (for example `--chunk-pages 4`). `PDF2EPUB_GPU_BATCH_SIZE` can raise the
model batch sizes on GPUs with more memory. `DETECTOR_BATCH_SIZE` alone does not
control Marker's internal detection batch in the pinned Marker version.

PyTorch's allocator variable is named `PYTORCH_CUDA_ALLOC_CONF` (including
`CUDA`), for example:

```bash
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True python main.py book.pdf
```

### Output Structure

```
Expand Down
14 changes: 12 additions & 2 deletions main.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,16 @@ def main():
'--start-page',
type=int,
default=None,
help='Page number to start from'
help='Zero-based page index to start from'
)
parser.add_argument(
'--chunk-pages',
type=int,
default=pdf2md.DEFAULT_CHUNK_PAGES,
help=(
'Pages processed at once (default: 10; lower this if GPU memory is '
'still exhausted, or use 0 to disable chunking)'
)
)
parser.add_argument(
'--skip-epub',
Expand Down Expand Up @@ -93,6 +102,7 @@ def main():
markdown_dir,
args.max_pages,
args.start_page,
args.chunk_pages,
)

# Convert Markdown to EPUB unless skipped
Expand All @@ -115,4 +125,4 @@ def main():
sys.exit(1)

if __name__ == '__main__':
main()
main()
Loading