I tried this two times with a big PDF (500 pages), every step goes on succesfully but at "text recognition" I get a "killed" message after a little while:
Processing: my book.pdf
Converting PDF to Markdown...
Recognizing Layout: 100%|██████████████████████████████████████████████████████████| 526/526 [3:19:16<00:00, 22.73s/it]
Running OCR Error Detection: 100%|███████████████████████████████████████████████████| 132/132 [00:06<00:00, 21.95it/s]
Detecting bboxes: 100%|██████████████████████████████████████████████████████████████| 132/132 [27:41<00:00, 12.59s/it]
Recognizing Text: 0%| | 0/41327 [00:00<?, ?it/s]Killed
The PDF has a layout with columns and plenty of images, it's kinda a school book.
Is this common? Is there some log to check out which I can paste here?
I'm on Linux Debian 13, AMD Ryzen CPU
I tried this two times with a big PDF (500 pages), every step goes on succesfully but at "text recognition" I get a "killed" message after a little while:
The PDF has a layout with columns and plenty of images, it's kinda a school book.
Is this common? Is there some log to check out which I can paste here?
I'm on Linux Debian 13, AMD Ryzen CPU