Repository navigation
Deliverable 3 pilot findings: 234 issues across 30 documents — prioritized improvement plan #82
Description
Activity
Update: Hyperlink Extraction — Docling Status & pypdfium2 Fallback Plan
Docling Status (as of March 2026)
The
DoclingDocumentschema does have aTextItem.hyperlink: Optional[Union[AnyUrl, Path]]field (docling-core source), and the low-level parser (docling-parse v5) extracts PDF link annotations. However, the propagation step — matching hyperlink bounding boxes to text items — was never implemented for the PDF pipeline.- Issue #828 (open since Jan 2025): Original report. Maintainer confirmed hyperlinks are identified in docling-parse but need propagation to DoclingDocument.
- PR #3131: External contributor added spatial matching of hyperlink annotations to text clusters. CI passes, one maintainer approved, but it's stuck awaiting a second reviewer and is not merged.
Result: Every
TextItemin our Docling JSON response hashyperlink=nullfor PDF documents. URLs that appear as literal text in the PDF body are preserved, but clickable hyperlinks where display text (e.g., "click here") links to a URL are lost entirely.pypdfium2 Fallback Implementation Plan
Since we already use
pypdfium2.rawinpdf_classifier.pyfor annotation inspection (form field widget counting), we can use the same API pattern to extract hyperlinks. Here's the plan:Step 1: Extract link annotations from PDF bytes
Add a new function
extract_hyperlinks()to a new module (e.g.,src/services/hyperlink_extractor.py) or extendpdf_classifier.py:import pypdfium2 import pypdfium2.raw as pdfium_c import ctypes @dataclass class PdfHyperlink: page_number: int # 1-indexed url: str # Target URI bbox: tuple[float, float, float, float] # (left, bottom, right, top) in PDF coords def extract_hyperlinks(file_content: bytes) -> list[PdfHyperlink]: """Extract hyperlink annotations from PDF using pypdfium2.""" pdf = pypdfium2.PdfDocument(file_content) links: list[PdfHyperlink] = [] for page_idx in range(len(pdf)): page = pdf.get_page(page_idx) try: # Enumerate link annotations on the page start_pos = ctypes.c_int(0) link_ptr = ctypes.c_void_p() while pdfium_c.FPDFLink_Enumerate(page.raw, ctypes.byref(start_pos), ctypes.byref(link_ptr)): # Get link bounding rectangle rect = pdfium_c.FS_RECTF() if not pdfium_c.FPDFLink_GetAnnotRect(link_ptr, ctypes.byref(rect)): continue # Get action and check if it's a URI action (type 1) action = pdfium_c.FPDFLink_GetAction(link_ptr) if not action: continue if pdfium_c.FPDFAction_GetType(action) != 1: # PDFACTION_URI continue # Extract the URI string buf_size = pdfium_c.FPDFAction_GetURIPath(pdf.raw, action, None, 0) if buf_size <= 0: continue buf = ctypes.create_string_buffer(buf_size) pdfium_c.FPDFAction_GetURIPath(pdf.raw, action, buf, buf_size) url = buf.value.decode("utf-8", errors="replace") links.append(PdfHyperlink( page_number=page_idx + 1, url=url, bbox=(rect.left, rect.bottom, rect.right, rect.top), )) finally: page.close() pdf.close() return links
This follows the exact same pattern as
_extract_metadata()inpdf_classifier.py(lines 278-302), which already usespdfium_c.FPDFPage_GetAnnotCount,FPDFPage_GetAnnot, andFPDFAnnot_GetSubtype.Step 2: Spatial-join hyperlinks to Docling text
The Docling JSON response includes bounding boxes for text items (
body→ items →prov→bbox). We can match hyperlink rectangles to the text they cover:- For each
PdfHyperlink, find text items on the same page whose bounding box overlaps with the link rectangle - The overlapping text becomes the display text for the markdown link
- Inject
[display text](url)into the per-page markdown at the corresponding position
This spatial join would happen in
_step_docling()(pipeline_viewer.py~line 700), right after extracting figures and before the pipeline proceeds to AI processing steps.Step 3: Inject into per-page markdown
After identifying which text spans are linked, modify the per-page markdown (v0) to wrap linked text in markdown link syntax. This runs as a deterministic post-processing step — no LLM needed.
Step 4: Future-proof for Docling native support
Structure the code so that when PR #3131 merges and
TextItem.hyperlinkbecomes populated:- Check
json_contentfor populatedhyperlinkfields first - Fall back to pypdfium2 extraction only if Docling doesn't provide them
- Eventually remove the pypdfium2 fallback entirely
def get_hyperlinks(file_content: bytes, json_content: dict) -> list[PdfHyperlink]: """Get hyperlinks from Docling JSON if available, else fall back to pypdfium2.""" docling_links = extract_hyperlinks_from_docling_json(json_content) if docling_links: return docling_links return extract_hyperlinks_from_pypdfium2(file_content)
Where this fits in the pipeline
_step_docling(): 1. Docling extraction + pypdfium2 page rendering (parallel) — existing 2. Split markdown by page — existing 3. Extract figures from JSON + crop — existing 4. Replace image placeholders — existing 5. ✨ NEW: Extract hyperlinks (pypdfium2) + spatial-join + inject into markdown 6. Detect column layout — existingThe hyperlink injection happens once, deterministically, at the end of
_step_docling(). All subsequent pipeline steps (structure analysis, page correction, boundary fixes, cleanup) will see the hyperlinks already in the markdown and preserve them.- added a commit that references this issue
on Apr 22, 2026
Overview
Deliverable 3 pilot tested 30 real-world UIC documents (174 pages, $13.45 total cost). Manual quality review found 234 issues (45 critical, 101 major, 88 minor). This issue organizes findings by topic, maps them to codebase locations, and proposes changes in priority order.
Summary stats: 30 docs, 174 pages, avg 2m44s conversion, avg $0.45/doc, 1,597 edits applied.
Issue Categories by Priority
1. 🔴 Footnote & Endnote Handling (~30 issues, 6+ documents)
The single most impactful improvement for academic content.
Current code:
src/agents/prompts/footnote_relocation.py— fuzzy matching via_find_body_context()and_find_marker_context()struggles with multi-column academic chapters and dense endnote sections (40+ footnotes).Proposed changes:
build_footnote_user_message()to include page images for visual verification of footnote content2. 🔴 Hyperlink Extraction (~12 issues, 8+ documents)
Systematic gap — no dedicated hyperlink handling exists.
Current code: No hyperlink extraction.
src/utils/text_cleanup.pyhasfix_url_formatting()(protocol fix only) andvalidate_urls()(logging only). Docling does not populate hyperlinks from PDFs (confirmed via GitHub discussions #771 and #2337).Proposed changes:
pypdfium2.raw(FPDFLink_Enumerate+FPDFAction_GetURIPath) — the project already uses this API pattern inpdf_classifier.pyFPDFLink_LoadWebLinks)fix_url_formatting()to fix escaped underscores (\_→_) and&→&in URLs3. 🔴 Low-Quality Scan Detection & Gating (~20 issues, 3 documents)
La Opinion Latina alone accounts for 8 critical issues from catastrophic OCR.
Current code:
src/services/pdf_classifier.pyalready has:FINDING_SCANNEDwhenchars_per_page < 50,FINDING_LOW_TEXT_DENSITYwhen< 200)enrich_classification()adds warnings but doesn't blockProposed changes:
FINDING_LOW_OCR_CONFIDENCE— after Docling extraction, measure OCR quality signals (garbled character ratios, average word length anomalies, non-dictionary word frequency)4. 🟡 Figures & Alt Text Quality (~35 issues, 15+ documents)
Proposed changes:
IMAGE_DESCRIBER_SYSTEM_PROMPTinimage_description.py: add explicit instructions for text-heavy images (transcribe visible text verbatim), QR codes, social media icons_replace_image_placeholders()— mismatch between<!-- image -->placeholder count and extracted figure count5. 🟡 Heading Hierarchy & Missing Titles (~15 issues, 10+ documents)
Proposed changes:
structure_analysis.py— always identify document title even when embedded in image/banner6. 🟡 Table Formatting (~10 issues, 5+ documents)
Proposed changes:
table_reconstruction.py: add ToC detection — output as ordered lists, not tables7. 🟡 Multi-Column Reading Order (~15 issues, 5+ documents)
Hardest problem. Scanned academic chapters are marked out-of-scope for phase 1.
Proposed changes:
double_column.mdprompt with more explicit column-ordering instructions8. 🟠 Presentation & Poster Layouts (~12 issues, 4 documents)
Proposed changes:
procedures/page_correction/layout/presentation.md(slide separators, repeated header handling)procedures/page_correction/layout/poster.md(zone groupings, visual hierarchy)9. 🟢 Minor Formatting (~20 issues)
HTML entities (
&), missing emphasis, page number artifacts, running header remnants.Proposed changes:
text_cleanup.pyQuality by Document Type (for reference)
Best/Worst Conversions
Best: State Infrastructure (3 issues, 0 critical), CEDA Utility (4 issues), RELS 225 (4 issues, all minor)
Worst: Latino Cultural Center (20 issues, 7 critical), La Opinion Latina (13 issues, 8 critical), Future of Work (13 issues, 3 critical)
Pilot Details