Skip to content

Deliverable 3 pilot findings: 234 issues across 30 documents — prioritized improvement plan #82

Description

@dylan-isaac

Overview

Deliverable 3 pilot tested 30 real-world UIC documents (174 pages, $13.45 total cost). Manual quality review found 234 issues (45 critical, 101 major, 88 minor). This issue organizes findings by topic, maps them to codebase locations, and proposes changes in priority order.

Summary stats: 30 docs, 174 pages, avg 2m44s conversion, avg $0.45/doc, 1,597 edits applied.


Issue Categories by Priority

1. 🔴 Footnote & Endnote Handling (~30 issues, 6+ documents)

The single most impactful improvement for academic content.

Pattern Severity Affected Documents
Missing footnotes (gaps in numbering) Critical Future of Work, Survival Migration, Who Is a Refugee
Duplicated footnote content Major Survival Migration, Future of Work
Swapped/out-of-order footnotes Major Future of Work (6 footnote pairs)
Footnotes contain body text instead of citations Critical Who Is a Refugee, Survival Migration
Inconsistent notation styles Major Boxing and Masculinity

Current code: src/agents/prompts/footnote_relocation.py — fuzzy matching via _find_body_context() and _find_marker_context() struggles with multi-column academic chapters and dense endnote sections (40+ footnotes).

Proposed changes:

  • Improve build_footnote_user_message() to include page images for visual verification of footnote content
  • Strengthen structure analysis footnote body extraction for multi-column layouts
  • Add post-processing validation: count footnote markers in body vs definitions in Notes section, flag mismatches
  • Consider deterministic footnote reconciliation pass

2. 🔴 Hyperlink Extraction (~12 issues, 8+ documents)

Systematic gap — no dedicated hyperlink handling exists.

Pattern Severity Affected Documents
URLs rendered as plain text Major BIOS 343, State Infrastructure, ICMJE, Patient Complaint, Graduate Rates
Email addresses not formatted Minor Advising Millennial, Graduate Rates
Broken URLs (escaped underscores, HTML entities) Major BIOS 343, Graduate Rates

Current code: No hyperlink extraction. src/utils/text_cleanup.py has fix_url_formatting() (protocol fix only) and validate_urls() (logging only). Docling does not populate hyperlinks from PDFs (confirmed via GitHub discussions #771 and #2337).

Proposed changes:

  • New hyperlink extraction using pypdfium2.raw (FPDFLink_Enumerate + FPDFAction_GetURIPath) — the project already uses this API pattern in pdf_classifier.py
  • Extract link bounding boxes, spatial-join to Docling text items by page coordinates
  • Handle both annotation links (invisible rectangles over text) and plain-text URLs (FPDFLink_LoadWebLinks)
  • Extend fix_url_formatting() to fix escaped underscores (\_ → _) and & → & in URLs
  • Inject extracted hyperlinks into markdown during cleanup step

3. 🔴 Low-Quality Scan Detection & Gating (~20 issues, 3 documents)

La Opinion Latina alone accounts for 8 critical issues from catastrophic OCR.

Pattern Severity Affected Documents
Catastrophic OCR on low-quality newsprint Critical La Opinion Latina (8 critical)
Garbled text in scanned multi-column pages Critical Who Is a Refugee, Boxing and Masculinity
Bilingual content OCR failures Critical La Opinion Latina

Current code: src/services/pdf_classifier.py already has:

  • Hard blockers (forms, encrypted, empty → ERROR = rejected)
  • Soft warnings (FINDING_SCANNED when chars_per_page < 50, FINDING_LOW_TEXT_DENSITY when < 200)
  • Post-extraction enrich_classification() adds warnings but doesn't block

Proposed changes:

  • Add FINDING_LOW_OCR_CONFIDENCE — after Docling extraction, measure OCR quality signals (garbled character ratios, average word length anomalies, non-dictionary word frequency)
  • Add severity thresholds: WARNING for moderate degradation, ERROR for catastrophic (like La Opinion Latina)
  • Consider treating catastrophic scans like forms — hard-reject with message: "Document scan quality is too low for reliable conversion. Please provide a higher-quality scan or born-digital PDF."
  • For multi-column scanned academic chapters: add WARNING-level flag noting reduced confidence, but still process

4. 🟡 Figures & Alt Text Quality (~35 issues, 15+ documents)

Pattern Severity Affected Documents
Empty alt text on informational figures Major Language Data, Mural Tours, Mesoamerican, Rafael Cintron Ortiz
Text embedded in images not extracted Critical Education of Alice Hamilton, Transgender Youth
Orphaned figures (extracted but not referenced) Major Heat Stress (11/22), Women's Rights (5/7)
QR codes / social media icons not described Minor-Major CEDA, Mural Tours, Mesoamerican

Proposed changes:

  • Update IMAGE_DESCRIBER_SYSTEM_PROMPT in image_description.py: add explicit instructions for text-heavy images (transcribe visible text verbatim), QR codes, social media icons
  • Investigate orphaned figures bug in _replace_image_placeholders() — mismatch between <!-- image --> placeholder count and extracted figure count
  • Add post-processing validation: flag figures with empty alt text that aren't marked decorative

5. 🟡 Heading Hierarchy & Missing Titles (~15 issues, 10+ documents)

Pattern Severity Affected Documents
Document title completely missing Critical Mesoamerican, Transgender Youth
Plain text rendered as heading Major Senate Letter, Heat Stress, Electricity Prices
Flat/inconsistent heading levels Major ICMJE, Professional Licensure

Proposed changes:

  • Strengthen title detection in structure_analysis.py — always identify document title even when embedded in image/banner
  • Add cross-reference: if no h1 detected, check filename/metadata for title candidates
  • Add heading reconciliation rule for image-only titles

6. 🟡 Table Formatting (~10 issues, 5+ documents)

Pattern Severity Affected Documents
ToC/glossary mangled into multi-column table Critical Electricity Prices (3 issues)
Nested headers flattened Major Graduate Rates
Missing table rows Critical Graduate Rates

Proposed changes:

  • Update table_reconstruction.py: add ToC detection — output as ordered lists, not tables
  • Add table row count validation (compare output rows vs image)

7. 🟡 Multi-Column Reading Order (~15 issues, 5+ documents)

Hardest problem. Scanned academic chapters are marked out-of-scope for phase 1.

Pattern Severity Affected Documents
Columns interleaved incorrectly Critical Survival Migration, Who Is a Refugee, La Opinion Latina
Paragraph reordering at transitions Major Survival Migration, Who Is a Refugee
Duplicated paragraphs Critical Who Is a Refugee

Proposed changes:

  • Strengthen double_column.md prompt with more explicit column-ordering instructions
  • Explore deterministic reading order from Docling JSON bounding boxes as verification
  • For now: rely on scan quality gating (item 3) to flag worst cases

8. 🟠 Presentation & Poster Layouts (~12 issues, 4 documents)

Proposed changes:

  • New layout fragment: procedures/page_correction/layout/presentation.md (slide separators, repeated header handling)
  • New layout fragment: procedures/page_correction/layout/poster.md (zone groupings, visual hierarchy)

9. 🟢 Minor Formatting (~20 issues)

HTML entities (&amp;), missing emphasis, page number artifacts, running header remnants.

Proposed changes:

  • Add HTML entity decoding to text_cleanup.py
  • Improve running header/footer stripping

Quality by Document Type (for reference)

Type Avg Issues Primary Problems
Academic book chapters 10-13 Footnotes, reading order, duplicate endnotes
Infographics 6-7 Lost spatial layout, orphaned figures
Policy documents 5-10 Heading hierarchy, missing hyperlinks, tables
Posters / flyers 4-8 Missing informational content in images
Presentation slides 8-11 Slide boundaries, embedded text in images

Best/Worst Conversions

Best: State Infrastructure (3 issues, 0 critical), CEDA Utility (4 issues), RELS 225 (4 issues, all minor)

Worst: Latino Cultural Center (20 issues, 7 critical), La Opinion Latina (13 issues, 8 critical), Future of Work (13 issues, 3 critical)


Pilot Details

  • 30 documents, 174 total pages
  • Total cost: $13.45 ($0.45 avg/doc, $0.08/page)
  • Total tokens: 10,499,347
  • Total edits: 1,597
  • Avg conversion time: 2m44s (range: 15s to 8m13s)
  • Full review notes available in Deliverable 3 Progress Report folder

Activity

  1. dylan-isaac commented on Mar 23, 2026

    @dylan-isaac
    CollaboratorAuthor

    Update: Hyperlink Extraction — Docling Status & pypdfium2 Fallback Plan

    Docling Status (as of March 2026)

    The DoclingDocument schema does have a TextItem.hyperlink: Optional[Union[AnyUrl, Path]] field (docling-core source), and the low-level parser (docling-parse v5) extracts PDF link annotations. However, the propagation step — matching hyperlink bounding boxes to text items — was never implemented for the PDF pipeline.

    • Issue #828 (open since Jan 2025): Original report. Maintainer confirmed hyperlinks are identified in docling-parse but need propagation to DoclingDocument.
    • PR #3131: External contributor added spatial matching of hyperlink annotations to text clusters. CI passes, one maintainer approved, but it's stuck awaiting a second reviewer and is not merged.

    Result: Every TextItem in our Docling JSON response has hyperlink=null for PDF documents. URLs that appear as literal text in the PDF body are preserved, but clickable hyperlinks where display text (e.g., "click here") links to a URL are lost entirely.

    pypdfium2 Fallback Implementation Plan

    Since we already use pypdfium2.raw in pdf_classifier.py for annotation inspection (form field widget counting), we can use the same API pattern to extract hyperlinks. Here's the plan:

    Step 1: Extract link annotations from PDF bytes

    Add a new function extract_hyperlinks() to a new module (e.g., src/services/hyperlink_extractor.py) or extend pdf_classifier.py:

    import pypdfium2
    import pypdfium2.raw as pdfium_c
    import ctypes
    
    @dataclass
    class PdfHyperlink:
        page_number: int       # 1-indexed
        url: str               # Target URI
        bbox: tuple[float, float, float, float]  # (left, bottom, right, top) in PDF coords
    
    def extract_hyperlinks(file_content: bytes) -> list[PdfHyperlink]:
        """Extract hyperlink annotations from PDF using pypdfium2."""
        pdf = pypdfium2.PdfDocument(file_content)
        links: list[PdfHyperlink] = []
        
        for page_idx in range(len(pdf)):
            page = pdf.get_page(page_idx)
            try:
                # Enumerate link annotations on the page
                start_pos = ctypes.c_int(0)
                link_ptr = ctypes.c_void_p()
                
                while pdfium_c.FPDFLink_Enumerate(page.raw, ctypes.byref(start_pos), ctypes.byref(link_ptr)):
                    # Get link bounding rectangle
                    rect = pdfium_c.FS_RECTF()
                    if not pdfium_c.FPDFLink_GetAnnotRect(link_ptr, ctypes.byref(rect)):
                        continue
                    
                    # Get action and check if it's a URI action (type 1)
                    action = pdfium_c.FPDFLink_GetAction(link_ptr)
                    if not action:
                        continue
                    if pdfium_c.FPDFAction_GetType(action) != 1:  # PDFACTION_URI
                        continue
                    
                    # Extract the URI string
                    buf_size = pdfium_c.FPDFAction_GetURIPath(pdf.raw, action, None, 0)
                    if buf_size <= 0:
                        continue
                    buf = ctypes.create_string_buffer(buf_size)
                    pdfium_c.FPDFAction_GetURIPath(pdf.raw, action, buf, buf_size)
                    url = buf.value.decode("utf-8", errors="replace")
                    
                    links.append(PdfHyperlink(
                        page_number=page_idx + 1,
                        url=url,
                        bbox=(rect.left, rect.bottom, rect.right, rect.top),
                    ))
            finally:
                page.close()
        
        pdf.close()
        return links

    This follows the exact same pattern as _extract_metadata() in pdf_classifier.py (lines 278-302), which already uses pdfium_c.FPDFPage_GetAnnotCount, FPDFPage_GetAnnot, and FPDFAnnot_GetSubtype.

    Step 2: Spatial-join hyperlinks to Docling text

    The Docling JSON response includes bounding boxes for text items (body → items → prov → bbox). We can match hyperlink rectangles to the text they cover:

    • For each PdfHyperlink, find text items on the same page whose bounding box overlaps with the link rectangle
    • The overlapping text becomes the display text for the markdown link
    • Inject [display text](url) into the per-page markdown at the corresponding position

    This spatial join would happen in _step_docling() (pipeline_viewer.py ~line 700), right after extracting figures and before the pipeline proceeds to AI processing steps.

    Step 3: Inject into per-page markdown

    After identifying which text spans are linked, modify the per-page markdown (v0) to wrap linked text in markdown link syntax. This runs as a deterministic post-processing step — no LLM needed.

    Step 4: Future-proof for Docling native support

    Structure the code so that when PR #3131 merges and TextItem.hyperlink becomes populated:

    1. Check json_content for populated hyperlink fields first
    2. Fall back to pypdfium2 extraction only if Docling doesn't provide them
    3. Eventually remove the pypdfium2 fallback entirely
    def get_hyperlinks(file_content: bytes, json_content: dict) -> list[PdfHyperlink]:
        """Get hyperlinks from Docling JSON if available, else fall back to pypdfium2."""
        docling_links = extract_hyperlinks_from_docling_json(json_content)
        if docling_links:
            return docling_links
        return extract_hyperlinks_from_pypdfium2(file_content)

    Where this fits in the pipeline

    _step_docling():
        1. Docling extraction + pypdfium2 page rendering (parallel) — existing
        2. Split markdown by page — existing
        3. Extract figures from JSON + crop — existing  
        4. Replace image placeholders — existing
        5. ✨ NEW: Extract hyperlinks (pypdfium2) + spatial-join + inject into markdown
        6. Detect column layout — existing
    

    The hyperlink injection happens once, deterministically, at the end of _step_docling(). All subsequent pipeline steps (structure analysis, page correction, boundary fixes, cleanup) will see the hyperlinks already in the markdown and preserve them.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions