Skip to content

feat(ai-service): XLSX, DOCX, MSG and scanned PDFs #23

Description

@Fluory

Goal

All document formats named in the customer request are parsed into segments with stable locators, including scanned PDFs, so evidence always points to a findable place.

Acceptance criteria

  • XLSX → segments with Sheet!A1 locators (one segment per non-empty cell or row – documented choice)
  • DOCX → paragraph and table-cell locators
  • MSG → body plus attachments, parsed recursively with the same parsers
  • Scanned PDF → OCR via docling; OCR segments are flagged ocr: true; a field whose only evidence is OCR text is at most uncertain
  • A failing attachment does not fail the request: per-document parse error is recorded and shown, the other documents are processed
  • Docling model downloads and CPU latency are measured and noted in the PR (ADR-0001 open point)

Not part of this task

  • Line-item review UI, eval gate

Affected areas

  • services/ai/src/**/parsing/
  • services/ai/tests/fixtures/

Test plan

Criterion Check
each format pytest with small synthetic fixtures generated by a committed script
partial failure pytest: one corrupt attachment, others still parsed
OCR cap pytest: OCR-only evidence never found

Security/Privacy affected?

Yes – untrusted documents: size limits, no macro execution, parser errors never crash the service; security rule parsing/** applies.

Epic: #17 · Architecture: docs/decisions/ADR-0001-pilot-architecture.md

Activity

  1. added
    featureNew capability
    securitySecurity or privacy relevant
    aiLLM, prompts, evals
    readyDefinition of Ready met – may be claimed
    on Sep 22, 2026
  2. self-assigned this
    on Sep 23, 2026
  3. Fluory commented on Sep 23, 2026

    @Fluory
    OwnerAuthor

    Claimed by @Fluory (Claude Code, overnight run) on branch claude/feat-ai-formats-23 – stacked on #39. AI-service part runs in a parallel worktree; draft PR follows.


    Generated by Claude Code

  4. Fluory commented on Sep 23, 2026

    @Fluory
    OwnerAuthor

    decision-needed – OCR models in the image (PR #40)

    OCR for scanned PDFs uses docling's models: the layout model (164 MB, Hugging Face) and the OCR models (~31 MB), which come from a second external host, modelscope.cn. I took the most reversible option:

    • OCR is off by default (AI_PDF_OCR=off). With auto, only pages without a text layer are OCRed.
    • The Dockerfile gets an opt-in build argument, PREFETCH_OCR_MODELS, that downloads the models into the image at build time, so runtime needs no outbound access. The image was not built in this run.

    To decide: enable OCR in the pilot, bake the models into the image (and allow modelscope.cn at build time), or mirror the models to an EU host you control.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

aiLLM, prompts, evalsfeatureNew capabilityreadyDefinition of Ready met – may be claimedsecuritySecurity or privacy relevant

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions