Skip to content

feat(ai-service): XLSX, DOCX, MSG and scanned PDFs - #40

Merged
Fluory merged 136 commits into
mainfrom
claude/feat-ai-formats-23
Sep 23, 2026
Merged

Fluory merged 136 commits into
mainfrom
claude/feat-ai-formats-23

Conversation

@Fluory

@Fluory Fluory commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Warum

Fixes #23 · Epic #17 · gestapelt auf #39 (Basis claude/feat-ai-full-fields-22).

Arbeitsstand

Was ist passiert (Klartext)

Die KI liest jetzt auch Excel-Tabellen (zeilenweise), Word-Dokumente (Absätze und Tabellenzellen) und Outlook-Nachrichten mit ihren Anhängen. Jede Fundstelle zeigt genau, wo ein Wert steht. Gescannte PDFs können per Texterkennung gelesen werden – solche Werte gelten höchstens als „unsicher"; bei sehr langen Scans werden die ersten 10 Seiten gelesen und der Rest mit Hinweis übersprungen. Ein kaputter Anhang bringt nicht mehr die ganze Anfrage zum Scheitern – die Prüfansicht zeigt, welcher Anhang nicht gelesen werden konnte. Sicherheitsfund unterwegs: Eine präparierte 2-KB-Outlook-Datei konnte den Dienst minutenlang mit Gigabytes Speicher blockieren; das ist jetzt abgefangen (sofortige Ablehnung), zusammen mit weiteren Grenzen für übergroße Dateien.

Plan-Pflicht (SYSTEM.md §4)

  • Kein Auslöser – keine Modulgrenze, öffentliche API, Migration, Auth/Rechte, kein Zahlungs-/Daten-/Infrapfad, höchstens zwei Module, keine Architekturvarianten, umkehrbar
  • Auslöser zutreffend – Impact Manifest ausgefüllt (Plan vor Code)

Impact Manifest

  • Betroffene Module: AI-Service (parsing/: document, xlsx, docx, msg, ooxml, budget; pdf OCR; Verifier-Deckel; API/Schemas), Vertrag contracts/ai-service.openapi.yaml (generiert), extraction (Client-Validierung, Speichern der Anhangsfehler), jobs (Worker sendet XLSX/DOCX/MSG), review (Anzeige je Dokument).
  • Schnittstellen / Datenänderungen: Vertrag additiv: documentKind + xlsx, docx, msg; Locator xlsx (sheet, row, cellRange), docx (part, paragraph/table, row, cell), msg (Kopf/Text wie E-Mail; Anhänge attachment + inner); pdf.ocr; Reason ocr_only; Warnungen attachment_failed, ocr_pages_skipped; attachments[] mit Fehlercodes inkl. budget_exceeded. Keine Migration (Anhangsfehler im bestehenden JSON extraction_runs.documents).
  • Akzeptanzkriterien: alle aus feat(ai-service): XLSX, DOCX, MSG and scanned PDFs #23 – siehe Nachweis.
  • Testplan: pytest je Format mit per Skript erzeugten synthetischen Fixtures, Teilfehler, OCR-Deckel, Zip-/Teil-/Verschachtelungsgrenzen, OLE-Bombe, Budget, HTTP-Ebene, Logs ohne Inhalt; Vitest/Integration: Worker verarbeitet XLSX, Anhangsfehler in der Prüfansicht.
  • Verifizierte Fakten: openpyxl 3.1.5, python-docx 1.2.0, python-oxmsg 0.0.2 (MIT), olefile 0.47 (BSD) – bereits über docling installiert, direkt gepinnt; extract-msg (GPL) nicht verwendet. Messung (4 vCPU): OCR 8–9 s/Seite, Modelle 164 MB (Hugging Face) + ~31 MB (modelscope.cn).
  • Offene Annahmen: decision-needed (feat(ai-service): XLSX, DOCX, MSG and scanned PDFs #23): opt-in PREFETCH_OCR_MODELS im Dockerfile (Infra, Image nicht gebaut) und Host modelscope.cn; OCR standardmäßig aus. Strenger olefile-Modus könnte echte, leicht fehlerhafte .msg ablehnen (keine echte Probe vorhanden).
  • Nicht-Ziele: siehe oben.
  • Risiken und Rollback: feindliche Dokumente → feste Grenzen + Regressionstests; keine Migration; Rollback per Revert.

Geändert

  • services/ai/** (Parser, Budget, Pipeline, Verifier, API, Tests, Fixtures, README, Dockerfile, pyproject.toml/uv.lock), contracts/ai-service.openapi.yaml
  • src/features/extraction/{ai-client,ai-service.contract,repository}.ts, src/features/jobs/process-request.ts, src/features/review/review.ts, src/app/requests/[id]/page.tsx
  • Tests: tests/integration/{processing,review}.test.ts
  • Doku: docs/technical/{architecture,operations}.md, CHANGELOG.md

Nachweis (SYSTEM.md §11)

  • verify:changed: AI-Service ruff/format/pyright grün, pytest 387 passed / 2 skipped (brauchen heruntergeladene Modelle)
  • verify: lokal grün (Exit 0, Node 24) – Unit 135/135, Integration 98/98, depcruise ohne Verstöße, Build, Audit; CI check: siehe Checks
  • verify:full / E2E-Spec: Stapelende (feat(observability): correlated structured logs and full health #45) mit diesem PR: verify:full lokal grün (siehe dort)
  • Akzeptanzkriterien feat(ai-service): XLSX, DOCX, MSG and scanned PDFs #23: XLSX mit Zeilen-Locator (dokumentierte Wahl: Zeile) → pytest xlsx; DOCX Absatz/Zelle → pytest docx; MSG Text + Anhänge rekursiv → pytest msg; Scan-PDF OCR, ocr: true, nur-OCR-Beleg höchstens unsicher → pytest OCR-Deckel (auch in Anhängen/Positionen); fehlender Anhang ohne Gesamtfehler, Fehler gespeichert und angezeigt → pytest Teilfehler + Integration „keeps failed attachments…"; Modell-Downloads und CPU-Latenz gemessen → oben und in operations.md
  • Manueller Prüfnachweis: –
  • Frischer Review (P1 vor Ready-for-review; Architektur/API/DB immer): unabhängiger Security-Review mit Probe-Dokumenten – 1 Blocker (OLE-Bombe), 4 should-fix, 4 nits – alle behoben/begründet (PR-Kommentar)

Doku-Entscheidung (genau eine)

  • Keine langlebige Doku betroffen – Begründung: –
  • Doku betroffen und im selben PR aktualisiert:
    • Produktdoku (P0: README): –
    • Technische Doku: docs/technical/operations.md, services/ai/README.md
    • Architekturkarte: docs/technical/architecture.md
    • ADR: –
    • CHANGELOG [Unreleased] (sichtbares Feature oder Verhalten – im selben PR, nie „später")

Entferntes oder Umbenanntes: Test test_outlook_msg_is_rejected_for_now durch MSG-Tests ersetzt; test_ocr_page_cap durch „erste 10 Seiten + Warnung" ersetzt (bewusste Verhaltensänderung aus dem Review)

Dateigrößen und neue Bausteine (SYSTEM.md §7)

Dateien über 500 Zeilen im Diff (Ausnahmen: generierter Code, Lockfiles, Fixtures, Migrationen, Schemas, Ressourcen, Doku, Konfiguration):

  • keine
  • bewusst belassen – Begründung: src/features/extraction/ai-service.contract.ts (592 Zeilen) ist von openpyxl-unabhängig generiert (pnpm contract:types aus dem OpenAPI-Vertrag) und wird nie von Hand bearbeitet; der Drift-Test sichert die Übereinstimmung.
  • im selben PR nach fachlicher Verantwortung geteilt
  • Folge-Issue –

Über 800 Zeilen mit neuer Fachlogik oder über 1000 Zeilen (P1/P2): nicht betroffen

Neue Shared-Komponente, Utility-Datei, Adapter oder fachlicher Service:

  • nein
  • ja – gesucht nach: docling-Readern für XLSX/DOCX/Mail; gefunden: liefern keine Zell-/Absatz-/Zeilenpositionen → eigene Parser auf openpyxl/python-docx/python-oxmsg

Subagent-Einsätze

  • AI-Service-Teil in eigenem Worktree (Subagent), Security-Review (read-only, Probes), Security-Fixes (Subagent auf diesem Branch) – vom Hauptagenten zusammengeführt und geprüft.

Risiken / offene Punkte

  • decision-needed: PREFETCH_OCR_MODELS (Dockerfile) und Host modelscope.cn.
  • Echte .msg-Dateien ungetestet (nur synthetische); RTF-only-Bodies liefern keine Textsegmente.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1

…ion access control

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…LS with integration proofs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…boss queues

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…nd download routes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…-only audit, composite FK

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
- sign-up requires the invitation id (link) plus the invited e-mail – no takeover by address alone
- one company per user (unique index), deterministic membership lookup, actor from membership
- organization plugin accepts only admin/clerk roles
- configurable client-IP source for the auth rate limit; local secret refused in production
- invite page: zod input, 404 for clerks, shows the invitation link; signup needs the link
- tests: wrong/missing invitation id, foreign set-active/list-members, last admin, roles
- docs: operations (rate limit/proxy, recovery), data model, exceptions register (admin plugin)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…e rejection

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
… docs for #5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…secret

The image sets NODE_ENV=production, so the previous check blocked the local compose stack.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…re script

Python 3.13 package requestflow_ai (src layout), docling/fastapi/google-genai
pinned, CPU-only torch via the PyTorch CPU index, dev tools ruff/pyright/pytest.
The synthetic PDF fixture is generated by scripts/make_fixtures.py (reportlab).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Test-first per ADR-0001 D8: quote normalisation (whitespace, case, NFKC,
hyphenation), German number and date formats, and the verifier rules that
turn unsupported model claims into unverified. Also adds the segment and
model-output types the tests build on; the verifier does not exist yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
normalize_text folds NFKC, soft hyphens, line-break hyphenation, typographic
dashes/quotes, whitespace and case. values parses German/ISO numbers and
DD.MM.YYYY, D.M.YY and ISO dates. verify_field only keeps or downgrades the
model's status: a quote not in the cited segment, an unknown segment, a value
inconsistent with the quote, found without evidence or value, and missing
with a value all become unverified with a reason.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ction (red)

EML: body lines with 1-based line locators, From/Subject header segments,
RFC 2047 and quoted-printable decoding, HTML-only bodies, header newline
collapse, no attachment payloads. PDF: textline segments with page + top-left
bbox via docling-parse (model-free); the layout-pipeline test only runs with
AI_TEST_DOCLING_MODELS=1. Detection by magic bytes; .msg rejected for now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
EML via the standard library email package (policy=default): From/Subject
header segments and one segment per non-empty body line, HTML-only bodies
reduced to text lines. docling's EMAIL backend was checked but emits
paragraphs without provenance, so it cannot give line locators.

PDF via docling: the default textlines pipeline reads docling-parse text
lines (page + top-left bbox, no ML models, no network); the opt-in layout
pipeline uses DocumentConverter (OCR and tables off) and the heron layout
model. docling is imported lazily.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ex model client (red)

The model client is exercised through the real google-genai SDK with an
httpx MockTransport replaying recorded generateContent bodies: eu multi-region
URL, bearer auth, structured-output config, token usage, fail-closed init
(no project, no credentials, dev flag without key), no silent switch to API
key mode, and schema-invalid output. Adds the env-based Settings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…lient

Prompt extract_header_v1.md (system instruction) plus render_document, which
puts segment-id-prefixed lines between <document> delimiters and neutralises
delimiter-like tags inside the document. GeminiModelClient calls Vertex via
google-genai with response_schema = ModelExtraction, JSON mime type,
temperature 0 and no tools; schema-invalid output raises ModelOutputError.
build_model_client needs VERTEX_PROJECT and credentials, loads ADC eagerly,
refuses API-key mode for Vertex, and allows the Gemini API only with
AI_ALLOW_GEMINI_API_DEV=true plus GEMINI_API_KEY.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Pipeline: PDF and EML end to end with recorded model responses, a model date
contradicting its own quote, the prompt-injection mail (the recorded model
obeys and invents a quote -> unverified), and the no-text PDF that skips the
model. API: bearer auth (401 + WWW-Authenticate), response shape, request id
handling, 400/413/415/422/429/502 mapping without echoing input, JSON logs
with IDs only. Contract: the committed OpenAPI file equals the app's schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ect dependencies

Already installed through docling[standard]; the XLSX/DOCX/MSG parsers of #23 import them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ment builders

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
… into pipeline and API

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ent and OCR counts

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…d of the deprecated flag

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…act (#23)

- documentKind gains xlsx, docx, msg; locators XlsxLocator, DocxLocator, MsgLocator (recursive
  inner locator for attachments); PdfLocator gains optional ocr flag
- MSG attachments parsed recursively with the same parsers; a failing attachment is reported in
  the new optional attachments list and warning attachment_failed, never fails the message
- AI_PDF_OCR=auto: docling + RapidOCR (torch) only on pages without text; OCR evidence is capped
  at uncertain (reason ocr_only)
- Contract regenerated; contract, parser, verifier and API tests

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ion to prefetch OCR models

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…oke of pages and PDF only

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ever read attachments beyond the cap

olefile only logs 'stream too large' by default, and python-oxmsg reads every
stream at load time, so a 2 KiB .msg declaring a gigabyte stream on a looped
FAT chain was read sector by sector. check_ole_container opens the directory
with DEFECT_INCORRECT and rejects any stream, the mini stream, or the sum of
streams declaring more bytes than the container has. Used by detect and
load_message.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…X block

A 60 MB document.xml of empty paragraphs (8.8 MB zipped) passed the 64 MiB/100:1
checks and cost 37 s / 1.3 GB for 0 segments. open_package now rejects any
entry declaring more than MAX_PART_BYTES (16 MiB) from the zip index, and the
DOCX walker counts every paragraph and table cell against MAX_BLOCKS (100,000),
not only emitted segments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…rest with a warning

More pages without text than the OCR cap used to fail the whole PDF (422).
Now the first ones are OCR'd (consecutive pages in one conversion), the rest
stay empty and the response carries the additive warning ocr_pages_skipped.
parse_pdf_document returns page count and OCR counts; parse_pdf stays a wrapper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Per-file caps multiplied across the attachments of a .msg. ParseBudget
(parsing/budget.py) is created once per uploaded document and shared by all
nested attachments: 100 attachments, 128 MiB unzipped OOXML, 100 PDF pages
(at least AI_MAX_PDF_PAGES), MAX_OCR_PAGES OCR pages. An attachment that would
overdraw it is reported with the additive error code budget_exceeded; the rest
of the message is still returned. BudgetExceededError subclasses
DocumentTooLongError, so a top-level document still maps to 422.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…EADME limits for the security fixes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ges_skipped

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
olefile follows the declared FAT sector count along the DIFAT chain in its
constructor (quadratic array copy per sector); a forged count on a looped
DIFAT chain hung it before the stream-size walk could run. The header's FAT,
DIFAT, mini FAT and directory sector counts must now fit the file size.
Also: stale docstring, README package list.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ored per document and shown in the review; contract types regenerated (#23 review)

Fluory commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Frischer Security-Review (unabhängiger Subagent, read-only, mit Probe-Dokumenten)

Ergebnis: 1 Blocker · 4 should-fix · 4 nits – alle behoben bzw. begründet. Bestätigt per Probe: XXE/Entity-Expansion in DOCX/XLSX sauber abgelehnt (0,01 s, nichts gelesen), Zip-Grenzen vor der Dekompression, keine Makros, Formeln nur gespeicherte Werte, Logs nur Klassen/Codes, Anhangsnamen nur als Text, OCR-Deckel rekursiv (auch Anhänge, Positionen), Segment-IDs eindeutig, Vertrag additiv.

# Schwere Befund Auflösung
1 blocker OLE-Stream-Bombe: 2-KB-.msg mit Riesen-Stream + FAT-Schleife → 3,4 GB RSS / 54 s, „geparst" Vor jedem Stream-Lesen: olefile mit DEFECT_INCORRECT, Ablehnung (422 document_unparseable) wenn ein Stream, der Mini-Stream oder die Summe mehr Bytes deklariert als die Datei hat, plus Plausibilität der Sektorzahlen im Header (fing eine zweite Hängevariante). Probe jetzt in 0,0 s abgelehnt. Anhänge über der Grenze werden nicht mehr dekodiert. Regressionstest
2 should-fix 60-MB-document.xml (8,8 MB gezippt) innerhalb der Grenzen → 37 s / 1,3 GB Einzelteile > 16 MiB unkomprimiert werden aus dem Zip-Index abgelehnt; DOCX zählt jeden Absatz/jede Zelle (auch leere) gegen 100.000. Test
3 should-fix kein Gesamtbudget je Dokument (viele gescannte Anhänge → Minuten) gemeinsames Budget je Dokument (parsing/budget.py): 100 Anhänge, 128 MiB entpackt, 100 PDF-Seiten, 10 OCR-Seiten; Überziehende Anhänge → budget_exceeded, Rest wird geliefert. Test
4 should-fix > 10 Seiten ohne Text ließ das ganze PDF scheitern erste 10 werden per OCR gelesen, der Rest übersprungen mit Warnung ocr_pages_skipped (Test ersetzt, Verhalten bewusst geändert)
5 should-fix Anhangsfehler wurden verworfen, nirgends angezeigt TS: fehlgeschlagene Anhänge (Name, Code) und die OCR-Warnung werden je Dokument gespeichert und in der Prüfansicht angezeigt („⚠ Anhang „…" konnte nicht gelesen werden – bitte im Original prüfen"). Integrationstest
nits – python-oxmsg _storage, doppeltes Öffnen, modelscope-Checksums, fehlende Probe-Tests Pin + Kommentar; doppeltes Öffnen bleibt (Bibliotheken brauchen Rohbytes); Modell-Host in decision-needed #23; Probe-Fälle als Regressionstests

Danach: AI-Service ruff/format/pyright grün, pytest 387 passed / 2 skipped; pnpm verify lokal grün (Unit 135/135, Integration 98/98). Restrisiko: der strengere olefile-Modus könnte echte, leicht fehlerhafte Outlook-Dateien ablehnen – ohne echte Probe ungetestet (README).


Generated by Claude Code

@Fluory
Fluory marked this pull request as ready for review September 23, 2026 09:08
@Fluory
Fluory changed the base branch from claude/feat-ai-full-fields-22 to main September 23, 2026 10:00
@Fluory
Fluory merged commit 18545f1 into main Sep 23, 2026
7 of 8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(ai-service): XLSX, DOCX, MSG and scanned PDFs

2 participants