Skip to content

feat(ai-service): stateless extraction service with grounding verifier - #33

Merged
Fluory merged 56 commits into
mainfrom
claude/feat-ai-service-6
Sep 23, 2026
Merged

Fluory merged 56 commits into
mainfrom
claude/feat-ai-service-6

Conversation

@Fluory

@Fluory Fluory commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Warum

Fixes #6 · Teil von Epic #2 · gestapelt auf #32 (Basis claude/feat-intake-upload-5).

Arbeitsstand

Was ist passiert (Klartext)

Es gibt jetzt den KI-Dienst: Er bekommt eine E-Mail oder ein PDF, zerlegt das Dokument in nummerierte Abschnitte mit genauer Fundstelle (Seite und Rechteck im PDF, Zeile in der Mail) und lässt das Sprachmodell drei Angaben heraussuchen – Firma, Ansprechpartner, gewünschter Liefertermin –, jeweils mit wörtlichem Zitat. Ein festes Prüfprogramm (kein KI-Modell) kontrolliert danach, ob das Zitat wirklich an der genannten Stelle steht und zum Wert passt (ganze Wörter, Daten als Kalenderdatum). Stimmt das nicht, wird der Wert als „nicht bestätigt" markiert. Das Modell entscheidet also nie allein, dass etwas „gefunden" ist.

Der Dienst hat keinen Zugriff auf Datenbank oder Speicher und kennt keine Firmen – er bekommt nur die Bytes und neutrale IDs. Anfragen ohne gültiges Token oder mit zu großer Länge werden abgewiesen, bevor auch nur ein Byte gelesen wird. Ein echter Aufruf von Vertex AI wurde nicht gemacht (keine Zugangsdaten in dieser Umgebung); die Anbindung ist gegen die SDK-Doku gebaut und mit handgeschriebenen Antworten im Vertex-Format getestet.

Plan-Pflicht (SYSTEM.md §4)

  • Kein Auslöser – keine Modulgrenze, öffentliche API, Migration, Auth/Rechte, kein Zahlungs-/Daten-/Infrapfad, höchstens zwei Module, keine Architekturvarianten, umkehrbar
  • Auslöser zutreffend – Impact Manifest ausgefüllt (Plan vor Code)

Impact Manifest

  • Betroffene Module: AI-Service services/ai/ (parsing, extraction, grounding, api, config, jsonlog), contracts/, extraction (TS-Typen).
  • Schnittstellen / Datenänderungen: neuer interner HTTP-Dienst POST /v1/extract (multipart: Datei + documentId, Header X-Request-Id, Bearer-Token), GET /healthz; Vertrag OpenAPI 3.1; keine Datenbank-/Speicheränderung.
  • Akzeptanzkriterien: alle aus feat(ai-service): stateless extraction service with grounding verifier #6 (Nachweis unten).
  • Testplan: pytest test-first: Verifier inkl. Normalisierung (Zahlen deutsch/englisch, Daten → ISO, Wortgrenzen, mehrere Daten, Unicode, Silbentrennung), Parser (EML-Zeilen, PDF Seite + BBox, Seitengrenze), Modell-Client gegen das echte SDK mit gemocktem HTTP-Transport (exakte eu-URL), Pipeline End-to-End, API (Auth vor dem Body, Längen, Fehlerformen, 500 mit Korrelation), Prompt-Injection, Vertrags-Drift; TS: Typen-Drift-Test.
  • Verifizierte Fakten: docling 2.130.0 (MIT) unterstützt .eml/.msg über mail-parser/python-oxmsg, liefert aber nur Absätze ohne Zeilennummern → EML über stdlib email (policy=default); PDF-Textebene ohne Modelle über docling-parse (Seite + BBox); der docling-Standardkonverter braucht auch ohne OCR das Layout-Modell docling-layout-heron (Download von Hugging Face klappte, ~12 s CPU für die erste Konvertierung); google-genai 2.25.0: location="eu" → https://aiplatform.eu.rep.googleapis.com/ (API v1beta1), per Test gegen das echte SDK geprüft; ein API-Key in der Umgebung schaltet Vertex nicht in den Key-Modus.
  • Offene Annahmen: gemini-3.5-flash in eu verfügbar (laut ADR, nicht live geprüft); Token-Zahlen/Latenz echter Aufrufe unbekannt.
  • Nicht-Ziele: siehe oben.
  • Risiken und Rollback: Dokumentinhalt geht an ein externes LLM → nur eu-Endpunkt, nur synthetische Daten, kein Inhalt in Logs; Free Tier nur mit explizitem Dev-Flag und nie zusammen mit VERTEX_PROJECT; Rollback per Revert.

Akzeptanzkriterien → Nachweis

Kriterium Nachweis
OpenAPI 3.1 committed; TS-Client/Typen generiert contracts/ai-service.openapi.yaml (aus FastAPI exportiert, test_contract.py Drift-Test); src/features/extraction/ai-service.contract.ts via pnpm contract:types, Vitest-Drift-Test (kann fehlschlagen – geprüft). Laufzeit-Client: #7 (PR #34)
Stabile Locators (PDF Seite + BBox, EML Zeile) test_parsing_eml.py, test_parsing_pdf.py
Pro Feld Wert/Status/Evidenz für company, contact_person, requested_delivery_date test_pipeline.py, test_api.py
Verifier test-first → unverified test_grounding_verifier.py, test_grounding_values.py, test_grounding_normalize.py (rote Commits vor dem Code)
Keine DB-/Storage-Credentials; Bearer; Logs ohne Inhalt Settings ohne DB/S3-Variablen; test_api.py (401 vor dem Body, konstante Vergleichszeit, Log ohne Dokumentinhalt)
Modell/Region aus Konfiguration; Free Tier nur mit Dev-Flag, fail-closed test_extraction_model_client.py (Start ohne Projekt/Credentials bricht ab; Key ohne Dev-Flag ignoriert; Dev-Flag + Projekt → Start verweigert)
docling-EML/MSG-Abdeckung und eu-SDK-Konfiguration dokumentiert services/ai/README.md + „Verifizierte Fakten" oben

Bekannte Grenze (per Test festgenagelt): Zitiert das Modell den eingeschleusten Satz selbst wörtlich, ist das Zitat echt im Dokument und gilt als belegt. Dagegen helfen Prompt, Schema und die menschliche Prüfung vor dem Export; der Eval-Satz (#24) bekommt Injection-Fälle.

Neue Dependencies (Begründung, Lizenzen geprüft – kein GPL/AGPL, kein PyMuPDF)

Paket Version Lizenz Warum
docling 2.130.0 MIT Parsing laut ADR D8
fastapi / uvicorn 0.141.1 / 0.53.0 MIT / BSD-3 HTTP-Dienst
pydantic / pydantic-settings 2.13.5 / 2.15.0 MIT Schema, Structured Output, Konfiguration
python-multipart 0.0.32 Apache-2.0 Upload
google-genai / google-auth 2.25.0 / 2.58.0 Apache-2.0 Vertex AI (eu)
torch / torchvision (CPU-Wheels) 2.14.0 / 0.29.0 BSD von docling benötigt; CPU-Index statt ~3 GB CUDA
dev: ruff, pyright, pytest, httpx, pyyaml, reportlab – MIT/BSD Checks, Tests, Fixture-Skript
TS dev: openapi-typescript 7.13.0 MIT Typen aus dem Vertrag

Installationsgröße .venv ≈ 1,5 GB (davon torch CPU ≈ 0,7 GB).

Geändert

  • services/ai/: Projekt (uv, gepinnt), Pakete parsing, extraction (Prompt extract_header_v1.md, Modell-Client), grounding (Normalisierung, Verifier), api (inkl. Middleware für Token/Länge vor dem Body), config, jsonlog, pipeline; Tests + synthetische Fixtures (Skript); Dockerfile (non-root); README.
  • contracts/ai-service.openapi.yaml.
  • src/features/extraction/ai-service.contract.ts (generiert), index.ts; tests/architecture/contract-types.test.ts; package.json (contract:types); .gitignore (.venv auch als Symlink).
  • Doku: Architekturkarte, CHANGELOG.

Nachweis (SYSTEM.md §11)

  • verify:changed: grün
  • verify: grün – TS-Seite lokal; AI-Service lokal: uv sync --frozen, ruff check ✔, ruff format --check ✔, pyright 0 Fehler, pytest -q 199 passed / 1 skipped (Layout-Modell-Test, mit AI_TEST_DOCLING_MODELS=1 grün). CI check inkl. Python-Job: siehe Checks (vorheriger Stand grün in 2 min 22 s).
  • verify:full / E2E-Spec: nicht betroffen
  • Manueller Prüfnachweis: Serverstart verweigert ohne Google-Credentials, mit zu kurzem Token und mit Dev-Flag + Projekt; realer Server: 1-TB-Anfrage ohne Token → 401 bei 0 übertragenen Bytes. Live-Aufruf unverified.
  • Frischer Review (P1 vor Ready-for-review; Architektur/API/DB immer): feat(ai-service): stateless extraction service with grounding verifier #33 (comment) – alle Befunde behoben

Doku-Entscheidung (genau eine)

  • Keine langlebige Doku betroffen – Begründung: –
  • Doku betroffen und im selben PR aktualisiert:
    • Produktdoku (P0: README): –
    • Technische Doku: services/ai/README.md (Betrieb, Env, Grenzen, verifiziert/unverifiziert)
    • Architekturkarte: docs/technical/architecture.md (AI service, Contracts, extraction)
    • ADR: –
    • CHANGELOG [Unreleased] (sichtbares Feature oder Verhalten – im selben PR, nie „später")

Entferntes oder Umbenanntes: nichts entfernt

Dateigrößen und neue Bausteine (SYSTEM.md §7)

Dateien über 500 Zeilen im Diff (Ausnahmen: generierter Code, Lockfiles, Fixtures, Migrationen, Schemas, Ressourcen, Doku, Konfiguration):

  • keine
  • bewusst belassen – Begründung: –
  • im selben PR nach fachlicher Verantwortung geteilt
  • Folge-Issue –

Über 800 Zeilen mit neuer Fachlogik oder über 1000 Zeilen (P1/P2): nicht betroffen

Neue Shared-Komponente, Utility-Datei, Adapter oder fachlicher Service:

  • nein
  • ja – gesucht nach: vorhandenem Python-/Parsing-/Verifier-Code; gefunden: nichts – services/ai ist laut ADR D8 der vorgesehene Dienst und entsteht hier erstmals

Subagent-Einsätze

  • Implementierungs-Subagent (gleiches Modell, isolierter Worktree, ohne Push-Rechte): baute services/ai test-first; Commits unverändert übernommen (cherry-pick).
  • Frischer Review: unabhängiger Subagent (read-only) mit eigenen Sonden – 4 Important, 4 Notes.
  • Fix-Subagent (isolierter Worktree, ohne Push): Review-Befunde test-first behoben; übernommen und hier geprüft.

Risiken / offene Punkte

  • Live-Aufruf von Vertex AI unverifiziert (keine Credentials; Free Tier verboten).
  • OCR und Tabellenstruktur aus – ein PDF ohne Text liefert alle Felder missing mit Warnung no_text (Scans: feat(ai-service): XLSX, DOCX, MSG and scanned PDFs #23).
  • Ob native stderr-Ausgaben von docling-parse Dokumenttext enthalten können: unbekannt (README).

🤖 Generated with Claude Code

https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1

…ion access control

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…LS with integration proofs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…boss queues

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…nd download routes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…-only audit, composite FK

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
- sign-up requires the invitation id (link) plus the invited e-mail – no takeover by address alone
- one company per user (unique index), deterministic membership lookup, actor from membership
- organization plugin accepts only admin/clerk roles
- configurable client-IP source for the auth rate limit; local secret refused in production
- invite page: zod input, 404 for clerks, shows the invitation link; signup needs the link
- tests: wrong/missing invitation id, foreign set-active/list-members, last admin, roles
- docs: operations (rate limit/proxy, recovery), data model, exceptions register (admin plugin)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…e rejection

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
… docs for #5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…secret

The image sets NODE_ENV=production, so the previous check blocked the local compose stack.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…re script

Python 3.13 package requestflow_ai (src layout), docling/fastapi/google-genai
pinned, CPU-only torch via the PyTorch CPU index, dev tools ruff/pyright/pytest.
The synthetic PDF fixture is generated by scripts/make_fixtures.py (reportlab).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Test-first per ADR-0001 D8: quote normalisation (whitespace, case, NFKC,
hyphenation), German number and date formats, and the verifier rules that
turn unsupported model claims into unverified. Also adds the segment and
model-output types the tests build on; the verifier does not exist yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
normalize_text folds NFKC, soft hyphens, line-break hyphenation, typographic
dashes/quotes, whitespace and case. values parses German/ISO numbers and
DD.MM.YYYY, D.M.YY and ISO dates. verify_field only keeps or downgrades the
model's status: a quote not in the cited segment, an unknown segment, a value
inconsistent with the quote, found without evidence or value, and missing
with a value all become unverified with a reason.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ction (red)

EML: body lines with 1-based line locators, From/Subject header segments,
RFC 2047 and quoted-printable decoding, HTML-only bodies, header newline
collapse, no attachment payloads. PDF: textline segments with page + top-left
bbox via docling-parse (model-free); the layout-pipeline test only runs with
AI_TEST_DOCLING_MODELS=1. Detection by magic bytes; .msg rejected for now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
EML via the standard library email package (policy=default): From/Subject
header segments and one segment per non-empty body line, HTML-only bodies
reduced to text lines. docling's EMAIL backend was checked but emits
paragraphs without provenance, so it cannot give line locators.

PDF via docling: the default textlines pipeline reads docling-parse text
lines (page + top-left bbox, no ML models, no network); the opt-in layout
pipeline uses DocumentConverter (OCR and tables off) and the heron layout
model. docling is imported lazily.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ex model client (red)

The model client is exercised through the real google-genai SDK with an
httpx MockTransport replaying recorded generateContent bodies: eu multi-region
URL, bearer auth, structured-output config, token usage, fail-closed init
(no project, no credentials, dev flag without key), no silent switch to API
key mode, and schema-invalid output. Adds the env-based Settings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…lient

Prompt extract_header_v1.md (system instruction) plus render_document, which
puts segment-id-prefixed lines between <document> delimiters and neutralises
delimiter-like tags inside the document. GeminiModelClient calls Vertex via
google-genai with response_schema = ModelExtraction, JSON mime type,
temperature 0 and no tools; schema-invalid output raises ModelOutputError.
build_model_client needs VERTEX_PROJECT and credentials, loads ADC eagerly,
refuses API-key mode for Vertex, and allows the Gemini API only with
AI_ALLOW_GEMINI_API_DEV=true plus GEMINI_API_KEY.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Pipeline: PDF and EML end to end with recorded model responses, a model date
contradicting its own quote, the prompt-injection mail (the recorded model
obeys and invents a quote -> unverified), and the no-text PDF that skips the
model. API: bearer auth (401 + WWW-Authenticate), response shape, request id
handling, 400/413/415/422/429/502 mapping without echoing input, JSON logs
with IDs only. Contract: the committed OpenAPI file equals the app's schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…rified facts

Covers the fail-closed startup, the env vars, the API and component schema
names, the locator scheme, the verifier rules and the injection limitation,
the logging policy, the recorded-response testing approach, what was verified
(docling EMAIL/.msg coverage, the model-free PDF path, the heron layout
model, the Vertex eu endpoint in the SDK source, install size) and what was
not (live call, image build, OCR/tables, .msg), plus dependency licenses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…y hyphenation

A model that quotes the injected sentence itself passes grounding by
construction (provenance, not intent); a test now pins this so the verifier
is never mistaken for an injection defence. The README now says that
line-break hyphenation only affects quotes, because parsed segments are
single lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ift test; docs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…heck, lock, streaming

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
- pg-boss client and queue install move to src/db (job-queue.ts); least-privilege grants; depcruise rule
- upload: Content-Length required (411), request cap (413), OOXML structure check (macros, foreign ZIPs, zip bombs)
- duplicate detection serialised per company (advisory lock); orphaned objects logged; streaming download
- composite same-company FKs declared in the Drizzle schema (migration 0005, fresh-DB tested)
- tests: job row proven inside the rolled-back transaction, audit rolled back, 411/413/too many files
- test helper: random IPv6 per call – no rate-limit bleed across test files
- docs: api.md limits, exceptions register (upload rate limit, .msg check)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…-db-connection rule

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
@Fluory Fluory mentioned this pull request Sep 23, 2026
7 tasks done
…tes (red)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ries

A verified value is returned normalised (dates as YYYY-MM-DD, text trimmed).
Text values must match the quote on word boundaries, so partial tokens are
unverified. A date quote holding several distinct dates caps the field at
uncertain (reason ambiguous_quote).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
The two existing dev-flag tests now set vertex_project=None in their setup
(make_settings defaults it); their assertions are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…are both set

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…quest id (red)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…inside request context

ExtractGuard (pure ASGI) rejects /v1/extract without a valid bearer token
(401), with a missing or non-numeric Content-Length (411, new error code
length_required) or with a declared length above the limit plus a 16 KiB
multipart allowance (413) before any body byte is read. Unexpected errors
are caught inside the request context: logged with requestId/documentId and
the exception type only, answered with requestId and X-Request-Id.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
AI_MAX_PDF_PAGES (default 50) rejects longer PDFs with 422
document_too_long before any page is parsed (model-free page count, also
before the layout model runs). With AI_PDF_PIPELINE=layout the converter
and its layout model are built in create_app, so a missing model stops
the service start (PdfPipelineInitError) instead of failing every request.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…p and native stderr

Fixtures are described as hand-written in the Vertex REST format (not
recorded). New env var AI_MAX_PDF_PAGES, 411/422 codes, dev-flag
ambiguity, layout startup check. docling-parse/qpdf stderr bypasses JSON
logging; no output observed, whether it can carry document text is unknown.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
New error codes length_required and document_too_long, 411 response,
reason ambiguous_quote.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Behind uvicorn --root-path the raw scope path carries the prefix, and the
guard would silently skip /v1/extract.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
… root path

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1

Fluory commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Frischer Review (unabhängiger Subagent) – Befunde und Auflösung

Review nach review-pr + ai-rag.md, security.md, api.md; der Reviewer hat zusätzlich eigene Sonden gegen Verifier, API und Parser laufen lassen (23 kaputte PDFs/EMLs → nur kontrollierte Parse-Fehler, nie ein Absturz). Fixes test-first (jeweils roter Test-Commit vor dem Fix), von einem Implementierungs-Subagenten im isolierten Worktree gebaut und hier übernommen.

# Befund Auflösung
Important Datumswerte nicht ISO (15.10.26 blieb found), Textwerte ungetrimmt behoben – bestätigte Werte kommen normalisiert zurück (YYYY-MM-DD, getrimmt); passt das Datum nicht zum Zitat → unverified
Important CI: Python-Schritte im 15-Minuten-Job könnten durch den torch-Download zu lang werden geprüft: der gesamte check-Lauf inkl. uv sync mit torch (CPU) dauerte 2 min 22 s (Run 35825667046) – kein Umbau nötig; wird beobachtet
Important (security) Multipart-Body wurde vor der Token-Prüfung gelesen (DoS für jeden, der den Dienst erreicht) behoben – ASGI-Middleware prüft vor dem ersten Body-Byte: Token (401), fehlende/ungültige Content-Length (411), zu groß (413); auch hinter einem Proxy-Pfadpräfix; Test mit einem Body, der beim Lesen fehlschlägt; realer Server: 1-TB-Anfrage ohne Token → 401 bei 0 übertragenen Bytes
Important 500 ohne Korrelations-ID behoben – Fehler werden im Request-Kontext gefangen; Log mit requestId/documentId und nur dem Fehlertyp; Antwort mit requestId + X-Request-Id; Tests
Note Textabgleich ohne Wortgrenzen ("G", "bau GmbH" = found); mehrere Daten im Zitat behoben – Wortgrenzen; mehrere verschiedene Daten im Zitat → höchstens uncertain (ambiguous_quote). Grenze: ein ganzes Wort wie "Max" aus „Max Mustermann" bleibt belegt (README)
Note (security) Dev-Flag + VERTEX_PROJECT gleichzeitig → Free Tier gewinnt still behoben – Start wird verweigert (fail-closed); Test
Note Layout-Pipeline ohne Modell scheitert erst pro Anfrage; keine Seitengrenze behoben – Modell wird beim Start gebaut (fail-closed); AI_MAX_PDF_PAGES (50) → 422 document_too_long
Note PR-Text nannte falsche Testdateien und „aufgezeichnete" Antworten; native stderr-Ausgabe behoben – PR-Text und README korrigiert: die Vertex-Antworten sind handgeschrieben im Vertex-Format, kein Live-Mitschnitt; ob native stderr-Ausgaben von docling-parse Dokumenttext enthalten können, ist unbekannt (README)

Vertrag und TS-Typen: nur additive Änderungen (length_required, document_too_long, ambiguous_quote, 411), neu generiert, Drift-Test grün. Lokal: ruff ✔, ruff format ✔, pyright 0 Fehler, pytest 199 passed / 1 skipped.

Hinweis: die Token-Prüfung ist jetzt Middleware (sicherheitsrelevant) → menschliche Freigabe nötig. Langsame Uploads (Slowloris) begrenzt nur ein Server-/Proxy-Timeout (README).


Generated by Claude Code

@Fluory
Fluory marked this pull request as ready for review September 23, 2026 06:47
@Fluory
Fluory changed the base branch from claude/feat-intake-upload-5 to main September 23, 2026 09:59
@Fluory
Fluory merged commit 8359fdb into main Sep 23, 2026
7 of 8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(ai-service): stateless extraction service with grounding verifier

2 participants