Skip to content

feat(observability): correlated structured logs and full health - #45

Merged
Fluory merged 185 commits into
mainfrom
claude/feat-observability-28
Sep 23, 2026
Merged

Fluory merged 185 commits into
mainfrom
claude/feat-observability-28

Conversation

@Fluory

@Fluory Fluory commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Warum

Fixes #28 · Epic #18 · Basis main (#44 ist gemergt; vorher gestapelt auf claude/feat-duplicates-27).

Arbeitsstand

  • Ziel: Betrieb kann eine Anfrage durch Web, Worker und KI-Dienst verfolgen; /api/health meldet alles, wovon der Pilot abhängt.
  • Nicht-Ziele: zentrales Log-Shipping, Alerting.
  • Erledigt: pino-JSON-Logs mit festem, zur Laufzeit gefiltertem Schlüsselsatz; alle console.* in src/** über logEvent (ESLint no-console); request.received beim Upload; Test-Hilfe captureLogs; Health um KI-Dienst-Erreichbarkeit und Queue-Rückstau erweitert (informativ, 10 s gecacht); Test eines kompletten synthetischen Durchlaufs ohne Inhalte/personenbezogene Daten in Logs; Doku; frischer Security-Review eingearbeitet; verify:full grün auf diesem Head; CI check grün.
  • Offen: Merge (freigegeben durch den Orchestrator, letzter PR des Stapels).
  • Annahmen: Korrelationsschlüssel ist die Anfrage-ID.
  • Nächster kleinster Schritt: Merge nach grünem check gegen main; danach Folge-Issues feat(export): send reviewed line items to the ERP (contract v2) #46–chore(observability): align the AI service log format with the web/worker logs #51.

Was ist passiert (Klartext)

Alle Protokollzeilen von Web und Worker sind jetzt einheitlich strukturiert und enthalten nur Kennungen und Fehlercodes – nie Dokumentinhalte oder Namen; das wird nicht nur beim Programmieren, sondern auch beim Schreiben jeder Zeile erzwungen. Eine Anfrage lässt sich über ihre Kennung vom Hochladen über den Worker bis zum KI-Dienst verfolgen. Ein Test spielt eine komplette Anfrage vom Hochladen bis zum Export durch und prüft jede Protokollzeile. Die Gesundheitsseite zeigt zusätzlich, ob der KI-Dienst erreichbar ist und wie viele Aufträge warten – ohne dass der Web-Dienst deswegen als „krank" gilt; diese Zusatzangaben werden 10 Sekunden zwischengespeichert, damit häufige Abfragen nichts kosten.

Plan-Pflicht (SYSTEM.md §4)

  • Kein Auslöser – keine Modulgrenze, öffentliche API, Migration, Auth/Rechte, kein Zahlungs-/Daten-/Infrapfad, höchstens zwei Module, keine Architekturvarianten, umkehrbar
  • Auslöser zutreffend – Impact Manifest ausgefüllt (Plan vor Code)

Impact Manifest

  • Betroffene Module: observability (pino-Logger, captureLogs, Health-Erweiterung, cachedFor), db (countWaitingJobs und onError im pg-boss-Modul), extraction (pingAiService), intake/app (request.received, /api/health), ESLint-Konfiguration (no-console).
  • Schnittstellen / Datenänderungen: /api/health additiv: dependencies.aiService, backlog (informativ, ändern den HTTP-Status nie); neue Abhängigkeit pino 10.3.1 (MIT). Keine Migration.
  • Akzeptanzkriterien: alle aus feat(observability): correlated structured logs and full health #28 – siehe Nachweis.
  • Testplan: Unit Logger (Schlüssel-Filter), Health-Aggregation (Abhängigkeiten, Gauges, Timeouts, keine Fehlertexte), Cache; Integration Health gegen echte Dienste + lokalen KI-Stub + echten Job-Zähler; Integration kompletter Durchlauf mit Log-Prüfung.
  • Verifizierte Fakten: KI-Dienst loggt JSON mit requestId aus X-Request-Id (seit feat(ai-service): stateless extraction service with grounding verifier #6); pino steht auf Next.js' Liste der server-externen Pakete.
  • Offene Annahmen: Korrelationsschlüssel ist die Anfrage-ID (keine eigene HTTP-Request-ID im Web).
  • Nicht-Ziele: siehe oben.
  • Risiken und Rollback: neue Laufzeitabhängigkeit; /api/health ist öffentlich (vor öffentlicher Bereitstellung hinter Proxy/Auth, in api.md dokumentiert); Rollback per Revert.

Geändert

  • src/features/observability/{log,health,index}.ts (+ Tests), src/db/job-queue-client.ts, src/features/extraction/{ai-client,index}.ts, src/app/api/health/route.ts, Upload-/Intake-/Worker-Einstiege (Logzeilen), eslint.config.mjs
  • package.json/pnpm-lock.yaml (pino)
  • Tests: tests/integration/{health,log-privacy,processing}.test.ts
  • Doku: docs/technical/{operations,api,architecture}.md, CHANGELOG.md

Nachweis (SYSTEM.md §11)

  • verify:changed: Unit 152/152 (inkl. Logger-/Cache-Tests), Integration Health/Log-Privacy/Upload/Intake 20/20
  • verify: grün auf diesem Head (09af010, frisch migrierte DB): lint, typecheck, Unit 154/154, Integration 114/114, depcruise (keine Verstöße), build, pnpm audit --audit-level high (1 moderate, kein high); AI-Service ruff/format/pyright grün, pytest 423 passed / 2 skipped; CI check grün
  • verify:full / E2E-Spec: grün – review-smoke und line-items-smoke, Eval-Gate „PASSED (replay, 15 cases, threshold 5 points)"
  • Akzeptanzkriterien feat(observability): correlated structured logs and full health #28: JSON-Logs mit requestId/jobId/companyId (pino; KI-Dienst JSON) → log.ts + log.test.ts + Log-Privacy-Test prüft Schlüsselsatz; Korrelation Web → Worker → KI (X-Request-Id) → request.received, Anfrage-ID im Job, X-Request-Id im AI-Client, KI-Dienst loggt sie; Health mit Rückstau + KI-Erreichbarkeit → Unit + Integration Health; Test: kompletter synthetischer Lauf ohne Inhalte/personenbezogene Daten → log-privacy.test.ts
  • Manueller Prüfnachweis: –
  • Frischer Review: erledigt, inkl. Security-Regel (0 Blocker · 4 should-fix · 3 nits, alle eingearbeitet oder begründet) – siehe PR-Kommentar

Doku-Entscheidung (genau eine)

  • Keine langlebige Doku betroffen – Begründung: –
  • Doku betroffen und im selben PR aktualisiert:
    • Produktdoku (P0: README): –
    • Technische Doku: docs/technical/operations.md, docs/technical/api.md
    • Architekturkarte: docs/technical/architecture.md
    • ADR: –
    • CHANGELOG [Unreleased] (sichtbares Feature oder Verhalten – im selben PR, nie „später")

Entferntes oder Umbenanntes: console-basierte Logzeilen durch pino ersetzt (gleiche Schlüssel)

Dateigrößen und neue Bausteine (SYSTEM.md §7)

Dateien über 500 Zeilen im Diff (Ausnahmen: generierter Code, Lockfiles, Fixtures, Migrationen, Schemas, Ressourcen, Doku, Konfiguration):

  • keine
  • bewusst belassen – Begründung: –
  • im selben PR nach fachlicher Verantwortung geteilt
  • Folge-Issue –

Über 800 Zeilen mit neuer Fachlogik oder über 1000 Zeilen (P1/P2): nicht betroffen

Neue Shared-Komponente, Utility-Datei, Adapter oder fachlicher Service:

  • nein
  • ja – gesucht nach: vorhandenem Logger/Health; gefunden: logEvent, runHealthChecks – erweitert, API gleich

Subagent-Einsätze

  • Frischer Security-Review: unabhängiger read-only Subagent.

Risiken / offene Punkte


🤖 Generated with Claude Code

https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1

…ion access control

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…LS with integration proofs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…boss queues

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…nd download routes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…-only audit, composite FK

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
- sign-up requires the invitation id (link) plus the invited e-mail – no takeover by address alone
- one company per user (unique index), deterministic membership lookup, actor from membership
- organization plugin accepts only admin/clerk roles
- configurable client-IP source for the auth rate limit; local secret refused in production
- invite page: zod input, 404 for clerks, shows the invitation link; signup needs the link
- tests: wrong/missing invitation id, foreign set-active/list-members, last admin, roles
- docs: operations (rate limit/proxy, recovery), data model, exceptions register (admin plugin)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…e rejection

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
… docs for #5

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…secret

The image sets NODE_ENV=production, so the previous check blocked the local compose stack.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…re script

Python 3.13 package requestflow_ai (src layout), docling/fastapi/google-genai
pinned, CPU-only torch via the PyTorch CPU index, dev tools ruff/pyright/pytest.
The synthetic PDF fixture is generated by scripts/make_fixtures.py (reportlab).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Test-first per ADR-0001 D8: quote normalisation (whitespace, case, NFKC,
hyphenation), German number and date formats, and the verifier rules that
turn unsupported model claims into unverified. Also adds the segment and
model-output types the tests build on; the verifier does not exist yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
normalize_text folds NFKC, soft hyphens, line-break hyphenation, typographic
dashes/quotes, whitespace and case. values parses German/ISO numbers and
DD.MM.YYYY, D.M.YY and ISO dates. verify_field only keeps or downgrades the
model's status: a quote not in the cited segment, an unknown segment, a value
inconsistent with the quote, found without evidence or value, and missing
with a value all become unverified with a reason.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ction (red)

EML: body lines with 1-based line locators, From/Subject header segments,
RFC 2047 and quoted-printable decoding, HTML-only bodies, header newline
collapse, no attachment payloads. PDF: textline segments with page + top-left
bbox via docling-parse (model-free); the layout-pipeline test only runs with
AI_TEST_DOCLING_MODELS=1. Detection by magic bytes; .msg rejected for now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
EML via the standard library email package (policy=default): From/Subject
header segments and one segment per non-empty body line, HTML-only bodies
reduced to text lines. docling's EMAIL backend was checked but emits
paragraphs without provenance, so it cannot give line locators.

PDF via docling: the default textlines pipeline reads docling-parse text
lines (page + top-left bbox, no ML models, no network); the opt-in layout
pipeline uses DocumentConverter (OCR and tables off) and the heron layout
model. docling is imported lazily.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ex model client (red)

The model client is exercised through the real google-genai SDK with an
httpx MockTransport replaying recorded generateContent bodies: eu multi-region
URL, bearer auth, structured-output config, token usage, fail-closed init
(no project, no credentials, dev flag without key), no silent switch to API
key mode, and schema-invalid output. Adds the env-based Settings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…lient

Prompt extract_header_v1.md (system instruction) plus render_document, which
puts segment-id-prefixed lines between <document> delimiters and neutralises
delimiter-like tags inside the document. GeminiModelClient calls Vertex via
google-genai with response_schema = ModelExtraction, JSON mime type,
temperature 0 and no tools; schema-invalid output raises ModelOutputError.
build_model_client needs VERTEX_PROJECT and credentials, loads ADC eagerly,
refuses API-key mode for Vertex, and allows the Gemini API only with
AI_ALLOW_GEMINI_API_DEV=true plus GEMINI_API_KEY.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Pipeline: PDF and EML end to end with recorded model responses, a model date
contradicting its own quote, the prompt-injection mail (the recorded model
obeys and invents a quote -> unverified), and the no-text PDF that skips the
model. API: bearer auth (401 + WWW-Authenticate), response shape, request id
handling, 400/413/415/422/429/502 mapping without echoing input, JSON logs
with IDs only. Contract: the committed OpenAPI file equals the app's schema.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
Per-file caps multiplied across the attachments of a .msg. ParseBudget
(parsing/budget.py) is created once per uploaded document and shared by all
nested attachments: 100 attachments, 128 MiB unzipped OOXML, 100 PDF pages
(at least AI_MAX_PDF_PAGES), MAX_OCR_PAGES OCR pages. An attachment that would
overdraw it is reported with the additive error code budget_exceeded; the rest
of the message is still returned. BudgetExceededError subclasses
DocumentTooLongError, so a top-level document still maps to 422.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…EADME limits for the security fixes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ng requests show a cause, export attempts for export errors), export retries record the next attempt, reprocess buttons named per request (#26 review)
…ges_skipped

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ejected duplicate carries no stale error, list shows the decision; tests for ERROR branches and the skipped job (#27 review)
…elist, config variable names only), no-console lint rule, web-side request.received, pg-boss errors as class only, cached informational health (#28 review)

Fluory commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Frischer Security-Review (unabhängiger Subagent, read-only aus Git-Refs)

Ergebnis: 0 Blocker · 4 should-fix · 3 nits. Bestätigt: pino ohne pid/hostname, synchron auf stdout; countWaitingJobs parametrisiert; pingAiService ohne SSRF (URL nur aus Konfiguration, 2 s Timeout); Health-Report nur Namen; KI-Dienst nimmt X-Request-Id nur nach striktem Muster.

# Schwere Befund Auflösung
1 should-fix Web-Seite loggte beim Upload nichts → Korrelation „web → worker" nur halb submitUpload loggt request.received (requestId, companyId, Anzahl); Log-Privacy-Test verlangt request.received, job.processed, request.exported
2 should-fix verbliebene console.*-Aufrufe umgingen pino und den Test (u. a. rohe pg-boss-Fehlermeldung) alle über logEvent (Upload-Route, Intake-Waise, Worker-Start, Health-Konfig); pg-boss-Fehler nur als Klasse (onError aus der Composition Root); ESLint no-console für src/** (Ausnahme: Operator-Skripte seed.ts, setup.ts, begründet)
3 should-fix Schlüsselsatz nur per TypeScript erzwungen pick() filtert zur Laufzeit gegen eine feste Liste; Unit-Test mit „einschmuggelten" Feldern; neues Feld names nur für Konfigurations-Variablennamen
4 should-fix /api/health öffentlich, jetzt teurer (Zählungen + ausgehender Ping), Rückstau zeigt Gesamtaktivität informative Teile 10 s gecacht (cachedFor, Unit-Test inkl. Fehler-Retry); Kosten + Risiko in api.md dokumentiert (vor öffentlicher Bereitstellung hinter Proxy/Auth)
5 nit captureLogs im öffentlichen Modul-Index bleibt (Tests brauchen es; Architekturregel verbietet Deep-Imports) – dokumentiert als Test-Hilfe
6 nit KI-Dienst nutzt ts/Großbuchstaben-Level bleibt; in operations.md als Formatunterschied zu vermerken bei Log-Shipping (Folgeaufgabe)
7 nit event ist freier String bleibt; alle Aufrufer nutzen Konstanten

Danach: Unit 152/152 (inkl. neuer Logger-/Cache-Tests), Integration Health/Log-Privacy/Upload/Intake 20/20.


Generated by Claude Code

olefile follows the declared FAT sector count along the DIFAT chain in its
constructor (quadratic array copy per sector); a forged count on a looped
DIFAT chain hung it before the stream-size walk could run. The header's FAT,
DIFAT, mini FAT and directory sector counts must now fit the file size.
Also: stale docstring, README package list.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DJ5vaKvTYiMvdngT4d3xo1
…ored per document and shown in the review; contract types regenerated (#23 review)
…e/feat-review-line-items-25

# Conflicts:
#	src/app/requests/[id]/page.tsx
#	src/features/review/review.ts
#	tests/integration/review.test.ts
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(observability): correlated structured logs and full health

2 participants