fix(ai-service): retry transient model errors within one extraction call - #72
Merged
Merged
Conversation
Symptom: a showcase request stayed "In Verarbeitung" with ai.unavailable. Cause: the Gemini API free tier answers about one call in three with 503 UNAVAILABLE (high demand), measured 2026-09-27; one 503 failed the whole extraction and left the job to the queue backoff. Fix: the SDK retries 429/503 up to 3 attempts in total with exponential backoff and jitter; every other status still fails at once, and the contract after the last attempt is unchanged (502 model_error). Fixes #69 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Refs #69 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
4 tasks
6 of 15 tasks
Owner
Author
Frischer Review (/review-pr) – unabhängiger Reviewer-Agent ohne Umsetzungskontext, 2026-09-27Review PR #72 (fix(ai-service): retry transient model errors), Kopf 38e9ff3 Übrige Punkte ohne Befund:
Urteil: changes requested. Wiederholte Timeouts sprengen das Zeitbudget des AC und den Worker-Timeout. Außerdem sind README und CHANGELOG nicht in diesem PR nachgezogen. |
… into claude/fix-ai-model-retry-69
The fresh review of #72 showed that the SDK's HttpRetryOptions also retry timeouts and connection errors, whatever status codes are set: a hanging model could take three full timeouts (~185 s) while the worker gave up after 60 s and pg-boss started the job again - duplicate model calls and cost. A small retry loop now repeats only 429/503, never a timeout, gives each attempt only the budget that is left and skips a retry when less than 10 s remain. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
README still said the SDK does not retry and lacked the two new settings; the CHANGELOG entry for #69 lived in the stacked #73 instead of the PR that changes the behaviour. The showcase runbook now sets the model budget below the web app's AI timeout so the worker never gives up while a model call is still running. The initial delay gets an upper bound so a misconfiguration cannot silently exceed the budget. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Fluory
marked this pull request as ready for review
September 27, 2026 10:27
Fluory
changed the base branch from
claude/chore-showcase-deploy-67
to
main
September 27, 2026 10:30
Owner
Author
|
Merge auf ausdrückliche Anweisung des Orchestrators (2026-09-27: „Merge und deploy“). |
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Warum
Fixes #69 · gestapelt auf #68 (Showcase-Deployment)
Arbeitsstand
AI_MODEL_TIMEOUT_SECONDS=45im Vercel-Projektrequestflow-aigesetzt (wirkt ab dem nächsten Deploy).requestflow-aiausc0f1690in Produktion deployen und drei Test-Anfragen im Showcase hochladen.Was ist passiert (Klartext)
Im Showcase blieb eine Anfrage „In Verarbeitung“ hängen. Gemessen war die Ursache eindeutig: Googles Gemini-Free-Tier antwortet bei etwa jedem dritten Aufruf mit „Modell überlastet“. Bisher ließ schon eine einzige solche Antwort die ganze Extraktion scheitern. Jetzt versucht der KI-Dienst es bei genau diesen vorübergehenden Fehlern (Überlast 503, Rate-Limit 429) bis zu zweimal erneut, mit kurzer, wachsender Pause. Dabei gilt ein festes Zeitbudget: Alle Versuche zusammen dauern nie länger als die bisherige Zeitgrenze eines Modellaufrufs. Hängt das Modell (Zeitüberschreitung), gibt es keinen zweiten Versuch im selben Aufruf, denn die Zeit ist dann schon verbraucht. Die Anfrage wird später über die Warteschlange erneut verarbeitet, wie bisher. Andere Fehler, etwa ein falscher Schlüssel oder ein ungültiges Modell, scheitern sofort, damit echte Konfigurationsfehler nicht hinter Wartezeiten verschwinden.
Plan-Pflicht (SYSTEM.md §4)
Geändert
services/ai/src/requestflow_ai/extraction/model_client.py:RetryPolicyund Wiederholungsschleife imGeminiModelClient– nurAPIError429/503, höchstensAI_MODEL_RETRY_ATTEMPTSVersuche, Pause ab 1 s verdoppelt (max. 8 s, bis 25 % Jitter); jeder Versuch bekommt nur das Restbudget als Timeout; kein Retry mit weniger als 10 s Rest; Timeouts und Netzwerkfehler nie. Der SDK-Retry (HttpRetryOptions) ist entfernt, weil er Timeouts unabhängig von den Statuscodes wiederholt (google/genai/_api_client.py:578–581).services/ai/src/requestflow_ai/config.py:AI_MODEL_RETRY_ATTEMPTS(Standard 3, 1–5),AI_MODEL_RETRY_INITIAL_DELAY_SECONDS(Standard 1, ≤ 5);AI_MODEL_TIMEOUT_SECONDSist jetzt das Budget inklusive Wiederholungenservices/ai/tests/test_model_retry.py: 503 → Erfolg (2 Aufrufe), 429 → Erfolg, Dauer-503 → Fehler nach 3 Aufrufen, 400/401/403/404/500 → genau 1 Aufruf, Timeout → genau 1 Aufruf, zu wenig Restbudget → 1 Aufruf, zweiter Versuch mit kleinerem Timeout als der erste, Gemini-API-Weg wiederholt ebenfallsservices/ai/README.md: Budget-Semantik und die beiden neuen Variablendocs/technical/deployment-vercel.md:AI_MODEL_TIMEOUT_SECONDS=45für den Showcase (unter dem 60-s-Timeout der Web-App abzüglich Parsing)CHANGELOG.md: Fixed-Eintrag für fix(ai-service): retry transient model errors within one extraction call #69 (vorher fälschlich in fix(requests): same processing state in list and detail, retries picked up on page view #73)Nachweis (SYSTEM.md §11)
verify:changed: grün –pytest tests/test_model_retry.py12/12; zuerst rot: Timeout-Test mit 3 statt 1 Aufruf, Budget-Test mit 2 statt 1 Aufruf, Timeout pro Versuch unverändert 30,0 sverify: grün lokal – AI-Serviceruff check,ruff format --check,pyrightohne Befund,pytest452 bestanden, 2 übersprungen (Modell-Tests, feat(evals): run the scanned-PDF eval cases with OCR in CI #49); AI-Eval-Gate (Replay, 15 Fälle) bestanden; TS-Seite unverändert; CI am PRverify:full/ E2E-Spec: nicht betroffenDoku-Entscheidung (genau eine)
services/ai/README.md,docs/technical/deployment-vercel.md[Unreleased](sichtbares Feature oder Verhalten – im selben PR, nie „später"): Fixed-Eintrag fix(ai-service): retry transient model errors within one extraction call #69Entferntes oder Umbenanntes:
docs/+ README gegrept, Treffer bereinigt: „The SDK does not retry“ inservices/ai/README.mdersetztDateigrößen und neue Bausteine (SYSTEM.md §7)
Dateien über 500 Zeilen im Diff (Ausnahmen: generierter Code, Lockfiles, Fixtures, Migrationen, Schemas, Ressourcen, Doku, Konfiguration):
Über 800 Zeilen mit neuer Fachlogik oder über 1000 Zeilen (P1/P2): nicht betroffen
Neue Shared-Komponente, Utility-Datei, Adapter oder fachlicher Service:
Subagent-Einsätze
Risiken / offene Punkte
AI_MODEL_TIMEOUT_SECONDS, Showcase 45 s, lokal 60 s bei 120 s Worker-Timeout) statt bis zu 3 × Timeout.🤖 Generated with Claude Code