Skip to content

Latest commit

 

History

History
148 lines (121 loc) · 7.47 KB

File metadata and controls

148 lines (121 loc) · 7.47 KB

Operations

Last updated: 2026-08-03 (Asia/Taipei)

Health and readiness

GET /healthz reports process liveness only.

GET /readyz reports authentication-source and transport readiness. Direct mode checks local configuration only. Embedded-frontend mode also asks the local relay for its cached authentication state; that relay refreshes its upstream model check at most once per minute after success. A failed check expires after three seconds and transient 502/503/504 checks receive one short retry. A ready response is not a guarantee that the next private Web request will succeed; use GET /v1/models for a direct authenticated catalog smoke test.

Example readiness body:

{
  "ready": true,
  "checks": {
    "chatgpt_auth_configured": true,
    "requirements_provider_configured": false
  }
}

Metrics

GET /metrics exposes Prometheus-compatible text metrics. When GPTWEB_API_KEY is configured, the endpoint requires the same Bearer key as /v1/*.

The metric surface is low-cardinality and excludes request content and identity data.

Current metrics:

  • gptweb2api_process_start_time_seconds
  • gptweb2api_process_uptime_seconds
  • gptweb2api_go_goroutines
  • gptweb2api_http_requests_total
  • gptweb2api_http_requests_in_flight
  • gptweb2api_http_request_duration_seconds
  • gptweb2api_http_response_bytes_total
  • gptweb2api_http_responses_total
curl -H 'Authorization: Bearer local-key' http://127.0.0.1:8788/metrics

Transient upstream retries

GPTWEB_UPSTREAM_MAX_RETRIES controls direct conversation retries for upstream 502, 503, and 504 responses. The default is 2; valid values are 0..5. Retries use 100/200/400 ms bounded backoff. A Retry-After value up to two seconds is honored; a longer value is returned immediately to the client.

When a requirements broker is configured, challenge/forbidden retries request fresh broker headers for every attempt. The gateway does not generate local HTTP 429 responses. Upstream ChatGPT 429 responses are forwarded unchanged and are never auto-retried.

Embedded browser relay recovery

Ordinary frontend resources and conversation streams remain on Chromium's native request path; guarded short backend requests use fetch/fulfill. Long-running relay processes recycle stale browser contexts without replaying accepted work:

  • a recoverable text-generation browser failure closes only the stale worker page and retries the request once on a fresh page;
  • a terminal official-frontend upstream_frontend_unavailable 503 with no observed successful conversation response enters short throttled pacing and retries once on a clean page; no other post-submit class is automatically replayed;
  • an aborted client request immediately closes its in-flight worker page while retaining that scheduler slot until cleanup has completed;
  • the relay proactively recycles after GPTWEB_BROWSER_RECYCLE_REQUESTS completed operations per worker (default 100) or after RSS reaches GPTWEB_BROWSER_MAX_RSS_MIB (default 768);
  • GPTWEB_BROWSER_CONCURRENCY sets the physical page-worker ceiling (default 3), while GPTWEB_BROWSER_STABLE_CONCURRENCY sets healthy effective concurrency (default 2); a conversation-ID key keeps continuations ordered;
  • one isolated upstream, operation-timeout, or internal failure uses short throttled pacing without collapsing healthy capacity; a second terminal failure inside one minute atomically reduces effective concurrency to GPTWEB_BROWSER_DEGRADED_CONCURRENCY (default 1); hard-degraded starts use the long cooldown and consecutive-success threshold;
  • GPTWEB_BROWSER_MIN_WARM_WORKERS (default 1) keeps a small idle footprint; burst pages scale up on demand and retire after GPTWEB_BROWSER_IDLE_PAGE_SECONDS (default 60);
  • browser stream:true responses flush incrementally through the private loopback NDJSON transport; relay_streaming reports aggregate chunks, characters, completions, failures, and aborts without content;
  • current streaming DOM markers keep slow active generations out of idle-stall classification, while separate start/empty/overall deadlines and GPTWEB_BROWSER_MAX_OUTPUT_CHARS bound dead or pathological turns;
  • after the same idle interval, retained warm pages request one browser garbage collection per completed-use epoch without closing the authenticated page;
  • the bounded scheduler rejects queue/input-budget overflow or queue expiry with HTTP 503 and operation expiry with HTTP 504 while retaining slot/key ownership through cleanup;
  • the relay /healthz response is a cached, non-blocking snapshot reporting authentication/frontend readiness, every active-operation age, worker and queue utilization, buffered-input usage, completed_operations, browser_recycles, resident_pages, idle_page_retirements, idle_garbage_collections, RSS/heap/external memory, effective/maximum concurrency, the output ceiling, and a credential-free rolling reliability snapshot, relay_streaming, plus automatic_frontend_retry scheduled/attempted/recovered/exhausted counters.

The default reliability target is 100 ppm over 10,000 counted outcomes. The health snapshot reports slo_met: null until the whole rolling window exists; one failure is within budget and two are over budget. Valid completed relay operations and 5xx/timeout/overload outcomes count, while invalid input, caller cancellation, and shutdown do not. Treat this as a local service objective and adaptive control signal, not as a guarantee of the private ChatGPT Web upstream. The rolling window and retry counters are durably stored at GPTWEB_BROWSER_RELIABILITY_STATE_FILE (or inside the browser profile by default). Updates are coalesced, fsynced, and atomically replaced. The file is credential-free and contains no request or response data. Verify /healthz.reliability_persistence.restored after a planned relay restart.

Input-validation failures are not retried. Image failures recycle the assigned page for the next request but are not replayed automatically. Temporary official-frontend failure notices without a successful conversation response fail fast with HTTP 503 instead of being returned as assistant text; the relay first gives the native UI a bounded self-recovery grace period, and then transparently retries that exact failure class once after throttled-mode backoff. A second 503 is returned to the caller. Queue pressure, unknown 5xx, timeouts, input failures, and image requests are not included in this retry.

Use the credential-safe concurrent smoke after changing worker or queue limits:

python .\scripts\concurrent_gateway_smoke.py --requests 12 --concurrency 3 --timeout 300
python .\scripts\concurrent_gateway_smoke.py --mode responses-stream --requests 6 --concurrency 3 --timeout 300

The script reads only GPTWEB_API_KEY from the ignored .env, disables proxies and redirects, sends unique exact-token requests to loopback, and prints only a sanitized status/latency summary. Any failed or cross-wired response exits non-zero.

Logging

Request logs are structured JSON and contain request ID, method, path, status, duration, and response byte count. Browser-backed requests propagate the same validated ID into the relay's active-operation snapshot and structured retry/failure events. Logs exclude query strings, request bodies, authorization data, cookies, account identifiers, and conversation identifiers.

Shutdown

SIGINT and SIGTERM trigger a bounded graceful shutdown. Existing requests or streams receive up to 15 seconds to finish.