Last updated: 2026-08-03 (Asia/Taipei)
GET /healthz reports process liveness only.
GET /readyz reports authentication-source and transport readiness. Direct mode
checks local configuration only. Embedded-frontend mode also asks the local
relay for its cached authentication state; that relay refreshes its upstream
model check at most once per minute after success. A failed check expires after
three seconds and transient 502/503/504 checks receive one short retry. A ready
response is not a guarantee that the next private Web request will succeed; use GET /v1/models for a direct
authenticated catalog smoke test.
Example readiness body:
{
"ready": true,
"checks": {
"chatgpt_auth_configured": true,
"requirements_provider_configured": false
}
}GET /metrics exposes Prometheus-compatible text metrics. When GPTWEB_API_KEY is configured, the endpoint requires the same Bearer key as /v1/*.
The metric surface is low-cardinality and excludes request content and identity data.
Current metrics:
gptweb2api_process_start_time_secondsgptweb2api_process_uptime_secondsgptweb2api_go_goroutinesgptweb2api_http_requests_totalgptweb2api_http_requests_in_flightgptweb2api_http_request_duration_secondsgptweb2api_http_response_bytes_totalgptweb2api_http_responses_total
curl -H 'Authorization: Bearer local-key' http://127.0.0.1:8788/metricsGPTWEB_UPSTREAM_MAX_RETRIES controls direct conversation retries for
upstream 502, 503, and 504 responses. The default is 2; valid values are
0..5. Retries use 100/200/400 ms bounded backoff. A Retry-After value up to
two seconds is honored; a longer value is returned immediately to the client.
When a requirements broker is configured, challenge/forbidden retries request fresh broker headers for every attempt. The gateway does not generate local HTTP 429 responses. Upstream ChatGPT 429 responses are forwarded unchanged and are never auto-retried.
Ordinary frontend resources and conversation streams remain on Chromium's native request path; guarded short backend requests use fetch/fulfill. Long-running relay processes recycle stale browser contexts without replaying accepted work:
- a recoverable text-generation browser failure closes only the stale worker page and retries the request once on a fresh page;
- a terminal official-frontend
upstream_frontend_unavailable503 with no observed successful conversation response enters short throttled pacing and retries once on a clean page; no other post-submit class is automatically replayed; - an aborted client request immediately closes its in-flight worker page while retaining that scheduler slot until cleanup has completed;
- the relay proactively recycles after
GPTWEB_BROWSER_RECYCLE_REQUESTScompleted operations per worker (default100) or after RSS reachesGPTWEB_BROWSER_MAX_RSS_MIB(default768); GPTWEB_BROWSER_CONCURRENCYsets the physical page-worker ceiling (default3), whileGPTWEB_BROWSER_STABLE_CONCURRENCYsets healthy effective concurrency (default2); a conversation-ID key keeps continuations ordered;- one isolated upstream, operation-timeout, or internal failure uses short
throttled pacing without collapsing healthy capacity; a second terminal
failure inside one minute atomically reduces effective concurrency to
GPTWEB_BROWSER_DEGRADED_CONCURRENCY(default1); hard-degraded starts use the long cooldown and consecutive-success threshold; GPTWEB_BROWSER_MIN_WARM_WORKERS(default1) keeps a small idle footprint; burst pages scale up on demand and retire afterGPTWEB_BROWSER_IDLE_PAGE_SECONDS(default60);- browser
stream:trueresponses flush incrementally through the private loopback NDJSON transport;relay_streamingreports aggregate chunks, characters, completions, failures, and aborts without content; - current streaming DOM markers keep slow active generations out of idle-stall
classification, while separate start/empty/overall deadlines and
GPTWEB_BROWSER_MAX_OUTPUT_CHARSbound dead or pathological turns; - after the same idle interval, retained warm pages request one browser garbage collection per completed-use epoch without closing the authenticated page;
- the bounded scheduler rejects queue/input-budget overflow or queue expiry with HTTP 503 and operation expiry with HTTP 504 while retaining slot/key ownership through cleanup;
- the relay
/healthzresponse is a cached, non-blocking snapshot reporting authentication/frontend readiness, every active-operation age, worker and queue utilization, buffered-input usage,completed_operations,browser_recycles,resident_pages,idle_page_retirements,idle_garbage_collections, RSS/heap/external memory, effective/maximum concurrency, the output ceiling, and a credential-free rollingreliabilitysnapshot,relay_streaming, plusautomatic_frontend_retryscheduled/attempted/recovered/exhausted counters.
The default reliability target is 100 ppm over 10,000 counted outcomes. The
health snapshot reports slo_met: null until the whole rolling window exists;
one failure is within budget and two are over budget. Valid completed relay
operations and 5xx/timeout/overload outcomes count, while invalid input, caller
cancellation, and shutdown do not. Treat this as a local service objective and
adaptive control signal, not as a guarantee of the private ChatGPT Web upstream.
The rolling window and retry counters are durably stored at
GPTWEB_BROWSER_RELIABILITY_STATE_FILE (or inside the browser profile by
default). Updates are coalesced, fsynced, and atomically replaced. The file is
credential-free and contains no request or response data. Verify
/healthz.reliability_persistence.restored after a planned relay restart.
Input-validation failures are not retried. Image failures recycle the assigned page for the next request but are not replayed automatically. Temporary official-frontend failure notices without a successful conversation response fail fast with HTTP 503 instead of being returned as assistant text; the relay first gives the native UI a bounded self-recovery grace period, and then transparently retries that exact failure class once after throttled-mode backoff. A second 503 is returned to the caller. Queue pressure, unknown 5xx, timeouts, input failures, and image requests are not included in this retry.
Use the credential-safe concurrent smoke after changing worker or queue limits:
python .\scripts\concurrent_gateway_smoke.py --requests 12 --concurrency 3 --timeout 300
python .\scripts\concurrent_gateway_smoke.py --mode responses-stream --requests 6 --concurrency 3 --timeout 300The script reads only GPTWEB_API_KEY from the ignored .env, disables proxies
and redirects, sends unique exact-token requests to loopback, and prints only a
sanitized status/latency summary. Any failed or cross-wired response exits
non-zero.
Request logs are structured JSON and contain request ID, method, path, status, duration, and response byte count. Browser-backed requests propagate the same validated ID into the relay's active-operation snapshot and structured retry/failure events. Logs exclude query strings, request bodies, authorization data, cookies, account identifiers, and conversation identifiers.
SIGINT and SIGTERM trigger a bounded graceful shutdown. Existing requests or streams receive up to 15 seconds to finish.