Last updated: 2026-08-02 (Asia/Taipei)
gptweb2api implements only behavior that can be represented faithfully by the normal ChatGPT Web transport. A JSON field being part of the official OpenAI API does not mean it is silently ignored here: meaningful unsupported values return HTTP 400 with type: invalid_request_error, param, and code.
Every response includes X-Request-ID for correlation. Browser-backed requests
carry that same validated ID through the local relay; it is never sent as user
content. Upstream ChatGPT conversation identifiers remain available through the
documented X-GPTWeb-* headers.
| Endpoint | Status | Notes |
|---|---|---|
GET /healthz |
Supported | Local process health only. |
GET /v1/models |
Supported | Returns the authenticated account's real ChatGPT Web model catalog. No synthetic fallback list. |
GET /v1/models/{model} |
Supported | Returns a catalog model or model_not_found. |
POST /v1/completions |
Supported single-prompt text subset | Legacy streaming/non-streaming text for one string prompt and one choice. |
POST /v1/chat/completions |
Supported text + image-input subset | Streaming/non-streaming text output, image input, and local store:true. |
| Stored Chat Completion list/retrieve/update/delete/messages | Supported local subset | Independent local records; metadata values are strings. |
POST /v1/vision/analyze |
Supported convenience extension | Multipart one-to-ten images plus prompt to Chat Completion text/SSE. |
POST /v1/responses |
Supported text, image, and local-file subset | Text output, image/file input, streaming, previous_response_id, and local Conversations. |
POST /v1/images/generations |
Supported single-image subset | prompt, optional model, n=1, and base64 output. |
POST /v1/images/edits |
Supported image-edit subset | One to ten PNG/JPEG/GIF/WebP inputs, prompt, optional model, and one base64 output. |
POST /v1/images/variations |
Supported compatibility subset | One image to one base64 output using a fixed Web variation instruction. |
GET /v1/responses/{response_id} |
Supported for stored text Responses | Returns the completed local Response object. |
DELETE /v1/responses/{response_id} |
Supported for stored text Responses | Deletes the Response and its continuation mapping. |
GET /v1/responses/{response_id}/input_items |
Supported text subset | Stored message inputs with after, limit, and order pagination. |
GET /v1/responses/{response_id}/input_items/{input_item_id} |
Supported text subset | Retrieves one stored message input. |
POST/GET /v1/files |
Supported local storage subset | Bounded 25 MiB-per-file local store, listing and cursor pagination. |
GET/DELETE /v1/files/{file_id} |
Supported local storage subset | Metadata retrieval and deletion. |
GET /v1/files/{file_id}/content |
Supported local storage subset | Returns original bytes. |
| Conversations and item CRUD | Supported local subset | Durable local message context for Responses; not Platform cloud storage. |
| Upload create/part/complete/cancel | Supported local subset | One-hour uploads, 25 MiB final limit, disk-streamed parts. |
| Batch create/list/retrieve/cancel | Supported local subset | JSONL local Files, supported text endpoints, one sequential worker. |
| Responses cancel | Not applicable to synchronous responses | Background Responses are not implemented. |
| Remote Platform storage, embeddings, audio, fine-tuning, Realtime, Assistants/vector stores | Not implemented | Local storage is explicitly local and does not impersonate these backends. |
Unknown URLs and incorrect methods return JSON invalid_request_error responses with not_found or method_not_allowed, plus X-Request-ID.
On Windows, the optional persistent embedded frontend can carry text, image-input,
and local-file conversation requests through ChatGPT's official frontend lifecycle using a
local OAuth credential file. It supports auto and explicit catalog model IDs;
see BROWSER_RELAY.md. This changes the transport, not the
API feature matrix below.
model- one non-empty string
prompt streamstore:true|false; true stores the request/final response locally- string-to-string
metadatawith up to 16 entries - explicit Web continuation through the documented
X-GPTWeb-*headers
n: 1best_of: 1echo: falselogprobs: 0- zero frequency/presence penalties
- empty
logit_biasand stop values stream_options: {"include_usage":false}
Token-array or multi-prompt input, exact token limits, suffixes, sampling controls, non-empty stop sequences, and usage/logprob claims are rejected.
modelmessageswithdeveloper,system,user, andassistanttext- user-message
image_urlparts using base64 PNG/JPEG/GIF/WebP data URLs stream- basic text content arrays
- explicit Web continuation through
X-GPTWeb-Conversation-IDplus the previous assistant/parent message ID
Without both continuation headers, every Chat Completions request starts a new Web conversation with history/training disabled. Replayed API messages are carried as isolated system context plus the current user turn; they do not silently attach to the previously used browser conversation.
Fresh requests also receive a compact, lowest-priority stateless/output-only
baseline so Web answers behave more like direct API payloads. Caller
system/developer instructions follow that baseline and remain authoritative.
Set GPTWEB_FIDELITY_MODE=chatgpt to retain the native conversational style.
These values are accepted because they do not claim unsupported behavior:
n: 1tools: []functions: []tool_choice: "none"or"auto"when no tools are suppliedfunction_call: "none"or"auto"when no functions are suppliedparallel_tool_calls: true|falsewhen no tools are suppliedmodalities: []or["text"]response_format: {"type":"text"}stream_options: {"include_usage":false}logprobs: falsetop_logprobs: 0- zero frequency/presence penalties
- empty logit-bias and stop values
- custom tools/functions and tool calls
- JSON mode / JSON Schema structured outputs
- audio output and non-text modalities
- plain HTTP image URLs; remote image fetching requires HTTPS, public DNS/IPs, port 443, at most three redirects, supported image bytes, and the normal attachment size limits
- exact token limits or token usage
- sampling controls such as
temperature,top_p, andseed - non-empty stop sequences
- log probabilities
- service tiers, prompt caching, prediction, verbosity, and web-search options
- API reasoning-effort values; ChatGPT Web exposes a different model-specific effort catalog and the gateway does not guess a cross-product mapping
model- string
input - message-array
inputwith text content input_imagein user messages or top-level input items using base64 data URLsinput_file.file_idbacked by the local Files store and uploaded to the official frontend for that request- string
instructionson fresh responses streamprevious_response_idbacked by a bounded local continuation map persisted atomically across gateway restartsstore: true|false(defaulttrue), with retrieve/delete for stored responses- stored input-item listing with cursor pagination and single-item retrieval
- local
conversationID/reference, mutually exclusive withprevious_response_id - plain text configuration through
text.format.type = "text" - text-only modalities
background: falsetools: []include: []- empty metadata
tool_choice: "none"or"auto"when no tools are suppliedparallel_tool_calls: true|falsewhen no tools are suppliedmodalities/output_modalitiesset to text only- empty
stream_options top_logprobs: 0
- tools, function calls, computer use, file search, and hosted tools
- JSON Schema / structured text output
- background Responses
- maximum output/tool-call limits
- sampling and logprob controls
- prompt templates and prompt caching
- reasoning configuration and encrypted reasoning items
- service tiers, truncation, safety identifiers, and remote OpenAI Platform storage semantics
- Browser-backed
stream:trueis genuinely incremental: the relay forwards append-only rendered/network deltas and the gateway flushes each standard SSE event before browser completion. A small uncommitted tail permits frontend rewrites, and the terminal response is prefix-reconciled to prevent duplicate text. - Exact OpenAI Platform token accounting is not available from ChatGPT Web, so the gateway does not fabricate
usage. - Account-specific specialty quotas and safe reasoning progress may appear in the optional top-level
gptwebextension. - The resolved ChatGPT Web model is returned in
modelwhen the stream identifies it. - Response IDs map to upstream conversation/message IDs in the local
GPTWEB_RESPONSE_STATE_FILEfor 30 days (up to 4096 entries). With the defaultstore: true, completed response objects, generated output text, and normalized text input items are also retained for retrieve/delete/input-items. The file never contains OAuth tokens or cookies. Usestore:falsewhen local prompt/output persistence is not desired. - Internal/hidden assistant-to-tool traffic is filtered and is not exposed as user-visible output.
- Unknown fields and meaningful unsupported values on both Chat Completions and Responses return HTTP 400 with
type,param, andcode. - Empty bodies, malformed JSON, multiple top-level JSON values, and oversized bodies return explicit JSON request errors.
- Missing account models return HTTP 404 with
code: model_not_foundandparam: model. - Upstream authentication, Cloudflare Challenge Page, generic forbidden, rate-limit, request-size, and not-found failures are classified separately without returning arbitrary upstream HTML.
- Direct conversations retry transient 502/503/504 failures up to
GPTWEB_UPSTREAM_MAX_RETRIEStimes; 429 responses are never auto-retried.
- Startup rejects malformed numeric configuration values.
- Non-loopback listening requires
GPTWEB_API_KEY. - The server handles SIGINT/SIGTERM with a bounded graceful shutdown.
- Request logs are structured JSON and include only request ID, method, path, status, duration, and response size.
- Logs do not include query strings, request bodies, Authorization headers, cookies, OAuth tokens, or ChatGPT account IDs.