Skip to content

Latest commit

 

History

History
198 lines (162 loc) · 10.8 KB

File metadata and controls

198 lines (162 loc) · 10.8 KB

OpenAI API Compatibility

Last updated: 2026-08-02 (Asia/Taipei)

gptweb2api implements only behavior that can be represented faithfully by the normal ChatGPT Web transport. A JSON field being part of the official OpenAI API does not mean it is silently ignored here: meaningful unsupported values return HTTP 400 with type: invalid_request_error, param, and code.

Every response includes X-Request-ID for correlation. Browser-backed requests carry that same validated ID through the local relay; it is never sent as user content. Upstream ChatGPT conversation identifiers remain available through the documented X-GPTWeb-* headers.

Endpoints

Endpoint Status Notes
GET /healthz Supported Local process health only.
GET /v1/models Supported Returns the authenticated account's real ChatGPT Web model catalog. No synthetic fallback list.
GET /v1/models/{model} Supported Returns a catalog model or model_not_found.
POST /v1/completions Supported single-prompt text subset Legacy streaming/non-streaming text for one string prompt and one choice.
POST /v1/chat/completions Supported text + image-input subset Streaming/non-streaming text output, image input, and local store:true.
Stored Chat Completion list/retrieve/update/delete/messages Supported local subset Independent local records; metadata values are strings.
POST /v1/vision/analyze Supported convenience extension Multipart one-to-ten images plus prompt to Chat Completion text/SSE.
POST /v1/responses Supported text, image, and local-file subset Text output, image/file input, streaming, previous_response_id, and local Conversations.
POST /v1/images/generations Supported single-image subset prompt, optional model, n=1, and base64 output.
POST /v1/images/edits Supported image-edit subset One to ten PNG/JPEG/GIF/WebP inputs, prompt, optional model, and one base64 output.
POST /v1/images/variations Supported compatibility subset One image to one base64 output using a fixed Web variation instruction.
GET /v1/responses/{response_id} Supported for stored text Responses Returns the completed local Response object.
DELETE /v1/responses/{response_id} Supported for stored text Responses Deletes the Response and its continuation mapping.
GET /v1/responses/{response_id}/input_items Supported text subset Stored message inputs with after, limit, and order pagination.
GET /v1/responses/{response_id}/input_items/{input_item_id} Supported text subset Retrieves one stored message input.
POST/GET /v1/files Supported local storage subset Bounded 25 MiB-per-file local store, listing and cursor pagination.
GET/DELETE /v1/files/{file_id} Supported local storage subset Metadata retrieval and deletion.
GET /v1/files/{file_id}/content Supported local storage subset Returns original bytes.
Conversations and item CRUD Supported local subset Durable local message context for Responses; not Platform cloud storage.
Upload create/part/complete/cancel Supported local subset One-hour uploads, 25 MiB final limit, disk-streamed parts.
Batch create/list/retrieve/cancel Supported local subset JSONL local Files, supported text endpoints, one sequential worker.
Responses cancel Not applicable to synchronous responses Background Responses are not implemented.
Remote Platform storage, embeddings, audio, fine-tuning, Realtime, Assistants/vector stores Not implemented Local storage is explicitly local and does not impersonate these backends.

Unknown URLs and incorrect methods return JSON invalid_request_error responses with not_found or method_not_allowed, plus X-Request-ID.

On Windows, the optional persistent embedded frontend can carry text, image-input, and local-file conversation requests through ChatGPT's official frontend lifecycle using a local OAuth credential file. It supports auto and explicit catalog model IDs; see BROWSER_RELAY.md. This changes the transport, not the API feature matrix below.

Legacy Completions request contract

Implemented

  • model
  • one non-empty string prompt
  • stream
  • store:true|false; true stores the request/final response locally
  • string-to-string metadata with up to 16 entries
  • explicit Web continuation through the documented X-GPTWeb-* headers

Accepted standard no-op defaults

  • n: 1
  • best_of: 1
  • echo: false
  • logprobs: 0
  • zero frequency/presence penalties
  • empty logit_bias and stop values
  • stream_options: {"include_usage":false}

Token-array or multi-prompt input, exact token limits, suffixes, sampling controls, non-empty stop sequences, and usage/logprob claims are rejected.

Chat Completions request contract

Implemented

  • model
  • messages with developer, system, user, and assistant text
  • user-message image_url parts using base64 PNG/JPEG/GIF/WebP data URLs
  • stream
  • basic text content arrays
  • explicit Web continuation through X-GPTWeb-Conversation-ID plus the previous assistant/parent message ID

Without both continuation headers, every Chat Completions request starts a new Web conversation with history/training disabled. Replayed API messages are carried as isolated system context plus the current user turn; they do not silently attach to the previously used browser conversation.

Fresh requests also receive a compact, lowest-priority stateless/output-only baseline so Web answers behave more like direct API payloads. Caller system/developer instructions follow that baseline and remain authoritative. Set GPTWEB_FIDELITY_MODE=chatgpt to retain the native conversational style.

Accepted standard no-op defaults

These values are accepted because they do not claim unsupported behavior:

  • n: 1
  • tools: []
  • functions: []
  • tool_choice: "none" or "auto" when no tools are supplied
  • function_call: "none" or "auto" when no functions are supplied
  • parallel_tool_calls: true|false when no tools are supplied
  • modalities: [] or ["text"]
  • response_format: {"type":"text"}
  • stream_options: {"include_usage":false}
  • logprobs: false
  • top_logprobs: 0
  • zero frequency/presence penalties
  • empty logit-bias and stop values

Explicitly rejected until faithfully mapped

  • custom tools/functions and tool calls
  • JSON mode / JSON Schema structured outputs
  • audio output and non-text modalities
  • plain HTTP image URLs; remote image fetching requires HTTPS, public DNS/IPs, port 443, at most three redirects, supported image bytes, and the normal attachment size limits
  • exact token limits or token usage
  • sampling controls such as temperature, top_p, and seed
  • non-empty stop sequences
  • log probabilities
  • service tiers, prompt caching, prediction, verbosity, and web-search options
  • API reasoning-effort values; ChatGPT Web exposes a different model-specific effort catalog and the gateway does not guess a cross-product mapping

Responses request contract

Implemented

  • model
  • string input
  • message-array input with text content
  • input_image in user messages or top-level input items using base64 data URLs
  • input_file.file_id backed by the local Files store and uploaded to the official frontend for that request
  • string instructions on fresh responses
  • stream
  • previous_response_id backed by a bounded local continuation map persisted atomically across gateway restarts
  • store: true|false (default true), with retrieve/delete for stored responses
  • stored input-item listing with cursor pagination and single-item retrieval
  • local conversation ID/reference, mutually exclusive with previous_response_id
  • plain text configuration through text.format.type = "text"
  • text-only modalities

Accepted standard no-op defaults

  • background: false
  • tools: []
  • include: []
  • empty metadata
  • tool_choice: "none" or "auto" when no tools are supplied
  • parallel_tool_calls: true|false when no tools are supplied
  • modalities / output_modalities set to text only
  • empty stream_options
  • top_logprobs: 0

Explicitly rejected until faithfully mapped

  • tools, function calls, computer use, file search, and hosted tools
  • JSON Schema / structured text output
  • background Responses
  • maximum output/tool-call limits
  • sampling and logprob controls
  • prompt templates and prompt caching
  • reasoning configuration and encrypted reasoning items
  • service tiers, truncation, safety identifiers, and remote OpenAI Platform storage semantics

Response differences

  • Browser-backed stream:true is genuinely incremental: the relay forwards append-only rendered/network deltas and the gateway flushes each standard SSE event before browser completion. A small uncommitted tail permits frontend rewrites, and the terminal response is prefix-reconciled to prevent duplicate text.
  • Exact OpenAI Platform token accounting is not available from ChatGPT Web, so the gateway does not fabricate usage.
  • Account-specific specialty quotas and safe reasoning progress may appear in the optional top-level gptweb extension.
  • The resolved ChatGPT Web model is returned in model when the stream identifies it.
  • Response IDs map to upstream conversation/message IDs in the local GPTWEB_RESPONSE_STATE_FILE for 30 days (up to 4096 entries). With the default store: true, completed response objects, generated output text, and normalized text input items are also retained for retrieve/delete/input-items. The file never contains OAuth tokens or cookies. Use store:false when local prompt/output persistence is not desired.
  • Internal/hidden assistant-to-tool traffic is filtered and is not exposed as user-visible output.

Error behavior

  • Unknown fields and meaningful unsupported values on both Chat Completions and Responses return HTTP 400 with type, param, and code.
  • Empty bodies, malformed JSON, multiple top-level JSON values, and oversized bodies return explicit JSON request errors.
  • Missing account models return HTTP 404 with code: model_not_found and param: model.
  • Upstream authentication, Cloudflare Challenge Page, generic forbidden, rate-limit, request-size, and not-found failures are classified separately without returning arbitrary upstream HTML.
  • Direct conversations retry transient 502/503/504 failures up to GPTWEB_UPSTREAM_MAX_RETRIES times; 429 responses are never auto-retried.

Operational contract

  • Startup rejects malformed numeric configuration values.
  • Non-loopback listening requires GPTWEB_API_KEY.
  • The server handles SIGINT/SIGTERM with a bounded graceful shutdown.
  • Request logs are structured JSON and include only request ID, method, path, status, duration, and response size.
  • Logs do not include query strings, request bodies, Authorization headers, cookies, OAuth tokens, or ChatGPT account IDs.