diff --git a/devlog/_plan/260912_stream_recovery/092_resume.md b/devlog/_plan/260912_stream_recovery/092_resume.md new file mode 100644 index 0000000000..a6463ddfe4 --- /dev/null +++ b/devlog/_plan/260912_stream_recovery/092_resume.md @@ -0,0 +1,20 @@ +# Stream delivery verification update + +Five source fixes were delivered as independent dev-based PRs. PR4341 (terminal integrity) is merged; PR4354 (Console), PR4356 (search), PR4363 (Cursor) and PR4367 (live sideband) remain open at the 2026-09-12 verification update. This task does not merge integration branches. Runtime verification remains incomplete; shared Cline, history and Windows failures are not treated as passing baseline evidence. + +## Console label capture + +![Console recovery label with synthetic log data](093_console_recovery.png) + +The capture uses the dashboard-preview artifact from GitHub Actions run34675214816, artifact10294058486, recorded build commit f9dcd6449298821e6ba02d026037a677effe6aaa and GUI tree e0d22385336080717ad29a14d65a270a207a41e7. The artifact was built in hosted CI. No local product build, install, typecheck or test suite ran. + +The artifact GUI differs from Console source head2c63a5d4283936aa9d0525e39090afb5e492c0ec only in Combo workspace files and associated tests; Logs.tsx, its locale strings and styles are unchanged. The Logs page was served locally from the existing bundle with a synthetic API fixture. The screenshot shows the actual Console upload retry label in the request detail dialog. Values, request ID and model are synthetic, not live request or billing evidence. This is rendered label evidence, not a backend retry test. + +## Remaining acceptance + +- #3389 remains HOLD: zero observed bytes cannot establish upstream nonexecution. +- #4191 retains its native failing-stage evidence requirement; existing diagnostics are not a reproduced fix. +- #4312 has its open-tool status sub-defect addressed; the actual client nonretryable refusal contract remains unresolved. +- #3506 requires a redacted translation-fidelity exchange. No semantic-progress cutoff is added. + +Fresh source/security reviews and exact CI heads/runs are retained in the task-local handoff. Skipped, cancelled, failed and superseded runs do not certify completion. Local suites/build/typecheck/install: NOT RUN. diff --git a/devlog/_plan/260912_stream_recovery/093_console_recovery.png b/devlog/_plan/260912_stream_recovery/093_console_recovery.png new file mode 100644 index 0000000000..43e6fdf359 Binary files /dev/null and b/devlog/_plan/260912_stream_recovery/093_console_recovery.png differ diff --git a/devlog/_plan/260912_stream_recovery/094_console_fixture.md b/devlog/_plan/260912_stream_recovery/094_console_fixture.md new file mode 100644 index 0000000000..31e4e77c5b --- /dev/null +++ b/devlog/_plan/260912_stream_recovery/094_console_fixture.md @@ -0,0 +1,7 @@ +# Effective Console destination fixture correction + +Hosted macOS control run34693025423/job103551974484 failed the canonical-row other-host fixture. The log records that routing discarded the configured other-host URL and selected the canonical Go endpoint. The test therefore never reached its intended negative condition. This is not evidence of a noncanonical replay. + +The negative case now configures an unsupported generation path on the same canonical row. Both supported adapter path fields name `/unrelated`; the assertion requires exactly one outgoing URL at that path and preserves the original refusal. Existing custom-row other-host coverage remains. A separate control records the two exact `/responses` sends produced after canonical base normalization; the exact Muse model wire default selects Responses. No production code, policy, retry count or assertion skip changes. + +Local tests/build/typecheck/install are NOT RUN. Independent source review and updated exact-head hosted CI are required. Windows pnpm/Devin failures in the same run remain separate shared repair items; no baseline runtime pass is claimed. diff --git a/devlog/_plan/260912_thinking_contract/000_plan.md b/devlog/_plan/260912_thinking_contract/000_plan.md new file mode 100644 index 0000000000..86a87f9d36 --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/000_plan.md @@ -0,0 +1,31 @@ +# Preserve reasoning provenance and transport intent + +Readers: maintainers choosing whether to integrate the thinking lane. Raw reasoning must remain content, while provider-authored summaries can be displayed under an explicit provider default. The plan reconciles #4301 and #4287, separately reviews #3652 hint suppression, and carries #4130 Spark compatibility without retirement. + +Loop: satisfy-spec HOTL, triggered by authorized thinking-lane delivery. Goal: reviewable carry PRs and final-head hosted CI. Non-goals: merges, closure of source PRs, retirement #4334, releases, user service/config changes, other worktrees. All local product suites/build/typecheck/install are NOT RUN by instruction. Only available existing credentials/tools are used; no user token/time/agent ceiling was set. Stop: every source PR has a justified disposition and every delivered branch has exact-head hosted CI evidence. Outcomes: DONE on evidence, HOLD/NEEDS_HUMAN on explicit unresolved acceptance, never fake green. Escalation: real tool denial or requirement beyond scope; main reclaims after two distinct reviewer failures. Native architect selector is unavailable; inherited independent design review and reflection follow the user instruction, with a separate A audit. + +## Dependency map + +| Cycle | Artifact | Result | +| --- | --- | --- | +| roadmap | this file and all decade docs | docs-only plan lock | +| presentation | 010_presentation.md | raw/summary contract and provider opt-in | +| hint | 020_transport_hint.md | independent transport-hint disposition/carry | +| spark | 030_spark.md | independent Spark Lite carry | +| delivery | 040_delivery.md | final heads, review closure and hosted CI | + +Presentation combines two conflicting source proposals into one contract. Hint and Spark are independent and receive ordinary dev-based PRs, not artificial stack dependencies. Final review consumes all branches. No GitHub native stacks are requested. + +## Evidence and owner map + +Baseline origin/dev: 69e3dcda755a52feb1327edad6c8ea6cefd6e871. Source PR heads: #4301 5d6d1862a11da6e4d0c04eb7f35f9f48ae1285fd; #4287 fe13bdb7bf8403a2a2cdb10f258a68b649177953; #3652 13fb263778e9036e66ae86d41e29f9f47bbbed92; #4130 5d56f5461ea3d18668b85f6bb0d8a523920f2536. All open when inspected. Original authors: Robin Bially, yxr1995-maker, itismyfield, luvs01; exact Git trailers will be read from original commits before carrying. + +Current owners: src/bridge.ts:663 raw-reasoning finalization; src/adapters/google.ts:571 shared part classifier; src/server/responses/core.ts:2490 final-route normalization; src/responses/parser.ts:543 summary omission policy; src/types/request.ts:310 AdapterEvent. Reuse these boundaries; no new event enum or generic service layer. Structure INDEX maps shared areas to topical documents; main contracts are providers/chat-compat.md, providers/google.md, transports/responses.md and config.md with references from affected area owners. + +Verification: git diff --check was run at baseline and exited 0, checking diff whitespace only. GitHub ci.yml workflow_dispatch lane=all reads checkout source, typechecks, runs product suites and cross-platform jobs; NOT RUN locally. Every conditional scenario is named in decade docs and must be asserted in committed regression tests. Source inspection is not runtime proof. + +## Cycle records + +Roadmap P: requirements/source inspection and independent design review in progress. No product patch applied. + +Roadmap B: locked amended contract after independent A PASS and both design reflections ALIGNED. Product implementation starts in the next cycle. diff --git a/devlog/_plan/260912_thinking_contract/001_sources.md b/devlog/_plan/260912_thinking_contract/001_sources.md new file mode 100644 index 0000000000..f05813eae8 --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/001_sources.md @@ -0,0 +1,7 @@ +# Source decisions + +Public PR diffs and latest comments are the source proposal evidence. #4130's September 11 corrections pin Lite on for nonempty additional_tools bodies and off otherwise; adopting the earlier unconditional false version loses tools. #4334 is an explicit retirement HOLD and is not carried. + +#4301 removes automatic content-to-summary conversion. #4287 tests raw DeepSeek content as a visible summary; that expectation conflicts with provenance and will be replaced, not adopted. Google thought-summary API documentation distinguishes summaries from opaque thought signatures: https://ai.google.dev/gemini-api/docs/generate-content/thinking (opened 2026-09-12). CCA generationConfig/includeThoughts behavior is contributor probe evidence, not a newly performed live-service probe. + +Searches used: reasoning_raw_delta, thinking_delta, hideThinkingSummary, googlePartTextEvent, preserveReasoningContent, and the four PR numbers. Existing bridge event types can distinguish raw content and summary without a new enum. No-code/config-only options cannot repair the existing mislabeled content; raw rewrite deletion plus existing boundaries is the smallest change. diff --git a/devlog/_plan/260912_thinking_contract/010_presentation.md b/devlog/_plan/260912_thinking_contract/010_presentation.md new file mode 100644 index 0000000000..b1503a5f34 --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/010_presentation.md @@ -0,0 +1,322 @@ +# Presentation contract + +Class C4 public protocol contract. Depends on roadmap lock. MODIFY src/bridge.ts: closeCurrentRawReasoning and reasoning_raw_delta emit response.reasoning_text.delta/done with content_index:0; final items use summary:[] and content:[{type:reasoning_text,text}]. buildResponseJSONWithBudget mirrors this. Keep hidden txt-only replay envelopes intact. DELETE src/server/responses-reasoning-summary-rewrite.ts and its obsolete unit test; MODIFY core.ts to remove imports and SSE/JSON content-to-summary rewrites. MODIFY both layout manifests to remove that test. Adopt the exact #4301 hunks below except reporter video and historical verification record. + +MODIFY provider.ts, registry.ts, derive.ts, router.ts and auth-cors.ts to carry showThinkingSummary boolean (preserve explicit false). Seed only google-antigravity true. Creation: provider config/registry; serialization: providerConfigSeed and deriveKeyLoginMap; deserialization: config provider passthrough and management field policy; consumers: routedProviderConfig, final-route normalization, Google request builder. No new enum. + +MODIFY core.ts final-route normalization: apply provider default only when original reasoning.summary is omitted, never explicit none; recompute on each final route so fallback cannot inherit another provider default. Provider opt-in authorizes summary display, not raw-to-summary conversion. + +MODIFY google.ts shared part classifier to use existing thinking_delta only for Gemini thought summaries under verified Gemini model provenance; CCA Claude/gpt-oss thought text remains reasoning_raw_delta. Persist request-local Gemini identity using existing adapter state, used by both stream and buffered classifier calls. includeThoughts stays provider-opted, Gemini-only, non-image and explicit-hide aware. MODIFY google-wire-compiler.ts to retain only boolean true includeThoughts, independently of thinkingLevel. Do not claim raw text is an actual summary. + +MODIFY the #4287 end-to-end fixture: raw DeepSeek content remains content with empty summary even under provider opt-in; actual CCA Gemini thought parts use summary; omitted vs none vs auto, explicit provider false, saved-row enrichment, fallback reset, streaming/buffered paths. Extend existing Google tests and bridge raw tests; both layout manifests register responses-show-thinking-summary.test.ts. Update English providers docs and structure owners, keeping locale statements consistent. Source tests are authored but run only by hosted CI. + +Acceptance: raw event fixture => content delta and no summary delta; actual Gemini summary fixture => summary only when requested/provider-opted; explicit none => no synthesized summary and no includeThoughts request; false/unknown provider => no opt-in; fallback to unopted route => hidden behavior reset; replay envelope decodes same raw text and tool continuation remains valid; native Responses mixed content/summary remains byte-semantically native. No model prose synthesizer is introduced. + +## Source patch blueprint + +```diff +diff --git a/src/bridge.ts b/src/bridge.ts +index 20e7c3fe09..bc90f35b94 100644 +--- a/src/bridge.ts ++++ b/src/bridge.ts +@@ -663,16 +663,13 @@ export function bridgeToResponsesSSE( + const closeCurrentRawReasoning = () => { + if (!currentRawReasoning) return; + rawReasoningForNextToolCall = currentRawReasoning.text; +- emit("response.reasoning_summary_text.done", { +- item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, summary_index: 0, text: currentRawReasoning.text, +- }); +- emit("response.reasoning_summary_part.done", { +- item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, summary_index: 0, +- part: { type: "summary_text", text: currentRawReasoning.text }, ++ emit("response.reasoning_text.done", { ++ item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, content_index: 0, text: currentRawReasoning.text, + }); + const item = { + type: "reasoning", id: currentRawReasoning.itemId, +- summary: [{ type: "summary_text", text: currentRawReasoning.text }], ++ summary: [] as never[], ++ content: [{ type: "reasoning_text", text: currentRawReasoning.text }], + }; + emit("response.output_item.done", { output_index: currentRawReasoning.outputIndex, item }); + retainFinishedItem(item as OutputItem, currentRawReasoning.textBytes, "reasoning"); +@@ -1111,10 +1108,6 @@ export function bridgeToResponsesSSE( + const itemId = `rs_${uuid()}`; + const item = { type: "reasoning", id: itemId, summary: [] as { type: string; text: string }[] }; + emit("response.output_item.added", { output_index: outputIndex, item }); +- emit("response.reasoning_summary_part.added", { +- item_id: itemId, output_index: outputIndex, summary_index: 0, +- part: { type: "summary_text", text: "" }, +- }); + currentRawReasoning = { itemId, outputIndex, text: "", textBytes: 0 }; + } + ({ value: currentRawReasoning.text, bytes: currentRawReasoning.textBytes } = appendString( +@@ -1123,9 +1116,13 @@ export function bridgeToResponsesSSE( + event.text, + "reasoning", + )); +- emit("response.reasoning_summary_text.delta", { ++ // Raw reasoning (openai-chat reasoning_content, kiro tags) rides the CONTENT ++ // channel, matching native gpt-oss passthrough: Codex applies its own display ++ // policy, so the desktop band shows the "Thinking…" placeholder instead of the ++ // raw CoT (the #45 summary-channel display intent is intentionally reverted). ++ emit("response.reasoning_text.delta", { + item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, +- summary_index: 0, delta: event.text, ++ content_index: 0, delta: event.text, + }); + break; + } +@@ -1780,7 +1777,8 @@ function buildResponseJSONWithBudget( + } + pushOutput({ + type: "reasoning", id: `rs_${uuid()}`, +- summary: [{ type: "summary_text", text: currentRawReasoning }], ++ summary: [], ++ content: [{ type: "reasoning_text", text: currentRawReasoning }], + }, currentRawReasoningBytes, "reasoning"); + currentRawReasoning = ""; + currentRawReasoningBytes = 0; + +``` + +## Pinned source hunks (apply with corrections above) + +```diff +diff --git a/src/adapters/google-wire-compiler.ts b/src/adapters/google-wire-compiler.ts +index 88c482ba7d..aa835e50b4 100644 +--- a/src/adapters/google-wire-compiler.ts ++++ b/src/adapters/google-wire-compiler.ts +@@ -130,12 +130,20 @@ function compileGenerationConfig(value: unknown): JsonObject | undefined { + ))].slice(0, 5); + if (stopSequences.length > 0) out.stopSequences = stopSequences; + } +- if (isObject(value.thinkingConfig) && typeof value.thinkingConfig.thinkingLevel === "string") { +- const raw = value.thinkingConfig.thinkingLevel.toLowerCase(); +- const thinkingLevel = GOOGLE_THINKING_LEVELS.has(raw) +- ? raw +- : (["xhigh", "max", "ultra"].includes(raw) ? "high" : undefined); +- if (thinkingLevel) out.thinkingConfig = { thinkingLevel }; ++ if (isObject(value.thinkingConfig)) { ++ const thinking: JsonObject = {}; ++ if (typeof value.thinkingConfig.thinkingLevel === "string") { ++ const raw = value.thinkingConfig.thinkingLevel.toLowerCase(); ++ const thinkingLevel = GOOGLE_THINKING_LEVELS.has(raw) ++ ? raw ++ : (["xhigh", "max", "ultra"].includes(raw) ? "high" : undefined); ++ if (thinkingLevel) thinking.thinkingLevel = thinkingLevel; ++ } ++ // The one key that makes Google return `thought: true` text. Cloud Code Assist serves ++ // thinking either way (thoughtsTokenCount stays non-zero) but withholds the text unless the ++ // request opts in, so dropping it here silently reinstates the missing-thinking behavior. ++ if (value.thinkingConfig.includeThoughts === true) thinking.includeThoughts = true; ++ if (Object.keys(thinking).length > 0) out.thinkingConfig = thinking; + } + if (Array.isArray(value.responseModalities)) { + const valid = value.responseModalities.filter((m): m is string => typeof m === "string" && ["TEXT", "IMAGE", "AUDIO"].includes(m)); +diff --git a/src/adapters/google.ts b/src/adapters/google.ts +index 7fcc88ba59..5a6675f6e4 100644 +--- a/src/adapters/google.ts ++++ b/src/adapters/google.ts +@@ -866,11 +866,27 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte + ); + antigravityModel = wireModelId; + antigravitySession = sessionId; ++ // Gemini returns no chain-of-thought TEXT unless the request opts in. Probed against CCA ++ // 2026-09-12: `gemini-3.8-flash-high` answered with thoughtsTokenCount=321 and zero ++ // `thought` parts, then 358-652 chars of genuine reasoning once includeThoughts was set. ++ // Scoped to Gemini wire ids — Claude-on-CCA accepts the flag but never returns thought ++ // parts, and gpt-oss rejects it outright (400 INVALID_ARGUMENT, which would break every ++ // gpt-oss turn). Gated on the provider's visible-thinking opt-in so a user who wants ++ // thinking hidden does not pay conversation-history tokens for text nobody renders; ++ // `hideThinkingSummary !== true` is the same per-request gate the response path uses, so ++ // a client that explicitly asked for hidden thinking is not billed for the text either. ++ const includeThoughts = provider.showThinkingSummary === true ++ && parsed.options.hideThinkingSummary !== true ++ && /^gemini-/.test(wireModelId) ++ && !isImageCapableModel(parsed.modelId); + // Effort → thinkingConfig for CCA (CLIProxyAPI proven: request.generationConfig.thinkingConfig). + // Suffix/compat IDs return thinkingLevel=undefined — the suffix IS the effort, no contradiction. +- if (thinkingLevel) { ++ if (thinkingLevel || includeThoughts) { + const gc = (body.generationConfig ?? {}) as Record; +- gc.thinkingConfig = { thinkingLevel }; ++ gc.thinkingConfig = { ++ ...(thinkingLevel ? { thinkingLevel } : {}), ++ ...(includeThoughts ? { includeThoughts: true } : {}), ++ }; + body.generationConfig = gc; + } + // Reasoning continuity: Gemini models re-inject cached thoughtSignatures; Claude-on-Antigravity +diff --git a/src/providers/derive.ts b/src/providers/derive.ts +index 67e6c0522e..7edf28787b 100644 +--- a/src/providers/derive.ts ++++ b/src/providers/derive.ts +@@ -43,6 +43,7 @@ export interface DerivedKeyLoginProvider { + autoToolChoiceOnlyModels?: string[]; + preserveReasoningContentModels?: string[]; + requiresReasoningPlaceholderModels?: string[]; ++ showThinkingSummary?: boolean; + reasoningSplitModels?: string[]; + reasoningDetailsModels?: string[]; + thinkingToggleModels?: string[]; +@@ -267,6 +268,7 @@ export function providerConfigSeed(entry: ProviderRegistryEntry): OcxProviderCon + ...(entry.autoToolChoiceOnlyModels ? { autoToolChoiceOnlyModels: [...entry.autoToolChoiceOnlyModels] } : {}), + ...(entry.preserveReasoningContentModels ? { preserveReasoningContentModels: [...entry.preserveReasoningContentModels] } : {}), + ...(entry.requiresReasoningPlaceholderModels ? { requiresReasoningPlaceholderModels: [...entry.requiresReasoningPlaceholderModels] } : {}), ++ ...(entry.showThinkingSummary !== undefined ? { showThinkingSummary: entry.showThinkingSummary } : {}), + ...(entry.reasoningSplitModels ? { reasoningSplitModels: [...entry.reasoningSplitModels] } : {}), + ...(entry.reasoningDetailsModels ? { reasoningDetailsModels: [...entry.reasoningDetailsModels] } : {}), + ...(entry.thinkingToggleModels ? { thinkingToggleModels: [...entry.thinkingToggleModels] } : {}), +@@ -315,6 +317,7 @@ export function deriveKeyLoginMap(): Record { + ...(entry.autoToolChoiceOnlyModels ? { autoToolChoiceOnlyModels: [...entry.autoToolChoiceOnlyModels] } : {}), + ...(entry.preserveReasoningContentModels ? { preserveReasoningContentModels: [...entry.preserveReasoningContentModels] } : {}), + ...(entry.requiresReasoningPlaceholderModels ? { requiresReasoningPlaceholderModels: [...entry.requiresReasoningPlaceholderModels] } : {}), ++ ...(entry.showThinkingSummary !== undefined ? { showThinkingSummary: entry.showThinkingSummary } : {}), + ...(entry.reasoningSplitModels ? { reasoningSplitModels: [...entry.reasoningSplitModels] } : {}), + ...(entry.reasoningDetailsModels ? { reasoningDetailsModels: [...entry.reasoningDetailsModels] } : {}), + ...(entry.thinkingToggleModels ? { thinkingToggleModels: [...entry.thinkingToggleModels] } : {}), +@@ -567,6 +570,7 @@ export function enrichProviderFromRegistry(name: string, prov: OcxProviderConfig + if (!prov.thinkingToggleModels && seed.thinkingToggleModels) prov.thinkingToggleModels = [...seed.thinkingToggleModels]; + if (!prov.thinkingBudgetModels && seed.thinkingBudgetModels) prov.thinkingBudgetModels = [...seed.thinkingBudgetModels]; + if (prov.escapeBuiltinToolNames === undefined && seed.escapeBuiltinToolNames !== undefined) prov.escapeBuiltinToolNames = seed.escapeBuiltinToolNames; ++ if (prov.showThinkingSummary === undefined && seed.showThinkingSummary !== undefined) prov.showThinkingSummary = seed.showThinkingSummary; + if (prov.keyOptional === undefined && seed.keyOptional !== undefined) prov.keyOptional = seed.keyOptional; + if (prov.freeTier === undefined && seed.freeTier !== undefined) prov.freeTier = seed.freeTier; + if (prov.modelSuffixBracketStrip === undefined && seed.modelSuffixBracketStrip !== undefined) prov.modelSuffixBracketStrip = seed.modelSuffixBracketStrip; +diff --git a/src/providers/registry.ts b/src/providers/registry.ts +index f72bb7650b..e483e9db23 100644 +--- a/src/providers/registry.ts ++++ b/src/providers/registry.ts +@@ -343,6 +343,10 @@ export interface ProviderRegistryEntry { + autoToolChoiceOnlyModels?: string[]; + preserveReasoningContentModels?: string[]; + requiresReasoningPlaceholderModels?: string[]; ++ /** ++ * Opt this provider into visible thinking summaries (see OcxProviderConfig.showThinkingSummary). ++ */ ++ showThinkingSummary?: boolean; + reasoningSplitModels?: string[]; + reasoningDetailsModels?: string[]; + thinkingToggleModels?: string[]; +@@ -367,7 +371,7 @@ export type ProviderConfigSeed = Pick< + | "modelMaxInputTokens" | "defaultMaxOutputTokens" | "modelMaxOutputTokens" + | "reasoningEfforts" | "modelReasoningEfforts" | "modelDefaultReasoningEfforts" | "reasoningEffortMap" | "modelReasoningEffortMap" | "reasoningWireFormat" + | "noVisionModels" | "noReasoningModels" | "noTemperatureModels" | "noTopPModels" | "noPenaltyModels" +- | "autoToolChoiceOnlyModels" | "preserveReasoningContentModels" | "requiresReasoningPlaceholderModels" | "reasoningSplitModels" | "reasoningDetailsModels" | "thinkingToggleModels" | "thinkingBudgetModels" | "escapeBuiltinToolNames" | "openaiChatEofTolerance" ++ | "autoToolChoiceOnlyModels" | "preserveReasoningContentModels" | "requiresReasoningPlaceholderModels" | "reasoningSplitModels" | "reasoningDetailsModels" | "thinkingToggleModels" | "thinkingBudgetModels" | "escapeBuiltinToolNames" | "openaiChatEofTolerance" | "showThinkingSummary" + | "googleMode" | "project" | "location" | "headers" + >; + +@@ -2045,7 +2049,7 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [ + // path must stay RELATIVE: this row sets `allowBaseUrlOverride`, and an absolute `url` would + // retarget a user's custom base back to Google. A leading `./` is required because a bare + // `v1internal:` reads as a URL scheme and `providerModelDiscoverySpecError` rejects it. +- { id: "google-antigravity", alias: "agy", label: "Google Antigravity", adapter: "google", baseUrl: "https://daily-cloudcode-pa.googleapis.com", authKind: "oauth", allowBaseUrlOverride: true, dashboardUrl: "https://antigravity.google", models: ANTIGRAVITY_MODELS, liveModels: true, defaultModel: "gemini-3.8-flash", modelContextWindows: ANTIGRAVITY_MODEL_CONTEXT_WINDOWS, modelInputModalities: ANTIGRAVITY_MODEL_INPUT_MODALITIES, modelReasoningEfforts: ANTIGRAVITY_MODEL_EFFORTS, googleMode: "cloud-code-assist", jawcodeBundle: "google", extraMetadataAliases: ["antigravity", "gemini-antigravity"], modelDiscovery: { path: "./v1internal:fetchAvailableModels" } }, ++ { id: "google-antigravity", alias: "agy", label: "Google Antigravity", adapter: "google", baseUrl: "https://daily-cloudcode-pa.googleapis.com", authKind: "oauth", allowBaseUrlOverride: true, dashboardUrl: "https://antigravity.google", models: ANTIGRAVITY_MODELS, liveModels: true, defaultModel: "gemini-3.8-flash", modelContextWindows: ANTIGRAVITY_MODEL_CONTEXT_WINDOWS, modelInputModalities: ANTIGRAVITY_MODEL_INPUT_MODALITIES, modelReasoningEfforts: ANTIGRAVITY_MODEL_EFFORTS, googleMode: "cloud-code-assist", showThinkingSummary: true, jawcodeBundle: "google", extraMetadataAliases: ["antigravity", "gemini-antigravity"], modelDiscovery: { path: "./v1internal:fetchAvailableModels" } }, + { id: "azure-openai", label: "Azure OpenAI", adapter: "azure-openai", baseUrl: "https://{resource}.openai.azure.com/openai", authKind: "key", featured: true, dashboardUrl: "https://portal.azure.com" }, + { id: "ollama", label: "Ollama (local)", adapter: "openai-chat", baseUrl: "http://localhost:11434/v1", authKind: "local", allowPrivateNetworkByDefault: true, allowBaseUrlOverride: true, featured: true, note: "Local — key usually blank" }, + { id: "vllm", label: "vLLM (local)", adapter: "openai-chat", baseUrl: "http://localhost:8000/v1", authKind: "local", allowPrivateNetworkByDefault: true, allowBaseUrlOverride: true, featured: true, note: "Local — key usually blank" }, +diff --git a/src/router.ts b/src/router.ts +index 55a0326fce..bf2b9b4b98 100644 +--- a/src/router.ts ++++ b/src/router.ts +@@ -410,6 +410,13 @@ export function routedProviderConfig(providerName: string, provider: OcxProvider + ...(provider.preserveResponsesReasoningContent === undefined && registryEntry.preserveResponsesReasoningContent !== undefined + ? { preserveResponsesReasoningContent: registryEntry.preserveResponsesReasoningContent } + : {}), ++ // The request path resolves through routedProviderConfig() and never calls ++ // enrichProviderFromRegistry(), so a saved provider row written before the ++ // registry learned this flag must be backfilled here or route.provider never ++ // carries it and the showThinkingSummary opt-in stays dead. ++ ...(provider.showThinkingSummary === undefined && registryEntry.showThinkingSummary !== undefined ++ ? { showThinkingSummary: registryEntry.showThinkingSummary } ++ : {}), + // Registry-only client-facing repair policy (#938): fill only when the + // saved provider has no explicit policy; clone so runtime never aliases + // the registry constant. +diff --git a/src/server/auth-cors.ts b/src/server/auth-cors.ts +index 3a93246cd0..93377b5573 100644 +--- a/src/server/auth-cors.ts ++++ b/src/server/auth-cors.ts +@@ -885,6 +885,7 @@ const PROVIDER_CONFIG_FIELD_POLICY = { + autoToolChoiceOnlyModels: "editor", + preserveReasoningContentModels: "editor", + requiresReasoningPlaceholderModels: "editor", ++ showThinkingSummary: "editor", + retryOn429: "editor", + transientRetryOn5xx: "editor", + reasoningSplitModels: "editor", +diff --git a/src/server/responses/core.ts b/src/server/responses/core.ts +index cccd942026..852cd9f8b0 100644 +--- a/src/server/responses/core.ts ++++ b/src/server/responses/core.ts +@@ -2467,6 +2467,20 @@ async function resolveSubagentFallbackModelEligibility(args: { + }; + } + ++/** ++ * Whether the client explicitly asked for hidden thinking (`reasoning.summary: "none"`). ++ * ++ * Pinned: parseRequest collapses "omitted" and "none" into one hideThinkingSummary ++ * flag, so the raw request body is the ONLY place that still distinguishes them. ++ * Provider opt-ins like showThinkingSummary must consult this — never the flag ++ * alone — or a future caller that copies only the flag would silently unlock an ++ * explicit opt-out. ++ */ ++function clientExplicitlyHidThinking(parsed: OcxParsedRequest): boolean { ++ const rawReasoning = (parsed._rawBody as { reasoning?: { summary?: unknown } } | undefined)?.reasoning; ++ return typeof rawReasoning === "object" && rawReasoning !== null ++ && (rawReasoning as { summary?: unknown }).summary === "none"; ++} + /** + * Apply every route-dependent request mutation against the final selected route. + * Must run only after subagent fallback has settled the model/provider. +@@ -2508,6 +2522,15 @@ async function applyFinalRouteRequestNormalization(args: { + // this request will actually use (#404). + route.provider = resolveOpenCodeGoTransport(route.provider, getOrAllocateRequestSessionLane(req)); + route.provider = resolveWireProtocolOverride(route.providerName, route.modelId, route.provider, inboundWire); ++ // Provider-opted visible thinking (e.g. google-antigravity): parseRequest hides thinking ++ // whenever the client omits reasoning.summary, which is the Codex default. A provider that ++ // serves genuine user-facing reasoning opts back into the summary channel here, so thought ++ // parts (Gemini thought, content-channel reasoning_text) reach the client instead of only ++ // the hidden replay envelopes. An explicit client reasoning.summary "none" still wins. ++ if (route.provider.showThinkingSummary === true && parsed.options.hideThinkingSummary === true ++ && !clientExplicitlyHidThinking(parsed)) { ++ parsed.options.hideThinkingSummary = false; ++ } + if (preserveAnthropicResponseModel) parsed._responseModelId = responseModelId; + logCtx.model = route.modelId; + logCtx.provider = route.providerName; +diff --git a/src/types/provider.ts b/src/types/provider.ts +index e65130a4fa..b6374a991a 100644 +--- a/src/types/provider.ts ++++ b/src/types/provider.ts +@@ -746,6 +746,15 @@ export interface OcxProviderConfig { + * out explicitly (e.g. MiniMax, where low effort disables thinking). + */ + requiresReasoningPlaceholderModels?: string[]; ++ /** ++ * Opt-in: surface upstream thinking as visible reasoning summaries even when the ++ * client did not send `reasoning.summary`. parseRequest hides thinking by default ++ * (Codex omits the field), which strands genuine reasoning — e.g. Gemini `thought` ++ * parts on the google-antigravity (Cloud Code Assist) wire — in hidden replay ++ * envelopes. An explicit client `reasoning.summary: "none"` still wins. Set `false` ++ * to opt a seeded preset back out. ++ */ ++ showThinkingSummary?: boolean; + /** + * Opt-in same-target 429 retry policy. Codex itself never retries 429 (it retries 5xx only, + * openai/codex#30471), and single-key pools have no failover, so the proxy waits and replays + +``` + +## Reflection corrections accepted + +Explicit wire reasoning.summary:"none" wins. A client that serializes configured none as omission cannot be distinguished from unspecified preference. No client config rewrite or global catalog summary default changes. Summary classification is limited to built CCA Gemini requests; unknown/uninitialized, direct Google/Vertex and CCA Claude/gpt-oss remain raw. Streaming and buffered summary-to-tool continuations assert exact Google signature on correct call; never emit Google signatures as Anthropic thinking_signature. Hidden unsigned summaries may disappear but required tool replay state survives. Exercise final assistant text and terminal order, fallback in both directions, and remove replay-comparison rewrite alongside SSE/JSON rewrite. Desktop appearance remains client-controlled; source patch comments claiming an unconditional placeholder are replaced during adoption. + +## Presentation P revalidation + +Prior D: roadmap locked; next presentation implementation. Both source patches apply to baseline; combined application requires keeping the newer no-rewrite expectation. Shared classifier signatures remain current. Implement CCA-only provider default by recomputing parsed.options.hideThinkingSummary from raw summary each final route for inboundWire responses; other inbound types preserve their existing flag. Missing raw request leaves original hide flag authoritative. CCA Gemini classification records boolean in existing per-request adapter closure on each build; default false. + +C review corrections: fixture summary emission now depends on includeThoughts; provider false + client auto does not request upstream summaries. Signature regression exercises actual SSE and JSON serializers, visible/hidden modes, a tool-ending first turn, exact signature and real matching functionResponse. Final-answer behavior is covered separately by bridge and combo tests. diff --git a/devlog/_plan/260912_thinking_contract/020_transport_hint.md b/devlog/_plan/260912_thinking_contract/020_transport_hint.md new file mode 100644 index 0000000000..abfff918c7 --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/020_transport_hint.md @@ -0,0 +1,323 @@ +# Optional hint suppression + +Class C4 review because client metadata policy changes. Independent of presentation; depends only on roadmap. Adopt #3652 only after independent security/transport review. Public proposal removes exactly two x-codex-safety-buffering headers, metadata.type=safety_buffering events and top-level safety_buffering fields at the client relay boundary. Default false; malformed config must remain off and candidate validation rejects nonbooleans. This suppresses optional transport hints; provider safety decisions/refusals and upstream checks are unchanged. Compact and independent WS/other-provider pathways retain existing policy unless a directly exercised shared boundary already applies. + +MODIFY src/config.ts and src/types/config.ts for validated boolean/default; src/server/relay.ts for allowlisted header removal and SSE terminal-boundary transformation; relay-eager.ts for option forwarding; core.ts to compute option only for canonical OpenAI forward destination and pass it to all relevant headers/client output boundaries; index.ts exports if needed by existing test style. Do not apply to custom gateway/key providers. Preserve errors, response.failed/incomplete and terminal sentinel handling. + +MODIFY tests/responses/passthrough-headers.test.ts, openai-responses-passthrough.test.ts and tests/server/config.test.ts. Scenarios: absent/false/true/malformed config; uppercase headers; unrelated headers; split metadata frames; actual failure carrying hint must still fail; noncanonical provider has identical fields and retains them; eager/non-eager client paths. Add missing canonical route coverage if independent review identifies it. MODIFY English/ja/ko/ru/zh-cn server configuration docs and structure owners. Avoid unsupported claims about models being weaker or provider safety bypass. + +Before/after anchor: createSseTerminalOutputBoundary() -> createSseTerminalOutputBoundary(options?: CodexSafetyBufferingFilterOptions); sanitizePassthroughHeaders(upstream) -> sanitizePassthroughHeaders(upstream, options?); canonical true => filter option, every other provider => undefined. Full public source diff is pinned by #3652 head in 000_plan.md and inspected locally; any needed correction is recorded here before B. + +## Independent design corrections + +H1 accepted: policy rewrite and hint stripping compose. Build policyFailurePayload first, then remove top-level safety_buffering from the effective emitted payload, preserving response.failed/error data and retryable:false. H2 accepted: extend current relaySseWithFailedTail fourth options object with terminalBoundary; never replace upstreamError. Core passes both existing upstreamError and new terminalBoundary; update the existing source-contract assertion to preserve its original guarantee. H3 accepted: native WebSocket codex.response.metadata.headers and /responses/compact are explicitly excluded; their hints remain unfiltered. No new WS metadata filter. Docs must not claim the old WS allowlist excludes these headers. Regression fixtures cover CRLF/split/malformed input, policy error plus hint, EOF upstreamError, canonical true and noncanonical preservation. + +## Pinned source hunks (apply with corrections above) + +```diff +diff --git a/src/config.ts b/src/config.ts +index fdcda9547c..cd0641feb2 100644 +--- a/src/config.ts ++++ b/src/config.ts +@@ -1125,6 +1125,8 @@ const configSchema = z.object({ + configRebaseProvenance: z.unknown().optional(), + // A retry can be billable, so absence and malformed hand edits both stay off. + emptyCompletionRetry: z.boolean().optional().catch(false), ++ // Header suppression changes what Codex sees, so absence and malformed edits stay off. ++ dropCodexSafetyBuffering: z.boolean().optional().catch(false), + // A malformed hand edit must not silently stop opening the browser: fall back + // to undefined, which resolves to the historical auto-open behavior. + oauthOpenBrowser: z.boolean().optional().catch(undefined), +@@ -2613,6 +2615,14 @@ function emptyCompletionRetryError(value: unknown): string | null { + return "schema_invalid: emptyCompletionRetry: must be a boolean or omitted"; + } + ++function dropCodexSafetyBufferingError(value: unknown): string | null { ++ const raw = rawConfigRecord(value); ++ if (!raw || !Object.hasOwn(raw, "dropCodexSafetyBuffering")) return null; ++ const enabled = raw.dropCodexSafetyBuffering; ++ if (enabled === undefined || typeof enabled === "boolean") return null; ++ return "schema_invalid: dropCodexSafetyBuffering: must be a boolean or omitted"; ++} ++ + function oauthOpenBrowserError(value: unknown): string | null { + const raw = rawConfigRecord(value); + if (!raw || !Object.hasOwn(raw, "oauthOpenBrowser")) return null; +@@ -2718,6 +2728,7 @@ export function validateConfigCandidate(value: unknown): { ok: true; config: Ocx + ?? codexQuotaAutoRefreshError(value) + ?? codexAccountPickerEnabledError(value) + ?? emptyCompletionRetryError(value) ++ ?? dropCodexSafetyBufferingError(value) + ?? oauthOpenBrowserError(value) + ?? runtimeRoleError(value) + ?? remoteGuiConfigError(value) +@@ -3684,6 +3695,7 @@ export function getDefaultConfig(): OcxConfig { + return { + port: 10100, + emptyCompletionRetry: false, ++ dropCodexSafetyBuffering: false, + managementUsageMaxReadBytes: 64 * 1024 * 1024, + appOwnedMemoryBudgetMb: DEFAULT_APP_OWNED_MEMORY_BUDGET_BYTES / (1024 * 1024), + // Fresh/re-initialized configs are already written in the current three-tier +diff --git a/src/server/index.ts b/src/server/index.ts +index aedd6bf236..c6ce73b2f1 100644 +--- a/src/server/index.ts ++++ b/src/server/index.ts +@@ -142,6 +142,7 @@ import { + } from "./relay"; + export { + consumeForInspection, ++ codexSafetyBufferingFilterOptions, + relaySseWithFailedTail, + relaySseWithHeartbeat, + relayWithAbort, +diff --git a/src/server/relay-eager.ts b/src/server/relay-eager.ts +index 655997b813..a6e60d3d02 100644 +--- a/src/server/relay-eager.ts ++++ b/src/server/relay-eager.ts +@@ -26,6 +26,7 @@ + + import { + adapterEofIncompleteFrame, ++ type CodexSafetyBufferingFilterOptions, + createSseTerminalOutputBoundary, + doneFrame, + failedTailFrame, +@@ -83,6 +84,8 @@ export type EagerRelayOptions = { + postCancelDrainBytes?: number; + /** Injectable clock for tests. */ + now?: () => number; ++ /** Client output boundary filters (Codex safety-buffering hints). */ ++ terminalBoundary?: CodexSafetyBufferingFilterOptions; + }; + + const DEFAULT_MAX_QUEUE_BYTES = 8 * 1024 * 1024; +@@ -111,7 +114,7 @@ export function relaySseEagerBounded( + const terminalEncoder = new TextEncoder(); + const adapterEofFrame = adapterEofIncompleteFrame(terminalEncoder); + const terminalSentinel = doneFrame(terminalEncoder); +- const terminalBoundary = createSseTerminalOutputBoundary(); ++ const terminalBoundary = createSseTerminalOutputBoundary(opts?.terminalBoundary); + const activeRewrite: SseBlockRewrite | undefined = hooks.rewriteBlocks + ?? (hooks.rewritePayload ? payloadRewriteAsBlockRewrite(hooks.rewritePayload) : undefined); + const encodeFailedTail = (error: unknown): Uint8Array | null => { +diff --git a/src/server/relay.ts b/src/server/relay.ts +index 60b57ea025..d840b2e59c 100644 +--- a/src/server/relay.ts ++++ b/src/server/relay.ts +@@ -162,7 +162,10 @@ export type SseTerminalOutputBoundary = { + * terminal, and drops every later block/byte. A premature [DONE] is held until + * a terminal arrives so clean EOF can synthesize one terminal and one sentinel. + */ +-export function createSseTerminalOutputBoundary(): SseTerminalOutputBoundary { ++export function createSseTerminalOutputBoundary( ++ options?: CodexSafetyBufferingFilterOptions, ++): SseTerminalOutputBoundary { ++ const dropSafetyBuffering = options?.dropCodexSafetyBuffering === true; + const decoder = new TextDecoder(); + const encoder = new TextEncoder(); + const framer = new BoundedSseFrameBuffer(MAX_INSPECTION_SSE_FRAME_BYTES); +@@ -181,6 +184,10 @@ export function createSseTerminalOutputBoundary(): SseTerminalOutputBoundary { + const payload = sseDataPayload(decoder.decode(frame.block)); + const isDone = payload === "[DONE]"; + const parsed = payload === null ? undefined : parseSsePayload(payload); ++ const safetyBuffering = dropSafetyBuffering && parsed !== undefined ++ ? codexSafetyBufferingBlockAction(parsed) ++ : "keep"; ++ if (safetyBuffering === "drop") continue; + const policyError = parsed !== undefined && isPolicyRewriteType(parsed) + ? cyberPolicyTerminalError(parsed) + : undefined; +@@ -189,7 +196,9 @@ export function createSseTerminalOutputBoundary(): SseTerminalOutputBoundary { + decoder.decode(frame.block), + policyFailurePayload(policyError, parsed), + )) +- : frame.block; ++ : safetyBuffering === "strip" ++ ? encoder.encode(stripCodexSafetyBufferingField(decoder.decode(frame.block), parsed)) ++ : frame.block; + if (isDone) { + done = true; + if (responsesTerminal) { +@@ -260,10 +269,11 @@ export function relaySseWithFailedTail( + body: ReadableStream, + upstream: AbortController, + onClientGone?: (reason?: unknown) => void, ++ boundaryOptions?: CodexSafetyBufferingFilterOptions, + ): ReadableStream { + const reader = body.getReader(); + const encoder = new TextEncoder(); +- const terminalBoundary = createSseTerminalOutputBoundary(); ++ const terminalBoundary = createSseTerminalOutputBoundary(boundaryOptions); + let closed = false; + const relayChunk = ( + controller: ReadableStreamDefaultController, +@@ -438,6 +448,29 @@ function isPolicyRewriteType(parsed: unknown): boolean { + return type === "response.failed" || type === "response.incomplete" || type === "error"; + } + ++/** ++ * Codex emits its safety-buffering hint in the SSE body as well as in headers: ++ * a `response.metadata` event whose `metadata.type` is `safety_buffering`, or a ++ * `safety_buffering` field on another event. The metadata event is dropped whole; ++ * the field is stripped so the carrying event is otherwise relayed unchanged. ++ */ ++function codexSafetyBufferingBlockAction(parsed: unknown): "keep" | "drop" | "strip" { ++ const root = asJsonRecord(parsed); ++ if (!root) return "keep"; ++ if (root.type === "response.metadata") { ++ const metadata = asJsonRecord(root.metadata); ++ if (metadata?.type === "safety_buffering") return "drop"; ++ } ++ return Object.hasOwn(root, "safety_buffering") ? "strip" : "keep"; ++} ++ ++function stripCodexSafetyBufferingField(block: string, parsed: unknown): string { ++ const root = asJsonRecord(parsed); ++ if (!root) return block; ++ const { safety_buffering: _safetyBuffering, ...rest } = root; ++ return replaceSseDataPayload(block, JSON.stringify(rest)); ++} ++ + function rewritePolicyTerminalBlock(block: string, payload: string): string { + const newline = block.includes("\r\n") ? "\r\n" : "\n"; + const rewritten = replaceSseDataPayload(block, payload); +@@ -1422,7 +1455,31 @@ export function consumeForResponseLogMetadata( + * body makes the caller (Codex) double-decode / truncate → "stream error" on every gpt passthrough. + * Drop encoding + hop-by-hop headers; relay everything else (content-type, etc.) verbatim. + */ +-export function sanitizePassthroughHeaders(upstream: Headers): Headers { ++export const CODEX_SAFETY_BUFFERING_HEADERS = [ ++ "x-codex-safety-buffering-enabled", ++ "x-codex-safety-buffering-faster-model", ++] as const; ++ ++const CODEX_SAFETY_BUFFERING_HEADER_SET: ReadonlySet = new Set(CODEX_SAFETY_BUFFERING_HEADERS); ++ ++export interface CodexSafetyBufferingFilterOptions { ++ /** ++ * Drop Codex safety-buffering hints: the `x-codex-safety-buffering-*` response ++ * headers and the `safety_buffering` SSE metadata event / field. Absent and ++ * `false` relay everything unchanged. ++ */ ++ dropCodexSafetyBuffering?: boolean; ++} ++ ++/** Resolve the passthrough header policy from the loaded config (absent means "forward everything"). */ ++export function codexSafetyBufferingFilterOptions( ++ config: { dropCodexSafetyBuffering?: boolean }, ++): CodexSafetyBufferingFilterOptions { ++ return { dropCodexSafetyBuffering: config.dropCodexSafetyBuffering === true }; ++} ++ ++export function sanitizePassthroughHeaders(upstream: Headers, options?: CodexSafetyBufferingFilterOptions): Headers { ++ const dropSafetyBuffering = options?.dropCodexSafetyBuffering === true; + const DROP = new Set([ + "content-encoding", + "content-length", +@@ -1439,7 +1496,10 @@ export function sanitizePassthroughHeaders(upstream: Headers): Headers { + ]); + const out = new Headers(); + upstream.forEach((value, key) => { +- if (!DROP.has(key.toLowerCase())) out.set(key, value); ++ const lower = key.toLowerCase(); ++ if (DROP.has(lower)) return; ++ if (dropSafetyBuffering && CODEX_SAFETY_BUFFERING_HEADER_SET.has(lower)) return; ++ out.set(key, value); + }); + return out; + } +diff --git a/src/server/responses/core.ts b/src/server/responses/core.ts +index 9d0eea0d76..e199917968 100644 +--- a/src/server/responses/core.ts ++++ b/src/server/responses/core.ts +@@ -304,6 +304,7 @@ import { + markEagerRelaySseResponse, + markNativePassthroughSseResponse, + relaySseWithFailedTail, ++ codexSafetyBufferingFilterOptions, + relayWithAbort, + sanitizePassthroughHeaders, + } from "../relay"; +@@ -3850,6 +3851,9 @@ async function handleResponsesInner( + let hostAdmissionLease = pendingHostAdmissionLease; + pendingHostAdmissionLease = null; + try { ++ const codexSafetyBufferingOptions = isCanonicalOpenAiForwardProvider(route.provider) ++ ? codexSafetyBufferingFilterOptions(config) ++ : undefined; + const imageGenCallAliases = route.provider.authMode === "forward" + ? new Map() + : imageGenToolCallAliases(toolBridgeMaps.toolNsMap, parsed._rawBody, translatorBudget); +@@ -4732,7 +4736,7 @@ async function handleResponsesInner( + } + break; + } +- const headers = sanitizePassthroughHeaders(upstreamResponse.headers); ++ const headers = sanitizePassthroughHeaders(upstreamResponse.headers, codexSafetyBufferingOptions); + const resolvedModel = headers.get("openai-model")?.trim(); + if (resolvedModel && !logCtx.preserveResolvedModelFromRoute) logCtx.resolvedModel = resolvedModel; + if (isUsageDebugEnabled()) { +@@ -4824,7 +4828,7 @@ async function handleResponsesInner( + return new Response(upstreamResponse.body, { + status: upstreamResponse.status, + statusText: upstreamResponse.statusText, +- headers: sanitizePassthroughHeaders(upstreamResponse.headers), ++ headers: sanitizePassthroughHeaders(upstreamResponse.headers, codexSafetyBufferingOptions), + }); + } + if (!upstreamResponse.ok) { +@@ -5027,6 +5031,7 @@ async function handleResponsesInner( + onDone: () => unregisterTurn(turnAc), + }, { + clientGoneSignal: options.abortSignal, ++ terminalBoundary: codexSafetyBufferingOptions, + ...(inlineEagerRewrite ? { rewriteBudget: translatorBudget } : {}), + }); + // When selected, this relay closes response.completed even if upstream +@@ -5110,7 +5115,8 @@ async function handleResponsesInner( + const rewrittenBody = clientBlockRewrite !== undefined + ? relaySseWithBlockRewrite(nativeBody, clientBlockRewrite, translatorBudget) + : nativeBody; +- const clientBody = relaySseWithFailedTail(rewrittenBody, upstream, reason => clientGone.abort(reason)); ++ const clientBody = relaySseWithFailedTail(rewrittenBody, upstream, reason => clientGone.abort(reason), ++ codexSafetyBufferingOptions); + return markNativePassthroughSseResponse(new Response(clientBody, { + status: upstreamResponse.status, + headers, +@@ -5238,7 +5244,7 @@ async function handleResponsesInner( + } + throw error; + } +- const sseHeaders = sanitizePassthroughHeaders(headers); ++ const sseHeaders = sanitizePassthroughHeaders(headers, codexSafetyBufferingOptions); + sseHeaders.set("content-type", "text/event-stream"); + sseHeaders.set("cache-control", "no-store"); + return new Response(stream, { +diff --git a/src/types/config.ts b/src/types/config.ts +index 8cf1246979..4d2c63fdf1 100644 +--- a/src/types/config.ts ++++ b/src/types/config.ts +@@ -335,6 +335,16 @@ export interface OcxConfig { + client?: OcxClientConnectionConfig; + /** Opt in to one identical-turn retry when a Responses completion has no text or tool call. */ + emptyCompletionRetry?: boolean; ++ /** ++ * Drop the Codex safety-buffering hints from a Codex Responses passthrough: the ++ * `x-codex-safety-buffering-*` response headers, `response.metadata` SSE events of ++ * type `safety_buffering`, and the `safety_buffering` field on other SSE events. ++ * The Codex TUI turns those hints into a "retry with a faster model" prompt whose ++ * default action switches the session to a weaker model, so an unattended session ++ * can lose its model to a stray keystroke. Absent and `false` relay everything ++ * unchanged. ++ */ ++ dropCodexSafetyBuffering?: boolean; + /** + * Whether a login may open a browser on the machine running the proxy. + * + +``` + +## Hint P revalidation + +Previous D: presentation source complete; final hosted CI remains in delivery. This branch starts from the common docs checkpoint bd34120180 and baseline product 69e3dcda. Original #3652 does not apply cleanly because relay upstreamError handling changed. Carry nonconflicting hunks and manually adapt relay/core/config hunks, preserving cancellation and error capture. Independent H1-H3 plan reflection ALIGNED remains applicable. diff --git a/devlog/_plan/260912_thinking_contract/030_spark.md b/devlog/_plan/260912_thinking_contract/030_spark.md new file mode 100644 index 0000000000..3b013f75aa --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/030_spark.md @@ -0,0 +1,74 @@ +# Spark Lite metadata follows body shape + +Class C3 bounded compatibility. Independent of presentation/hint; depends on roadmap. MODIFY src/adapters/openai-responses.ts only inside canonical OpenAI forwarding and final wire model gpt-5.3-codex-spark. Add bodyCarriesLiteToolShape next to existing tool-shape helpers: Array.isArray(body.input) && body.input.some(item => isPlainObject(item) && item.type === "additional_tools" && Array.isArray(item.tools) && item.tools.length > 0). After final Spark body construction, delete all case variants of CODEX_RESPONSES_LITE_HEADER then set it to liteShaped ? "true" : "false". Existing prepareCodexWsRequest projects it onto native frame metadata. + +Before: Spark deletes the header, allowing stale native metadata to survive. After: tool-less/top-level-tool Spark frames advertise false; nonempty Lite catalog frames advertise true despite conflicting inherited header. No retirement, no changes to model availability, no user service changes. + +MODIFY tests/codex-integration/codex-metadata-integrity.test.ts: alias resolved final model, inherited true/false/mixed-case/absent header, Lite tool body true, empty Lite group false, malformed metadata keeps HTTP fallback/body, noncanonical remains unchanged. MODIFY tests/responses/ws-upstream-reuse.test.ts: legacy true socket retires when adapter produces false, replacement same identity reused, raw request immutable. MODIFY all eight existing architecture locale pages and structure/transports/responses.md, referencing body-shape rule from shared area owners. Adopt latest #4130 source diff, preserving author; do not import historical earlier heads. + +Verification: source diff review and final-branch hosted ci.yml lane=all. Tests NOT RUN locally. Success proves framing and connection identity, not a live backend EOF fix or all tool-bearing EOF cases. Remaining acceptance: broader tool-format conversion stays out of scope. + +## Pinned source hunks (apply with corrections above) + +```diff +diff --git a/src/adapters/openai-responses.ts b/src/adapters/openai-responses.ts +index c4aa523ee6..8fbe43816d 100644 +--- a/src/adapters/openai-responses.ts ++++ b/src/adapters/openai-responses.ts +@@ -864,6 +864,21 @@ function promoteClientLoadedTools(body: unknown): unknown { + } + + const MAX_RESPONSES_CALL_ID_LENGTH = 64; ++ ++/** ++ * Whether the outgoing body still delivers tools through the responses-lite shape. ++ * ++ * Lite carries the client catalog as an `additional_tools` input item; the non-Lite wire shape ++ * expects top-level `tools`. Anything that flips the Lite advertisement has to agree with the ++ * shape actually being sent, or the destination silently loses the tool surface. ++ */ ++function bodyCarriesLiteToolShape(body: Record): boolean { ++ if (!Array.isArray(body.input)) return false; ++ return body.input.some(item => ++ isPlainObject(item) && item.type === "additional_tools" ++ && Array.isArray(item.tools) && item.tools.length > 0 ++ ); ++} + const REPAIRED_CALL_ID_PREFIX = "call_ocx_"; + const REPAIRED_CALL_ID_DIGEST_LENGTH = MAX_RESPONSES_CALL_ID_LENGTH - REPAIRED_CALL_ID_PREFIX.length; + +@@ -2515,12 +2530,22 @@ export function createResponsesPassthroughAdapter(provider: OcxProviderConfig): + parsed.modelId, + ); + if (isCanonicalOpenAiForwardProvider(provider)) { +- // Spark closes Responses Lite streams before a terminal completion. Select compatibility +- // from the final wire model so aliases cannot leave the caller or a static header enabled. ++ // Select Spark's Lite compatibility from the final wire model, including aliases, and ++ // let the BODY decide it. The header also overrides native WS metadata downstream, so a ++ // forwarded or statically configured value must never contradict the shape being sent. ++ // ++ // The synchronized catalog keeps `use_responses_lite: true` for Spark precisely because ++ // it selects tool delivery (`input[].additional_tools` instead of top-level `tools`), and ++ // stripSparkCompatibility filters that group in place rather than promoting it. So a ++ // Lite-shaped body is pinned back ON — otherwise an inherited `false` advertises non-Lite ++ // while the tools exist only in the Lite shape, and Spark loses the tool surface. Only a ++ // body with no Lite tool group is downgraded, which is what the stream fix needs. + if (isPlainObject(finalBody) && finalBody.model === "gpt-5.3-codex-spark") { ++ const liteShaped = bodyCarriesLiteToolShape(finalBody); + for (const name of Object.keys(headers)) { + if (name.toLowerCase() === CODEX_RESPONSES_LITE_HEADER) delete headers[name]; + } ++ headers[CODEX_RESPONSES_LITE_HEADER] = liteShaped ? "true" : "false"; + } + const routingHeaders = new Headers(headers); + applyCodexRoutingHint(routingHeaders, finalBody); + +``` + +## Spark P revalidation + +Prior D: hint source/security review PASS, final hosted tests pending. This independent branch starts from bd34120180. Latest #4130 hunks still apply cleanly. CCA summary and hint branches do not modify this adapter. Body-dependent Lite true/false, canonical final wire model and no retirement remain the acceptance contract. + +## Spark design reflection amendments + +S1 accepted: apply the existing modelSuffixBracketStrip normalization to finalBody before deciding Lite, using the same immutable object serialized later. A canonical gpt-5.3-codex-spark[1m] request that strips to Spark gets the policy; a final non-Spark model does not. S2 accepted: all eight architecture paragraphs say nonempty additional_tools tools array, not merely group presence; source PR outstanding documentation finding is addressed. S3 accepted: tests cover catalog filtered empty, surviving functions group, only top-level tools, both alias directions and preserved noncanonical configured Lite. Check both actual serialized body and WS header metadata; shape detection is not tool-support validation. diff --git a/devlog/_plan/260912_thinking_contract/040_delivery.md b/devlog/_plan/260912_thinking_contract/040_delivery.md new file mode 100644 index 0000000000..a70e62bfb6 --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/040_delivery.md @@ -0,0 +1,7 @@ +# Final heads and handoff + +Class C3 delivery evidence. Depends on all dispositions. MODIFY branch-owned numbered completion docs and ignored .tmp/thinking/handoff.md. Read existing .github/PULL_REQUEST_TEMPLATE.md; write every section, credits and precise NOT RUN limitation. Publish only own codex/260912-60plus-thinking* branches with git push --no-verify; PR bases dev for independent units, ordinary parent branch only for actual dependencies. No merge/auto-merge/closures. + +NEW .tmp/thinking/*-ci.json captures gh run view JSON for final SHA plus all jobs. NEW .tmp/thinking/*-review.md captures independent implementation findings with accepted/rebutted disposition. Refresh head/base, native stack membership (unknown if API unsupported), current reviews and CI before handoff. Inspect .github/workflows/ci.yml and dispatch lane=all at each final branch where needed. Existing automatic runs stay untouched. If final-head CI fails, inspect failing logs, repair scoped source or fixtures, commit/push --no-verify and validate new final tip. Do not label skipped/cancelled/old-head runs passing. + +Final handoff fields: own worktree, branch per PR, source PR disposition, exact head, PR URL, dependency order, original author trailers, remaining acceptance, unresolved reviews, CI run id/url/head/result/job conclusions, own cycle records and local tests NOT RUN. Parent performs any subsequent integration. No evidence claims from peer commentary alone. diff --git a/devlog/_plan/260912_thinking_contract/050_refresh.md b/devlog/_plan/260912_thinking_contract/050_refresh.md new file mode 100644 index 0000000000..16d4aa80dd --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/050_refresh.md @@ -0,0 +1,5 @@ +# Integration conflict repair + +Parent explicitly requests own hint branch latest-dev integration with independent resolution-only audit and no-verify push. Latest fetched dev ca5ac39124671ee05349e7873231f672824ea26c also conflicts with presentation; preserve all three independent dev-based PRs. Most collisions are adjacent structure-document additions; core received continuation recovery changes that must survive. No source PR/other worktree modifications or merges into dev. Rebase only owned branches, preserve pre-rebase refs in ignored evidence and compare range-diff; use explicit expected old remote SHA with force-with-lease plus --no-verify. This is branch refresh, not native restacking. Product tests remain NOT RUN. + +Prior D: Spark source audit PASS; hosted tests remain pending. Refresh is a subtask of the already-active delivery cycle; no separate cycle is claimed. MODIFY conflict paths only, retaining source contracts and new dev changes. Verification: git range-diff, git diff --check, docs source validator, independent resolution-only review; hosted tests rerun only on refreshed final heads. diff --git a/docs-site/src/content/docs/fr/guides/combos.md b/docs-site/src/content/docs/fr/guides/combos.md index 9073372327..dc18f68971 100644 --- a/docs-site/src/content/docs/fr/guides/combos.md +++ b/docs-site/src/content/docs/fr/guides/combos.md @@ -218,20 +218,16 @@ Le basculement est intentionnellement limité. Il facilite la disponibilité, l' ## Effort de raisonnement par défaut -`defaultEffort` fournit `reasoning.effort` uniquement lorsque toutes ces conditions sont vraies : +`defaultEffort` complète un `reasoning.effort` absent si le combo possède une valeur par défaut non nulle et si la liste des niveaux acceptés par la cible est connue et non vide. La valeur configurée est conservée si elle est acceptée ; sinon, le niveau accepté le plus élevé ne la dépassant pas est choisi, ou le niveau le plus bas si aucun n’est inférieur. Une liste inconnue ou vide n’ajoute aucune valeur par défaut. -1. le combo a un défaut non nul ; -2. l'appelant n'a pas fait d'effort ; et -3. le catalogue de la cible sélectionnée annonce cet effort précis. +Cette étape conserve un effort existant et les autres champs reasoning. La normalisation des capacités ci-dessous peut supprimer séparément les paramètres effort/thinking non acceptés. Valeurs possibles : `low`, `medium`, `high`, `xhigh`, `max`, `ultra` ; l’absence du champ ou `null` désactive l’ajout. -Si la requête n'a pas d'objet `reasoning`, opencodex en crée un. Si `reasoning` existe sans -`effort`, il préserve les autres champs et ajoute la valeur par défaut. Un effort fourni par l’appelant n’est -jamais écrasé. -Lorsque la capacité cible est inconnue ou n'inclut pas l'effort configuré, opencodex omet le -par défaut et laisse le comportement de la cible inchangé. Les valeurs prises en charge sont `low`, `medium`, -`high`, `xhigh`, `max` et `ultra` ; omettez le champ ou réglez-le sur `null` pour laisser l'effort entièrement à -l'appelant et la cible. +## Capacités reasoning mixtes + +`reasoningEffortMode` vaut `"strict"` par défaut : le catalogue publie l’intersection des listes effort de toutes les cibles, y compris les listes explicitement vides. `"adaptive"` exclut ces listes vides pour conserver le sélecteur dans un combo mixte. Une liste inconnue ne limite l’intersection dans aucun des deux modes. + +À l’envoi, une liste explicitement vide supprime les paramètres effort et thinking dans les deux modes ; une liste inconnue les supprime uniquement en adaptive. `reasoning.summary` et les autres champs hors effort sont conservés. La résolution des cibles connues non vides reste inchangée. Les cibles inconnues en strict et les déclarations inconnues du native Chat ordinaire conservent les paramètres de l’appelant. L’ajout d’une valeur par défaut ne remplace pas un effort existant, mais cette normalisation peut supprimer les paramètres non pris en charge. ## Capacité d’entrée d’images / multimodale @@ -337,6 +333,7 @@ Les combos sont stockés dans l'objet `combos` de niveau supérieur, saisi par l | `strategy` | Non | `"failover"` | Valeurs autorisées : `"failover"`, `"round-robin"`, `"random"`, `"least-used"` et `"reset-window"`. | | `stickyLimit` | Non | `1` | Nombre entier de 1 à 100 requêtes réussies par sélection à tour de rôle. S’applique uniquement à `round-robin`. | | `defaultEffort` | Non | `null` | `low`, `medium`, `high`, `xhigh`, `max` ou `ultra` ; appliqué uniquement lorsque l'appelant omet ses efforts et que la cible annonce son soutien. | +| `reasoningEffortMode` | Non | `"strict"` | `strict` ou `adaptive` ; choisit l’intersection des capacités et la normalisation par cible. | | `imageInput` | Non | `"auto"` | `"auto"` ou `"disabled"`. `"auto"` publie les images uniquement si toutes les cibles les prennent en charge ; `"disabled"` impose le texte seul, retire les images des modalités publiées et rejette les requêtes qui en contiennent avant leur distribution. | | `alias` | Non | aucun | Identifiant de modèle public tronqué facultatif ; utilisez les règles d'alias ci-dessus. Une valeur vide est stockée sans alias. | | `nativeAlias` | Non | `false` | Autoriser explicitement un `alias` natif nu actuellement pris en charge à avoir la priorité sur le routage et le catalogue. Jamais déduit de l'alias. | diff --git a/docs-site/src/content/docs/fr/reference/architecture.md b/docs-site/src/content/docs/fr/reference/architecture.md index f197d85c11..746ea7f726 100644 --- a/docs-site/src/content/docs/fr/reference/architecture.md +++ b/docs-site/src/content/docs/fr/reference/architecture.md @@ -89,6 +89,15 @@ Par défaut, `server/index.ts` sert HTTP/SSE sur `/v1/responses`. Si Codex tente Indépendamment de ce réglage côté client, les requêtes canoniques transmises à ChatGPT avec `stream: true` à la racine peuvent utiliser le transport WebSocket en amont de Codex avec une version stable de Bun 1.4.0 ou ultérieure. La version intégrée Bun 1.3.14, les préversions et les identités de runtime impossibles à vérifier utilisent HTTP/SSE. Les réponses WS en amont qui réussissent conservent le contrat SSE en aval et contournent `tee()` au moyen d’un relais borné à lecteur unique et avide (4 MiB par trame brute/enveloppée et une file de production de 8 MiB). Le dépassement de la file ferme la connexion en amont et émet en aval un événement terminal `response.failed`, suivi de `[DONE]`. +Pour le modèle sortant final `gpt-5.3-codex-spark`, la transmission canonique à ChatGPT +désactive explicitement Responses Lite dans l’en-tête HTTP et les métadonnées natives des +trames WS, même lorsqu’un alias sélectionne Spark — uniquement si le corps sortant ne porte pas +de groupe `additional_tools` contenant un tableau `tools` non vide. Ce groupe EST la forme Lite de livraison des outils : un corps Spark +qui l’utilise conserve Lite ACTIF même si un en-tête appelant ou configuré disait l’inverse. Un changement d’identité Lite retire +l’ancien socket ; les requêtes admissibles suivantes ayant la même identité peuvent réutiliser +le nouveau socket. Les autres modèles et passerelles conservent leur politique Lite. +Des métadonnées natives mal formées entraînent toujours un repli HTTP, sans modifier le corps. + Le compactage du contexte Codex fonctionne avec les modèles routés. `server/responses/compact.ts` traite `POST /v1/responses/compact` en exécutant un tour interne de synthèse routé et en renvoyant un historique compacté, tandis que `responses/parser.ts` et `bridge.ts` traitent les tours de compactage distant v2 `compaction_trigger` en émettant exactement un élément de sortie synthétique `compaction`. ## Mise en cache et catalogue diff --git a/docs-site/src/content/docs/fr/reference/configuration/routing.md b/docs-site/src/content/docs/fr/reference/configuration/routing.md index 110d2fcd66..64713d645c 100644 --- a/docs-site/src/content/docs/fr/reference/configuration/routing.md +++ b/docs-site/src/content/docs/fr/reference/configuration/routing.md @@ -58,7 +58,8 @@ Chaque clé de combinaison est un identifiant conforme à `[A-Za-z0-9][A-Za-z0-9 | `targets` | `{ provider: string; model: string; weight?: number }[]` | requis | Routes concrètes ordonnées. `weight` est compris entre 1 et 10000 et vaut `1` par défaut. | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | Stratégie de sélection. L’ordre des cibles définit la priorité de `failover` ; les poids déterminent les sélections de `round-robin` et de `random` ; `least-used` suit les réussites enregistrées ; `reset-window` suit la réinitialisation de quota la plus proche. | | `stickyLimit?` | `number` | `1` | Nombre de requêtes réussies conservées dans un même lot de rotation. Plage de 1 à 100. | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | non défini | Appliqué uniquement lorsque l’appelant ne précise aucun effort et que la cible sélectionnée annonce le niveau demandé. | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | non défini | `defaultEffort` complète un `reasoning.effort` absent si le combo possède une valeur par défaut non nulle et si la liste des niveaux acceptés par la cible est connue et non vide. La valeur configurée est conservée si elle est acceptée ; sinon, le niveau accepté le plus élevé ne la dépassant pas est choisi, ou le niveau le plus bas si aucun n’est inférieur. Une liste inconnue ou vide n’ajoute aucune valeur par défaut. | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` calcule l’intersection des listes connues, y compris les listes vides ; `"adaptive"` exclut les listes vides. Les listes inconnues ne limitent l’intersection dans aucun des deux modes. À l’envoi, les listes explicitement vides suppriment les paramètres effort/thinking dans les deux modes ; les listes inconnues les suppriment seulement en adaptive. `reasoning.summary` est conservé. La résolution des listes connues non vides ainsi que le choix et l’ordre des cibles restent inchangés. | | `imageInput?` | `"auto" \| "disabled"` | `"auto"` | `"auto"` publie les images uniquement lorsque toutes les cibles les prennent en charge ; `"disabled"` impose le texte seul, retire les images des modalités publiées et rejette les requêtes qui en contiennent avant leur distribution. | | `alias?` | `string` | — | Identifiant public facultatif du modèle, à la place du slug canonique du sélecteur. | | `nativeAlias?` | `boolean` | `false` | Permet à un identifiant natif non qualifié actuellement pris en charge de prendre la priorité uniquement pour cet identifiant. Les identifiants non qualifiés `gpt-5.6-*` utilisent les identifiants Codex Pool/Direct. Les routes qualifiées par un compte restent distinctes. Les routes qualifiées par un fournisseur, telles que `openai-apikey/gpt-5.6-*`, utilisent la route configurée avec sa clé d’API et ne passent jamais par l’alias natif. | diff --git a/docs-site/src/content/docs/guides/combos.md b/docs-site/src/content/docs/guides/combos.md index ef5ccde16c..979a84ffcd 100644 --- a/docs-site/src/content/docs/guides/combos.md +++ b/docs-site/src/content/docs/guides/combos.md @@ -256,20 +256,10 @@ instead of growing memory without a bound. ## Default reasoning effort -`defaultEffort` supplies `reasoning.effort` only when all of these are true: +`defaultEffort` fills an absent `reasoning.effort` when the combo has a non-null default and the selected target has a known, nonempty supported ladder. If the target supports the configured value, it is retained; otherwise the highest supported rung at or below it is used, or the lowest supported rung when none is lower. Unknown or empty ladders omit the default. -1. the combo has a non-null default; -2. the caller did not set an effort; and -3. the selected target's catalog advertises that exact effort. +The default-injection step preserves existing effort and other reasoning fields. Capability normalization can separately remove unsupported effort/thinking controls as described below. Supported defaults are `low`, `medium`, `high`, `xhigh`, `max`, and `ultra`; omit the field or use `null` to disable default injection. -If the request has no `reasoning` object, opencodex creates one. If `reasoning` exists without an -`effort` property, it preserves the other fields and adds the default. A caller-provided effort is -never overwritten. - -When target capability is unknown or does not include the configured effort, opencodex omits the -default and leaves the target's own behavior unchanged. Supported values are `low`, `medium`, -`high`, `xhigh`, `max`, and `ultra`; omit the field or set it to `null` to leave effort entirely to -the caller and target. ### Mixed-capability groups (`reasoningEffortMode`) @@ -297,9 +287,12 @@ as wildcards in both modes. } ``` -The default is `"strict"`, which keeps the original behavior. This setting changes published -catalog metadata only — it does not change target order, failover policy, or which effort a given -target receives at dispatch. In the dashboard it is the **Adaptive reasoning ladder** switch in a +The default is `"strict"`, which keeps the original picker behavior. This setting does not change +target order or failover policy. At dispatch, an explicitly empty target ladder has its unsupported +effort/thinking controls removed in either mode while preserving supported non-effort reasoning fields +such as `reasoning.summary`; `"adaptive"` applies the same normalization to an unknown target +capability, while known non-empty targets keep their existing per-target effort resolution. +In the dashboard it is the **Adaptive reasoning ladder** switch in a combo's Capabilities section. ## Image / multimodal capability @@ -414,7 +407,7 @@ Combos are stored in the top-level `combos` object, keyed by combo id: | `cooldownMs` | No | unset → upstream fallback (5 s for request-rate 429 codes `1302`/`1305`, otherwise 60 s) | Integer from 1 to 600000. When set, applies as the per-target cooldown whenever no usable upstream `Retry-After` or Codex reset signal exists, including request-rate 429s; when unset, uses the upstream fallback. | | `waitForCooldownMs` | No | `0` | Integer from 0 to 600000. Maximum time to wait for the earliest eligible cooling target before returning `combo_unavailable`; abort cancels the wait. | | `defaultEffort` | No | `null` | `low`, `medium`, `high`, `xhigh`, `max`, or `ultra`; applied only when the caller omits effort and the target advertises support. | -| `reasoningEffortMode` | No | `"strict"` | `"strict"` intersects every known target ladder, so one target advertising no effort control empties the combo's picker. `"adaptive"` excludes those empty ladders from the published intersection. Metadata only; dispatch is unchanged. | +| `reasoningEffortMode` | No | `"strict"` | `"strict"` intersects every known target ladder, so one target advertising no effort control empties the combo's picker. `"adaptive"` excludes those empty ladders from the published intersection. At dispatch, explicit empty or adaptive unknown ladders remove unsupported effort/thinking controls while preserving supported non-effort reasoning fields such as `reasoning.summary`; known non-empty targets keep existing effort resolution. | | `imageInput` | No | `"auto"` | `"auto"` or `"disabled"`. `"auto"` publishes image support only when every target supports images; `"disabled"` forces text-only (drops image from published modalities and rejects image-bearing requests before dispatch). | | `alias` | No | none | Optional trimmed public model id; use the alias rules above. An empty value is stored as no alias. | | `nativeAlias` | No | `false` | Explicitly permit a currently supported bare native `alias` to take routing and catalog precedence. Never inferred from the alias. | diff --git a/docs-site/src/content/docs/ja/guides/combos.md b/docs-site/src/content/docs/ja/guides/combos.md index 70c902ec6e..7b2c62a9ec 100644 --- a/docs-site/src/content/docs/ja/guides/combos.md +++ b/docs-site/src/content/docs/ja/guides/combos.md @@ -139,15 +139,16 @@ ocx combo set balanced \ ## デフォルトの推論負荷 -`defaultEffort` は、次のすべてが当てはまる場合にのみ `reasoning.effort` を提供します。 +`defaultEffort` は、コンボの既定値が null でなく、対象の対応リストが既知で空でない場合に、省略された `reasoning.effort` を補います。設定値に対応していればその値を使い、そうでなければ設定値以下で最も高い段階を選びます。それもなければ最も低い対応段階を使います。不明または空のリストでは既定値を省略します。 -1. コンボには null 以外のデフォルトがあります。 -2. 呼び出し側は努力を設定しませんでした。そして -3. 選択したターゲットのカタログは、その正確な取り組みを宣伝します。 +既定値の補完は既存の effort と他の reasoning フィールドを保持します。以下の capability 正規化は、別途、非対応の effort/thinking 制御を削除できます。設定可能な既定値は `low`、`medium`、`high`、`xhigh`、`max`、`ultra` です。省略または `null` で補完を無効にします。 -リクエストに `reasoning` オブジェクトがない場合、opencodex はオブジェクトを作成します。 `reasoning` が `effort` プロパティなしで存在する場合、他のフィールドは保持され、デフォルトが追加されます。呼び出し元が提供した努力は決し​​て上書きされません。 -ターゲットの機能が不明な場合、または設定されたエフォートが含まれていない場合、opencodex はデフォルトを省略し、ターゲット自体の動作を変更しないままにします。サポートされている値は、`low`、`medium`、`high`、`xhigh`、`max`、および `ultra` です。このフィールドを省略するか、`null` に設定して、呼び出し元とターゲットに作業を完全に任せます。 +## 異なる reasoning capability の組み合わせ + +`reasoningEffortMode` の既定値は `"strict"` です。明示的な空リストを含む全対象の effort リストの共通部分を公開します。`"adaptive"` は空リストを共通部分の計算から除外し、混在するコンボでも選択肢を維持します。不明なリストは、どちらのモードでもカタログの共通部分を制限しません。 + +送信時には、明示的な空リストの対象で effort と thinking の制御を両モードとも削除します。不明な対象で削除するのは adaptive のみです。`reasoning.summary` と effort 以外のフィールドは保持し、既知の空でない対象は従来どおり effort を解決します。strict の不明な対象と通常の native Chat の不明な宣言は保持します。既定値の補完は既存の effort を上書きしませんが、この正規化は非対応の制御を削除できます。 ## 暗号化された v2 サブエージェント タスク @@ -237,6 +238,7 @@ ocx combo remove --yes | `cooldownMs` |いいえ | 未設定 → アップストリーム フォールバック(リクエストレート 429 コード `1302`/`1305` では 5 秒、それ以外では 60 秒) | 1 ~ 600000 の整数。設定時は、使用可能なアップストリーム `Retry-After` または Codex リセットシグナルがない場合に、リクエストレート 429 を含むターゲットごとのクールダウンとして適用されます。未設定時はアップストリーム フォールバックを使用します。 | | `waitForCooldownMs` |いいえ | `0` | 0 ~ 600000 の整数。最も早く利用可能になる冷却中のターゲットを待ってから `combo_unavailable` を返すまでの最大待機時間。中止すると待機はキャンセルされます。 | | `defaultEffort` |いいえ | `null` | `low`、`medium`、`high`、`xhigh`、`max`、または `ultra`;呼び出し元が努力を省略し、ターゲットがサポートをアドバタイズした場合にのみ適用されます。 | +| `reasoningEffortMode` | いいえ | `"strict"` | `strict` または `adaptive`。混在する capability の共通部分と対象別の制御正規化を選択します。 | | `alias` |いいえ |なし |オプションのトリミングされたパブリック モデル ID。上記のエイリアス ルールを使用します。空の値はエイリアスなしで保存されます。 | | `nativeAlias` |いいえ | `false` | 現在サポートされている bare native alias に routing/catalog の優先権を明示的に与えます。 | | `displayName` |いいえ |なし | catalog 表示専用ラベル。`nativeAlias` が true の場合は必須です。 | diff --git a/docs-site/src/content/docs/ja/reference/architecture.md b/docs-site/src/content/docs/ja/reference/architecture.md index cd2ce3db56..3500cbaf86 100644 --- a/docs-site/src/content/docs/ja/reference/architecture.md +++ b/docs-site/src/content/docs/ja/reference/architecture.md @@ -97,6 +97,15 @@ HTTP の境界は `server/index.ts` が担い、Responses データプレーン `server/index.ts` はデフォルトで `/v1/responses` を HTTP/SSE で提供します。`websockets` が `false` の状態で Codex が Responses WebSocket アップグレードを試みると、opencodex は `426 upgrade_required` を返し、Codex はそのセッションで HTTP にフォールバックします。`"websockets": true` を設定すると同じエンドポイントがアップグレードを受け入れ WebSocket ブリッジを使います。 +最終送信モデルが `gpt-5.3-codex-spark` の場合、canonical ChatGPT 転送は HTTP ヘッダーと +ネイティブ WS フレームのメタデータの両方で Responses Lite を明示的に無効にします。 +エイリアスで Spark を選択した場合も同様です。ただし無効化は、送信本文が空でない `tools` 配列を持つ `additional_tools` +項目を持たない場合に限ります。このグループ自体が Lite のツール受け渡し形式なので、それを +使う Spark 本文は呼び出し元や設定のヘッダーに関わらず Lite を有効のまま保ちます。Lite の識別値が変わると古いソケットは退役し、 +以後の条件を満たす同じ識別値のリクエストは新しいソケットを再利用できます。他のモデルと +ゲートウェイの Lite ポリシーは維持されます。不正なネイティブメタデータは引き続き、 +本文を変更せずに HTTP にフォールバックします。 + Codex コンテキスト compaction はルーティングされたモデルでも動作します。`server/responses/compact.ts` は `POST /v1/responses/compact` を内部ルーティング要約ターンとして扱い、圧縮されたヒストリーを返します。 `responses/parser.ts` と `bridge.ts` は remote compaction v2 の `compaction_trigger` ターンを扱い、合成 `compaction` 出力項目を正確に 1 つ送ります。 diff --git a/docs-site/src/content/docs/ja/reference/configuration/routing.md b/docs-site/src/content/docs/ja/reference/configuration/routing.md index 18b238c8bf..c6a6d80772 100644 --- a/docs-site/src/content/docs/ja/reference/configuration/routing.md +++ b/docs-site/src/content/docs/ja/reference/configuration/routing.md @@ -69,7 +69,8 @@ picker catalog の convergence だけが保留中で routing change は失われ | `targets` | `{ provider: string; model: string; weight?: number }[]` |必須 |具体的なルートを指示しました。 `weight` は 1 ~ 10000 で、デフォルトは `1` です。 | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` |選択戦略。ターゲットの順序は `failover` の優先順位となり、`weight` は `round-robin` と `random` の抽選に影響し、`least-used` は記録された成功数に従い、`reset-window` は最も早いクォータリセットに従います。 | | `stickyLimit?` | `number` | `1` |成功したリクエストは 1 つのラウンドロビン バッチに保持されます。範囲は 1 ~ 100。 | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` |設定を解除する |呼び出し元が努力を省略し、選択されたターゲットが要求されたラングをアドバタイズする場合にのみ適用されます。 | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` |設定を解除する | `defaultEffort` は、コンボの既定値が null でなく、対象の対応リストが既知で空でない場合に、省略された `reasoning.effort` を補います。設定値に対応していればその値を使い、そうでなければ設定値以下で最も高い段階を選びます。それもなければ最も低い対応段階を使います。不明または空のリストでは既定値を省略します。 | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` は空リストを含む既知の対応リストの共通部分を公開し、`"adaptive"` は空リストを除外します。不明なリストは両モードで共通部分を制限しません。送信時、明示的な空リストは両モードで effort/thinking 制御を削除し、不明なリストでは adaptive のみ削除します。`reasoning.summary` は保持されます。既知の空でない対象の effort 解決と対象の選択・順序は変わりません。 | | `alias?` | `string` | — |正規のピッカー スラグの代わりのオプションのパブリック モデル ID。 | | `nativeAlias?` | `boolean` | `false` | 現在サポートされている bare native id に限り、その未修飾 id で優先します。アカウント修飾およびプロバイダー修飾の OpenAI ルートは別のままです。 | | `displayName?` | `string` | — | catalog 表示専用ラベル。native alias では空でない値が必須です。 | diff --git a/docs-site/src/content/docs/ja/reference/configuration/server.md b/docs-site/src/content/docs/ja/reference/configuration/server.md index 0dd9cf59e4..e86387f80e 100644 --- a/docs-site/src/content/docs/ja/reference/configuration/server.md +++ b/docs-site/src/content/docs/ja/reference/configuration/server.md @@ -13,6 +13,7 @@ description: リスナー、リモート アクセス、アドミッション | `hostname?` | `string` | `"127.0.0.1"` |バインドアドレス。非ループバック バインドには `OPENCODEX_API_AUTH_TOKEN` が必要です。 | | `proxy?` | `string` | — |送信 HTTP(S) プロキシ URL または `${ENV_VAR}`。これらの変数が設定されていない場合にのみ、`HTTP_PROXY` / `HTTPS_PROXY` に適用されます。ループバックは `NO_PROXY` に残ります。 | | `emptyCompletionRetry?` | `boolean` | `false` | テキストもツール呼び出しもない Responses ターンを、ターミナルイベント前にストリームが終了した場合も含め、同一リクエストで 1 回再試行するよう明示的に有効化します。再試行は課金対象になる場合があります。`OCX_EMPTY_COMPLETION_RETRY=0` で設定を変更せず無効化できます。combo と routed-compaction turn は対象外です。 | +| `dropCodexSafetyBuffering?` | `boolean` | `false` | Codex Responses パススルーから Codex の safety-buffering ヒントを除去します。対象は `x-codex-safety-buffering-enabled` / `x-codex-safety-buffering-faster-model` 応答ヘッダー、`safety_buffering` 型の `response.metadata` SSE イベント、およびその他の SSE イベントにある `safety_buffering` フィールドです。Codex TUI はこれらを、既定の操作でセッションをより弱いモデルに切り替える「より高速なモデルで再試行」プロンプトとして表示します。その他の `x-codex-*` ヘッダーと SSE イベントの内容は、そのフィールドの除去を除いて変更せずに転送されます。既定ではオフです。 | | `stallTimeoutSec?` | `number` | `300` | `response.incomplete` より前にアップストリーム データがない秒数。最小 1。 | `connectTimeoutMs?` | `number` | `200000` |試行ごとの DNS/TCP/TLS/最終ヘッダーの期限。本体が生成される前に終了します。 | | `shutdownTimeoutMs?` | `number` | `5000` |アクティブなターンが中止される前の正常な排出期限。 | @@ -171,3 +172,5 @@ Anthropic OAuth サイドカーは、opencodex の既存のクロード コー ## Codex クォータのネットワーク診断 メイン Codex アカウント行の `quotaRefresh` はクォータ取得の診断情報であり、残量やモデルへのアクセス権を示すものではありません。キャッシュ利用時や取得を行わない場合は省略されることがあります。取得には操作中のシェルではなく、実行中のプロキシサービスの環境が使われます。`proxy` 未設定では既存の環境を維持し、`"auto"` は起動時に Windows の静的プロキシ設定だけを読みます。PAC/WPAD、SOCKS のみの設定、実行中の変更は自動反映されません。TUN での成功だけでは HTTP プロキシ経路の正常性は確認できません。[コマンドと状態の説明(英語)](/reference/configuration/server/#codex-quota-network-diagnostics)を参照してください。 + +`dropCodexSafetyBuffering`: プロバイダーの安全性の適用と拒否応答は変更しません。native `codex.response.metadata.headers` WebSocket メタデータと `/responses/compact` は対象外です。 diff --git a/docs-site/src/content/docs/ko/guides/combos.md b/docs-site/src/content/docs/ko/guides/combos.md index 80eac32c2d..e507889c6f 100644 --- a/docs-site/src/content/docs/ko/guides/combos.md +++ b/docs-site/src/content/docs/ko/guides/combos.md @@ -145,15 +145,16 @@ ocx combo set balanced \ ## 기본 reasoning effort -`defaultEffort`는 다음 조건이 모두 참일 때만 `reasoning.effort`를 채웁니다. +`defaultEffort`는 콤보 기본값이 null이 아니고, 선택한 대상의 지원 목록이 알려져 있으며 비어 있지 않을 때 생략된 `reasoning.effort`를 채웁니다. 설정값을 지원하면 그대로 사용합니다. 그렇지 않으면 설정값 이하의 가장 높은 지원 단계를 사용하고, 그런 단계가 없으면 가장 낮은 지원 단계를 사용합니다. 지원 목록이 없거나 비어 있으면 기본값을 생략합니다. -1. 콤보에 null이 아닌 기본값이 있습니다. -2. 호출자가 effort를 설정하지 않았습니다. -3. 선택된 대상의 카탈로그가 그 정확한 effort를 광고합니다. +기본값 주입은 기존 effort와 다른 reasoning 필드를 보존합니다. 아래의 capability 정규화는 별도로 지원되지 않는 effort·thinking 제어를 제거할 수 있습니다. 기본값은 `low`, `medium`, `high`, `xhigh`, `max`, `ultra`이며, 필드를 생략하거나 `null`로 설정하면 주입하지 않습니다. -요청에 `reasoning` 객체가 없으면 opencodex가 새로 만듭니다. `reasoning`은 있지만 `effort` 속성이 없으면 다른 필드는 그대로 두고 기본값만 추가합니다. 호출자가 준 effort는 절대 덮어쓰지 않습니다. -대상 기능을 알 수 없거나 설정한 effort를 포함하지 않으면 opencodex는 기본값을 생략하고 대상의 동작은 그대로 둡니다. 지원 값은 `low`, `medium`, `high`, `xhigh`, `max`, `ultra`입니다. effort를 호출자와 대상에 완전히 맡기려면 이 필드를 생략하거나 `null`로 설정하십시오. +## 서로 다른 reasoning capability + +`reasoningEffortMode`의 기본값은 `"strict"`입니다. 모든 대상의 effort 목록을 교집합으로 계산하므로 명시적 빈 목록도 반영합니다. `"adaptive"`는 빈 목록을 교집합에서 제외해 혼합 콤보에서도 선택기를 유지합니다. 알 수 없는 목록은 두 모드 모두 카탈로그 교집합을 제한하지 않습니다. + +전송 시 명시적 빈 목록은 두 모드 모두에서 effort·thinking 제어를 제거하고, 알 수 없는 목록은 adaptive에서만 제거합니다. `reasoning.summary`와 다른 비-effort 필드는 보존하며, 알려진 비어 있지 않은 대상은 기존 방식으로 effort를 결정합니다. strict의 unknown 대상과 일반 native Chat의 unknown 선언은 그대로 유지됩니다. 기본값 주입은 기존 effort를 덮어쓰지 않지만, 이 capability 정규화는 지원되지 않는 제어를 제거할 수 있습니다. ## 암호화된 v2 서브에이전트 작업 @@ -241,6 +242,7 @@ ocx combo remove --yes | `cooldownMs` | 아니요 | 미설정 → 업스트림 폴백(요청 속도 제한 429 코드 `1302`/`1305`는 5초, 그 외는 60초) | 1에서 600000 사이의 정수입니다. 설정하면 사용 가능한 업스트림 `Retry-After` 또는 Codex 재설정 신호가 없을 때 요청 속도 제한 429를 포함한 대상별 쿨다운으로 적용됩니다. 설정하지 않으면 업스트림 폴백을 사용합니다. | | `waitForCooldownMs` | 아니요 | `0` | 0에서 600000 사이의 정수입니다. `combo_unavailable`을 반환하기 전에 가장 먼저 적합해지는 쿨다운 중인 대상을 기다리는 최대 시간입니다. 중단하면 대기가 취소됩니다. | | `defaultEffort` | 아니요 | `null` | `low`, `medium`, `high`, `xhigh`, `max`, 또는 `ultra`입니다. 호출자가 effort를 생략하고 대상이 지원을 광고할 때만 적용됩니다. | +| `reasoningEffortMode` | 아니요 | `"strict"` | `strict` 또는 `adaptive`; 혼합 capability의 교집합과 대상별 제어 정규화를 선택합니다. | | `alias` | 아니요 | 없음 | 선택적으로 앞뒤 공백을 제거한 공개 모델 ID입니다. 위의 alias 규칙을 따릅니다. 빈 값은 alias 없음으로 저장됩니다. | | `nativeAlias` | 아니요 | `false` | 현재 지원되는 bare native alias가 routing/catalog 우선권을 갖도록 명시적으로 허용합니다. | | `displayName` | 아니요 | 없음 | catalog 표시 전용 label입니다. `nativeAlias`가 true이면 필수입니다. | diff --git a/docs-site/src/content/docs/ko/reference/architecture.md b/docs-site/src/content/docs/ko/reference/architecture.md index fa3056fb49..d95bf391ad 100644 --- a/docs-site/src/content/docs/ko/reference/architecture.md +++ b/docs-site/src/content/docs/ko/reference/architecture.md @@ -127,6 +127,15 @@ envelope를 각각 4 MiB로 제한하고 8 MiB producer queue 상한이 있는 b relay를 거칩니다. queue overflow 시 업스트림을 닫고 downstream에는 terminal `response.failed` 이벤트와 `[DONE]`을 내보냅니다. +최종 전송 모델이 `gpt-5.3-codex-spark`이면 canonical ChatGPT forward 경로는 HTTP 헤더와 +네이티브 WS 프레임 메타데이터 모두에서 Responses Lite를 명시적으로 끕니다. 별칭으로 Spark를 +선택해도 동일합니다. 다만 이 비활성화는 전송 본문에 비어 있지 않은 `tools` 배열을 가진 `additional_tools` 항목이 없을 때만 +적용됩니다. 이 그룹 자체가 Lite의 도구 전달 형식이므로, 그것을 사용하는 Spark 본문은 호출자나 +설정 헤더가 무엇이든 Lite를 켠 상태로 유지합니다. Lite 식별값이 바뀌면 기존 소켓은 사용을 종료하며, 이후 같은 식별값으로 +재사용 조건을 충족하는 요청은 새 소켓을 재사용할 수 있습니다. 다른 모델과 게이트웨이는 기존 +Lite 정책을 유지합니다. 네이티브 메타데이터 형식이 잘못된 경우에는 본문을 바꾸지 않고 +기존처럼 HTTP로 폴백합니다. + Codex 컨텍스트 compaction은 라우팅된 모델에서도 동작합니다. `server/responses/compact.ts`는 `POST /v1/responses/compact`를 내부 라우팅 요약 턴으로 처리해 압축된 히스토리를 반환합니다. `responses/parser.ts`와 `bridge.ts`는 remote compaction v2의 `compaction_trigger` 턴을 처리해 diff --git a/docs-site/src/content/docs/ko/reference/configuration/routing.md b/docs-site/src/content/docs/ko/reference/configuration/routing.md index cfc20d07fc..2f617695ce 100644 --- a/docs-site/src/content/docs/ko/reference/configuration/routing.md +++ b/docs-site/src/content/docs/ko/reference/configuration/routing.md @@ -68,7 +68,8 @@ Codex Auth 페이지에서 이 picker 동작을 opt-in할 수 있습니다. 비 | `targets` | `{ provider: string; model: string; weight?: number }[]` | required | 순서가 있는 concrete route입니다. `weight`는 1–10000이며 기본값은 `1`입니다. | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | 선택 전략입니다. 대상 순서는 `failover` 우선순위이고, 가중치는 `round-robin`과 `random` 추첨 비율을 결정하며, `least-used`는 기록된 성공 횟수를 따르고, `reset-window`는 가장 가까운 할당량 재설정을 따릅니다. | | `stickyLimit?` | `number` | `1` | 한 round-robin 배치에서 유지되는 성공 요청 수입니다. 범위는 1–100입니다. | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | 호출자가 effort를 생략했고 선택된 대상이 요청한 rung를 광고할 때만 적용됩니다. | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | `defaultEffort`는 콤보 기본값이 null이 아니고, 선택한 대상의 지원 목록이 알려져 있으며 비어 있지 않을 때 생략된 `reasoning.effort`를 채웁니다. 설정값을 지원하면 그대로 사용합니다. 그렇지 않으면 설정값 이하의 가장 높은 지원 단계를 사용하고, 그런 단계가 없으면 가장 낮은 지원 단계를 사용합니다. 지원 목록이 없거나 비어 있으면 기본값을 생략합니다. | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"`는 빈 목록을 포함한 알려진 대상 지원 목록의 교집합을 사용하고, `"adaptive"`는 빈 목록을 제외합니다. 알 수 없는 목록은 두 모드 모두 교집합을 제한하지 않습니다. 전송 시 명시적 빈 목록은 두 모드에서 effort·thinking 제어를 제거하고, 알 수 없는 목록은 adaptive에서만 제거합니다. `reasoning.summary`는 보존됩니다. 알려진 비어 있지 않은 대상의 effort 결정과 대상 선택·순서는 그대로입니다. | | `alias?` | `string` | — | 정규화된 picker slug 대신 쓰는 선택적 공개 model id입니다. | | `nativeAlias?` | `boolean` | `false` | 현재 지원되는 bare native id가 해당 비수식 id에만 우선하도록 합니다. 계정 또는 프로바이더로 수식된 OpenAI route는 별도로 유지됩니다. | | `displayName?` | `string` | — | catalog 표시 전용 label이며 native alias에서는 비어 있지 않아야 합니다. | diff --git a/docs-site/src/content/docs/ko/reference/configuration/server.md b/docs-site/src/content/docs/ko/reference/configuration/server.md index 577379fed9..5f60fc7c6c 100644 --- a/docs-site/src/content/docs/ko/reference/configuration/server.md +++ b/docs-site/src/content/docs/ko/reference/configuration/server.md @@ -13,6 +13,7 @@ description: 리스너, 원격 접근, admission 키, 타임아웃, 저장소, | `hostname?` | `string` | `"127.0.0.1"` | 바인드 주소입니다. 루프백이 아닌 바인드에는 데이터 admission 토큰이 필요하며, `OPENCODEX_API_AUTH_TOKEN` → `OCX_API_TOKEN_FILE` → 설치된 owner-only `service-api-token` 순서로 결정됩니다. 손으로 내보낼 값은 없습니다. [Remote access](#remote-access)를 보세요. | | `proxy?` | `string` | — | 송신용 HTTP(S) 프록시 URL 또는 `${ENV_VAR}`입니다. 해당 변수가 비어 있을 때만 `HTTP_PROXY` / `HTTPS_PROXY`에 적용되며, 루프백은 `NO_PROXY`에 그대로 남습니다. | | `emptyCompletionRetry?` | `boolean` | `false` | 텍스트나 도구 호출이 없는 Responses 턴을, 터미널 이벤트 전에 스트림이 종료된 경우를 포함해 동일한 요청으로 한 번 재시도하도록 선택합니다. 재시도에는 비용이 발생할 수 있습니다. `OCX_EMPTY_COMPLETION_RETRY=0`은 설정을 바꾸지 않고 비활성화하며, combo 및 routed-compaction turn은 제외됩니다. | +| `dropCodexSafetyBuffering?` | `boolean` | `false` | Canonical Codex Responses 응답의 선택적 safety-buffering 헤더 두 개와 SSE 힌트를 제거합니다. 공급자의 안전 정책이나 거절 응답은 바뀌지 않습니다. Native WS 메타데이터와 compact는 제외됩니다. | | `stallTimeoutSec?` | `number` | `300` | 업스트림 데이터가 없을 때 `response.incomplete`가 되기까지의 초 수입니다. 최소 1입니다. | | `connectTimeoutMs?` | `number` | `200000` | 시도별 DNS/TCP/TLS/최종 헤더 기한입니다. 본문 생성 전에 끝납니다. | | `shutdownTimeoutMs?` | `number` | `5000` | 진행 중인 turn을 중단하기 전에 허용하는 정상 종료 드레인 기한입니다. | diff --git a/docs-site/src/content/docs/reference/architecture.md b/docs-site/src/content/docs/reference/architecture.md index e0fbc8bcba..a9bdf48978 100644 --- a/docs-site/src/content/docs/reference/architecture.md +++ b/docs-site/src/content/docs/reference/architecture.md @@ -157,6 +157,15 @@ upstream WS responses keep the downstream SSE contract and bypass `tee()` throug single-reader relay (4 MiB per raw/enveloped frame and an 8 MiB producer queue). Queue overflow closes the upstream and emits a terminal downstream `response.failed` event followed by `[DONE]`. +For the final outgoing model `gpt-5.3-codex-spark`, canonical ChatGPT forwarding explicitly +disables Responses Lite in both the HTTP header and native WS frame metadata, including when +an alias selects Spark — but only when the outgoing body carries no `additional_tools` item with a nonempty `tools` array. +That group IS the Lite tool-delivery shape, so a Spark body that still uses it keeps Lite ON even +if a caller or configured header said otherwise; otherwise the frame would advertise non-Lite +while the tools exist only in the Lite shape. A changed Lite identity retires the old socket; subsequent eligible +requests with the same identity can reuse the new socket. Other models and gateways keep +their existing Lite policy. Malformed native metadata still falls back to HTTP with its body unchanged. + When a provider rejects a streaming request with HTTP 413 before SSE begins, OpenCodex emits one terminal `response.failed` event with `context_length_exceeded` instead of relaying the retryable unknown status. This lets Codex stop its reconnect loop and apply its own context-compaction policy diff --git a/docs-site/src/content/docs/reference/configuration/providers.md b/docs-site/src/content/docs/reference/configuration/providers.md index 67e072177c..eb61194c94 100644 --- a/docs-site/src/content/docs/reference/configuration/providers.md +++ b/docs-site/src/content/docs/reference/configuration/providers.md @@ -208,6 +208,7 @@ Providers can expose a built-in shorthand, such as `agy` for `google-antigravity | `preserveReasoningContentModels?` | `string[]` | Models requiring prior assistant `reasoning_content` in chat history. | | `reasoningDetailsModels?` | `string[]` | Models whose endpoint returns thinking as a structured `reasoning_details` array (MiniMax M-series with `reasoning_split`); stream deltas are cumulative snapshots that are prefix-diffed, and preserved reasoning replays as a `reasoning_details` array instead of a `reasoning_content` string. | | `requiresReasoningPlaceholderModels?` | `string[]` | Models whose upstream rejects a tool_call continuation missing `reasoning_content` (DeepSeek thinking mode); a minimal placeholder is injected when the replay cache misses. Defaults to `preserveReasoningContentModels`; set `[]` to opt out. | +| `showThinkingSummary?` | `boolean` | Display provider-authored summaries when a Responses client omits `reasoning.summary`. Explicit wire `"none"` wins; a client that serializes its preference as omission cannot be distinguished. Raw reasoning remains content and is never relabeled as a summary. The `google-antigravity` preset defaults to `true`; explicit `false` disables that default. CCA Gemini requests also opt into `generationConfig.thinkingConfig.includeThoughts` when display is enabled; image, Claude and gpt-oss requests do not. This does not change client configuration or global catalog summary defaults. | | `thinkingToggleModels?` | `string[]` | Chat models using `thinking.enabled` rather than an effort ladder. | | `thinkingBudgetModels?` | `string[]` | Chat models using integer `thinking_budget`; effort maps to a budget fraction. | | `noVisionModels?` | `string[]` | Text-only models sent through the vision sidecar; matching tolerates an Ollama `:size` tag. | diff --git a/docs-site/src/content/docs/reference/configuration/routing.md b/docs-site/src/content/docs/reference/configuration/routing.md index 25bd32f64e..2934c337a6 100644 --- a/docs-site/src/content/docs/reference/configuration/routing.md +++ b/docs-site/src/content/docs/reference/configuration/routing.md @@ -90,8 +90,8 @@ namespace, and cannot use reserved bare native families such as `gpt-*`, `o1-*`, | `stickyLimit?` | `number` | `1` | Successful requests retained in one round-robin batch. Range 1–100. Applies only to round-robin. | | `cooldownMs?` | `number` | unset → upstream fallback (5 s for request-rate 429 codes `1302`/`1305`, otherwise 60 s) | Range 1–600000. When set, applies whenever no usable upstream `Retry-After` or Codex reset signal exists, including request-rate 429s; when unset, uses the upstream fallback. Upstream signals take precedence and all cooldowns are capped at 10 minutes. | | `waitForCooldownMs?` | `number` | `0` | Maximum wait for the earliest eligible cooling target on each selection attempt before returning `combo_unavailable`. Range 0–600000; an abort cancels the wait. | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | Applied only when the caller omits effort and the selected target advertises the requested rung. | -| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` intersects every known target effort ladder, so a target advertising no effort control empties the combo's picker. `"adaptive"` excludes those empty ladders from the published intersection. Picker metadata only; target selection and dispatch are unchanged. | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | `defaultEffort` fills an absent `reasoning.effort` when the combo has a non-null default and the selected target has a known, nonempty supported ladder. If the target supports the configured value, it is retained; otherwise the highest supported rung at or below it is used, or the lowest supported rung when none is lower. Unknown or empty ladders omit the default. | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` intersects all known target ladders, including empty ones; `"adaptive"` excludes empty ladders. Unknown ladders are catalog wildcards in both modes. At dispatch, explicit empty ladders remove effort/thinking controls in both modes; unknown ladders do so only in adaptive. `reasoning.summary` is preserved. Known nonempty targets retain their effort resolution, and target selection/order is unchanged. | | `imageInput?` | `"auto" \| "disabled"` | `"auto"` | `"auto"` publishes image only when every target supports images; `"disabled"` forces text-only (drops image from published modalities and rejects image-bearing requests before dispatch). | | `alias?` | `string` | — | Optional public model id in place of the canonical picker slug. | | `nativeAlias?` | `boolean` | `false` | Let a currently supported bare native id take precedence only for that unqualified id. Bare `gpt-5.6-*` ids use Codex Pool/Direct credentials. Account-qualified routes remain distinct. Provider-qualified routes such as `openai-apikey/gpt-5.6-*` use their configured API-key route and never fall through to the native alias. | diff --git a/docs-site/src/content/docs/reference/configuration/server.md b/docs-site/src/content/docs/reference/configuration/server.md index 108d2e0925..3781d5846e 100644 --- a/docs-site/src/content/docs/reference/configuration/server.md +++ b/docs-site/src/content/docs/reference/configuration/server.md @@ -15,6 +15,7 @@ runs helper features around provider requests. | `proxy?` | `string` | — | Outbound HTTP(S) proxy URL, `${ENV_VAR}`, or `"auto"`. Applied to `HTTP_PROXY` / `HTTPS_PROXY` only when those variables are unset; loopback remains in `NO_PROXY`. `"auto"` reads the Windows system proxy (WinINET `ProxyEnable`/`ProxyServer`, `https=` then `http=` entry) once at process start and logs the host it chose. On other platforms, or when the system proxy is off, SOCKS-only, or unreadable, it uses direct egress and says so. PAC/WPAD and live proxy changes are not followed; restart the service after changing the system proxy. | | `noProxy?` | `string \| string[]` | — | Hosts that bypass `proxy`, merged with inherited `NO_PROXY` and loopback entries. A string may use comma-separated `NO_PROXY` syntax or `${ENV_VAR}`. | | `emptyCompletionRetry?` | `boolean` | `false` | Opt in to one identical Responses retry when a turn has no text or tool call, including a stream that ends before a terminal event. The retry may be billable. `OCX_EMPTY_COMPLETION_RETRY=0` disables it without changing config; combo and routed-compaction turns remain excluded. | +| `dropCodexSafetyBuffering?` | `boolean` | `false` | Remove optional client-facing hints from canonical Codex Responses passthrough: the two `x-codex-safety-buffering-enabled` / `x-codex-safety-buffering-faster-model` response headers, `response.metadata` events whose metadata type is `safety_buffering`, and top-level `safety_buffering` fields. Other headers, response data, policy refusals and failures are preserved. This does not disable provider safety enforcement or upstream buffering. Native `codex.response.metadata.headers` WebSocket metadata and `/responses/compact` are outside this filter. | | `stallTimeoutSec?` | `number` | `300` | Seconds without upstream data before `response.incomplete`. Minimum 1. | | `oauthOpenBrowser?` | `boolean` | `true` | Whether a login may open a browser on the machine running the proxy. Absent and `true` both open, so an existing install is unchanged; only an explicit `false` declines. Decline when you need the authorization link in a different browser profile, or when the dashboard is not on the proxy's machine — the login still starts and the URL is still returned and displayed. `POST /api/oauth/login` and `POST /api/codex-auth/login` accept a per-request `openBrowser` boolean that overrides this, and the dashboard exposes the same choice beside the login button. Device-code flows never open a browser either way. | | `connectTimeoutMs?` | `number` | `200000` | Per-attempt DNS/TCP/TLS/final-header deadline; it ends before body generation. | diff --git a/docs-site/src/content/docs/reference/proxy-formats.md b/docs-site/src/content/docs/reference/proxy-formats.md index 7cec1af766..4c1a9fba21 100644 --- a/docs-site/src/content/docs/reference/proxy-formats.md +++ b/docs-site/src/content/docs/reference/proxy-formats.md @@ -24,6 +24,14 @@ should select among several targets. Credential-bearing model, image, video, and search requests do not automatically follow HTTP redirects, including same-origin redirects. Configure the final upstream API URL instead of a redirecting alias. A redirect does not cause the server to resend credentials or the request body to its destination. The response owner retains its existing error or relay behavior; native Responses and compact routes can return the original 3xx and `Location` to the client. Client redirect behavior is separate from this server transport policy. +## Console upload rejections + +An exact Console or Console Go `Invalid upload request.` HTTP 400 from a canonical +OpenCode Zen/Go generation endpoint receives one retry after 800 ms. The proxy reuses +the same serialized request and records the recovery in Logs. Other 400 errors, +custom destinations, cancellations and repeated upload rejections remain failures. +This does not retry filtered model responses or interrupted streams. + ## Endpoint overview | Client surface | Endpoint | Successful non-stream result | Successful stream or socket result | diff --git a/docs-site/src/content/docs/ru/guides/combos.md b/docs-site/src/content/docs/ru/guides/combos.md index b873f4ed03..ddd16f6175 100644 --- a/docs-site/src/content/docs/ru/guides/combos.md +++ b/docs-site/src/content/docs/ru/guides/combos.md @@ -177,20 +177,16 @@ Failover намеренно ограничен. Он помогает при п ## Effort по умолчанию -`defaultEffort` подставляет `reasoning.effort` только если одновременно выполняются все условия: +`defaultEffort` заполняет отсутствующий `reasoning.effort`, если задан `defaultEffort`, отличный от `null`, и список поддерживаемых уровней цели известен и непуст. Поддерживаемое настроенное значение сохраняется; иначе выбирается максимальный поддерживаемый уровень не выше него, а если такого нет — минимальный поддерживаемый уровень. При неизвестном или пустом списке default не добавляется. -1. у combo задано ненулевое значение по умолчанию; -2. вызывающая сторона сама не указала effort; и -3. каталог выбранной цели объявляет поддержку именно этого effort. +Подстановка сохраняет существующий effort и остальные поля reasoning. Нормализация возможностей ниже может отдельно удалить неподдерживаемые параметры effort/thinking. Значения default: `low`, `medium`, `high`, `xhigh`, `max`, `ultra`; отсутствие поля или `null` отключает подстановку. -Если в запросе нет объекта `reasoning`, opencodex создаёт его. Если `reasoning` есть, но в нём нет -свойства `effort`, остальные поля сохраняются, а значение по умолчанию добавляется. Effort, -заданный вызывающей стороной, никогда не перезаписывается. -Если возможности цели неизвестны или не включают настроенный effort, opencodex опускает значение -по умолчанию и оставляет нативное поведение цели без изменений. Поддерживаются `low`, `medium`, -`high`, `xhigh`, `max` и `ultra`; опустите поле или задайте `null`, чтобы полностью оставить выбор -effort вызывающей стороне и цели. +## Разные возможности reasoning в одном combo + +`reasoningEffortMode` по умолчанию равен `"strict"`: публикуется пересечение списков effort всех целей, включая явно пустые списки. `"adaptive"` исключает пустые списки из пересечения, сохраняя выбор effort для смешанного combo. Неизвестные списки не ограничивают пересечение каталога в обоих режимах. + +При отправке явно пустой список удаляет параметры effort и thinking в обоих режимах; неизвестный список — только в adaptive. `reasoning.summary` и остальные поля, не задающие effort, сохраняются. Для известных непустых списков разрешение effort не меняется. Неизвестные цели в strict и неизвестные объявления обычного native Chat сохраняют параметры вызывающей стороны. Подстановка значения по умолчанию не заменяет существующий effort, но нормализация возможностей может удалить неподдерживаемые параметры. ## Шифрованные задачи подагентов v2 @@ -292,6 +288,7 @@ Combo хранятся в объекте верхнего уровня `combos`, | `cooldownMs` | No | не задано → fallback upstream (5 с для rate-limit 429 с кодами `1302`/`1305`, иначе 60 с) | Целое число от 1 до 600000. Если задано, применяется как cooldown каждой цели, когда нет пригодного upstream `Retry-After` или сигнала сброса Codex, включая rate-limit 429; если не задано, используется fallback upstream. | | `waitForCooldownMs` | No | `0` | Целое число от 0 до 600000. Максимальное время ожидания самой ранней подходящей цели в cooldown перед возвратом `combo_unavailable`; отмена запроса отменяет ожидание. | | `defaultEffort` | No | `null` | `low`, `medium`, `high`, `xhigh`, `max` или `ultra`; применяется только когда вызывающая сторона не указала effort, а цель объявляет поддержку. | +| `reasoningEffortMode` | Нет | `"strict"` | `strict` или `adaptive`; задаёт пересечение возможностей и нормализацию параметров конкретной цели. | | `alias` | No | none | Необязательный обрезанный публичный id модели; используйте правила alias выше. Пустое значение хранится как отсутствие alias. | | `nativeAlias` | No | `false` | Явно разрешает поддерживаемому сейчас bare native alias перехватить приоритет routing/catalog только для неквалифицированного id. Bare `gpt-5.6-*` использует учётные данные Codex Pool/Direct; маршруты с квалификатором аккаунта сохраняют свою идентичность, а provider-qualified `openai-apikey/gpt-5.6-*` использует API-ключ и никогда не переходит на native alias. | | `displayName` | No | none | Метка только для отображения в catalog; обязательна при `nativeAlias: true`. | diff --git a/docs-site/src/content/docs/ru/reference/architecture.md b/docs-site/src/content/docs/ru/reference/architecture.md index c46da773c0..cd3f776efb 100644 --- a/docs-site/src/content/docs/ru/reference/architecture.md +++ b/docs-site/src/content/docs/ru/reference/architecture.md @@ -151,6 +151,15 @@ loopback; настроенные записи `corsAllowOrigins` расширя `426 upgrade_required`; Codex тогда откатывается на HTTP для этой сессии. Когда установлено `"websockets": true`, та же конечная точка принимает апгрейд и использует WebSocket-мост. +Для итоговой исходящей модели `gpt-5.3-codex-spark` каноническая пересылка в ChatGPT явно +отключает Responses Lite в HTTP-заголовке и нативных метаданных WS-кадра, в том числе при +выборе Spark через псевдоним — но только если в исходящем теле нет элемента `additional_tools` с непустым массивом `tools`. +Эта группа и ЕСТЬ Lite-форма доставки инструментов, поэтому тело Spark, которое её использует, +сохраняет Lite ВКЛЮЧЁННЫМ независимо от заголовка вызывающего клиента или конфигурации. Изменение идентичности Lite выводит старый сокет из использования; +последующие подходящие запросы с той же идентичностью могут повторно использовать новый сокет. +Другие модели и шлюзы сохраняют прежнюю политику Lite. Некорректные нативные метаданные +по-прежнему приводят к откату на HTTP без изменения тела запроса. + Compaction контекста Codex работает для маршрутизируемых моделей. `server/responses/compact.ts` обрабатывает `POST /v1/responses/compact`, выполняя внутренний маршрутизируемый ход суммаризации и возвращая сжатую историю, а `responses/parser.ts` и `bridge.ts` обрабатывают ходы diff --git a/docs-site/src/content/docs/ru/reference/configuration/routing.md b/docs-site/src/content/docs/ru/reference/configuration/routing.md index dbc7181044..595916aebd 100644 --- a/docs-site/src/content/docs/ru/reference/configuration/routing.md +++ b/docs-site/src/content/docs/ru/reference/configuration/routing.md @@ -87,7 +87,8 @@ selector-qualified строки и возвращает обычные GPT-ст | `targets` | `{ provider: string; model: string; weight?: number }[]` | required | Упорядоченные конкретные маршруты. `weight` находится в диапазоне 1–10000 и по умолчанию равен `1`. | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | Стратегия выбора. Порядок целей задаёт приоритет `failover`; значения `weight` определяют взвешивание выборов `round-robin` и `random`; `least-used` следует числу зарегистрированных успешных запросов; `reset-window` следует ближайшему сбросу квоты. | | `stickyLimit?` | `number` | `1` | Число успешных запросов, удерживаемых в одной партии round-robin. Диапазон 1–100. | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | Применяется, только если вызывающая сторона не задала effort, а выбранная цель объявляет эту ступень. | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | `defaultEffort` заполняет отсутствующий `reasoning.effort`, если задан `defaultEffort`, отличный от `null`, и список поддерживаемых уровней цели известен и непуст. Поддерживаемое настроенное значение сохраняется; иначе выбирается максимальный поддерживаемый уровень не выше него, а если такого нет — минимальный поддерживаемый уровень. При неизвестном или пустом списке default не добавляется. | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` вычисляет пересечение известных списков уровней, включая пустые; `"adaptive"` исключает пустые списки. Неизвестные списки не ограничивают пересечение в обоих режимах. При отправке явно пустой список удаляет параметры effort/thinking в обоих режимах, а неизвестный — только в adaptive. `reasoning.summary` сохраняется. Разрешение effort для известных непустых списков, выбор и порядок целей не меняются. | | `alias?` | `string` | — | Необязательный публичный id модели вместо канонического slug в селекторе. | | `nativeAlias?` | `boolean` | `false` | Даёт поддерживаемому bare native id приоритет только для этого неквалифицированного id. Bare `gpt-5.6-*` использует учётные данные Codex Pool/Direct. Маршруты с квалификатором аккаунта остаются отдельными. Провайдер-квалифицированные маршруты, например `openai-apikey/gpt-5.6-*`, используют настроенный API-ключ и никогда не переходят на native alias. | | `displayName?` | `string` | — | Метка только для catalog; для native alias обязательна и не может быть пустой. | diff --git a/docs-site/src/content/docs/ru/reference/configuration/server.md b/docs-site/src/content/docs/ru/reference/configuration/server.md index 84b3caf183..000a194b5c 100644 --- a/docs-site/src/content/docs/ru/reference/configuration/server.md +++ b/docs-site/src/content/docs/ru/reference/configuration/server.md @@ -14,6 +14,7 @@ description: Listener, удалённый доступ, admission key, тайм | `hostname?` | `string` | `"127.0.0.1"` | Адрес bind'а. Не-loopback bind требует `OPENCODEX_API_AUTH_TOKEN`. | | `proxy?` | `string` | — | URL исходящего HTTP(S)-прокси или `${ENV_VAR}`. Применяется к `HTTP_PROXY` / `HTTPS_PROXY` только когда эти переменные не заданы; loopback всегда остаётся в `NO_PROXY`. | | `emptyCompletionRetry?` | `boolean` | `false` | Явно включает один идентичный повтор Responses, если в turn нет ни текста, ни tool call, включая случай, когда stream завершается до terminal event. Повтор может тарифицироваться. `OCX_EMPTY_COMPLETION_RETRY=0` отключает его без изменения config; combo и routed-compaction turn исключены. | +| `dropCodexSafetyBuffering?` | `boolean` | `false` | Удаляет подсказки Codex safety-buffering из passthrough-ответов Codex Responses: заголовки `x-codex-safety-buffering-enabled` / `x-codex-safety-buffering-faster-model`, SSE-события `response.metadata` типа `safety_buffering` и поле `safety_buffering` в других SSE-событиях. Codex TUI отображает их как предложение повторить запрос с более быстрой моделью, действие по умолчанию в котором переключает сессию на более слабую модель. Остальные заголовки `x-codex-*` и содержимое других SSE-событий передаются без изменений, кроме удаления этого поля. По умолчанию выключено. | | `stallTimeoutSec?` | `number` | `300` | Секунды без upstream-данных до `response.incomplete`. Минимум 1. | | `connectTimeoutMs?` | `number` | `200000` | Дедлайн одной попытки DNS/TCP/TLS/final-header; он завершается до генерации тела ответа. | | `shutdownTimeoutMs?` | `number` | `5000` | Дедлайн graceful-drain до принудительного прерывания активных turn'ов. | @@ -219,3 +220,5 @@ opencodex. Перед использованием прогоните soak-test ## Сетевая диагностика квоты Codex Поле `quotaRefresh` в строке основного аккаунта Codex описывает получение квоты, а не её остаток или право доступа к модели. Оно может отсутствовать при чтении кэша или если запрос не выполнялся. Используется окружение работающего прокси-сервиса, а не текущего терминала. Если `proxy` не задан, существующее окружение сохраняется; `"auto"` читает только статические настройки прокси Windows при запуске. PAC/WPAD, настройки только SOCKS и изменения во время работы автоматически не учитываются. Успех через TUN сам по себе не подтверждает исправность пути HTTP-прокси. См. [команды и состояния на английском](/reference/configuration/server/#codex-quota-network-diagnostics). + +`dropCodexSafetyBuffering`: не меняет проверки безопасности провайдера или отказы. Native WebSocket `codex.response.metadata.headers` и `/responses/compact` не входят в область фильтра. diff --git a/docs-site/src/content/docs/tr/guides/combos.md b/docs-site/src/content/docs/tr/guides/combos.md index b8cd5bad0d..4ac9e0febf 100644 --- a/docs-site/src/content/docs/tr/guides/combos.md +++ b/docs-site/src/content/docs/tr/guides/combos.md @@ -249,22 +249,16 @@ veya politika retlerini gizlemez. ## Varsayılan akıl yürütme çabası -`defaultEffort`, yalnızca bunların tümü doğru olduğunda `reasoning.effort` -sağlar: +`defaultEffort`, combo varsayılanı null değilse ve hedefin desteklenen seviye listesi bilinen ve boş olmayan bir listeyse eksik `reasoning.effort` değerini doldurur. Yapılandırılmış değer destekleniyorsa korunur; değilse bu değeri aşmayan en yüksek desteklenen seviye, böyle bir seviye yoksa en düşük desteklenen seviye kullanılır. Liste bilinmiyor veya boşsa varsayılan eklenmez. -1. kombonun boş olmayan (non-null) bir varsayılanı vardır; -2. arayan bir çaba ayarlamamıştır; ve -3. seçilen hedefin kataloğu tam olarak bu çabayı bildirmektedir. +Varsayılan ekleme mevcut effort ve diğer reasoning alanlarını korur. Aşağıdaki yetenek normalizasyonu desteklenmeyen effort/thinking denetimlerini ayrıca kaldırabilir. Desteklenen varsayılanlar: `low`, `medium`, `high`, `xhigh`, `max`, `ultra`; alanı atlamak veya `null` kullanmak eklemeyi kapatır. -İstekte bir `reasoning` nesnesi yoksa opencodex bir tane oluşturur. Bir `effort` -özelliği olmadan `reasoning` varsa diğer alanları korur ve varsayılanı ekler. -Arayan tarafından sağlanan bir çabanın üzerine asla yazılmaz. -Hedef yeteneği bilinmediğinde veya yapılandırılan çabayı içermediğinde opencodex -varsayılanı atlar ve hedefin kendi davranışını değiştirmeden bırakır. -Desteklenen değerler `low`, `medium`, `high`, `xhigh`, `max` ve `ultra`'dır; -çabayı tamamen arayana ve hedefe bırakmak için alanı atlayın veya `null` olarak -ayarlayın. +## Farklı reasoning yetenekleri + +`reasoningEffortMode` varsayılan olarak `"strict"` kullanır: açıkça boş listeler dahil tüm hedeflerin effort listelerinin kesişimi yayımlanır. `"adaptive"`, karma kombolarda seçiciyi korumak için boş listeleri kesişimden çıkarır. Bilinmeyen listeler her iki modda da katalog kesişimini sınırlamaz. + +Gönderim sırasında açıkça boş liste her iki modda effort ve thinking denetimlerini kaldırır; bilinmeyen liste bunları yalnızca adaptive modunda kaldırır. `reasoning.summary` ve effort dışındaki alanlar korunur. Bilinen, boş olmayan hedeflerin effort çözümü değişmez. strict modundaki bilinmeyen hedefler ve normal native Chat bilinmeyen bildirimleri çağıranın denetimlerini korur. Varsayılan değer ekleme mevcut effort değerini değiştirmez; yetenek normalizasyonu desteklenmeyen denetimleri kaldırabilir. ## Şifrelenmiş v2 alt ajan görevleri @@ -373,6 +367,7 @@ saklanır: | `strategy` | Hayır | `"failover"` | İzin verilen değerler: `"failover"`, `"round-robin"`, `"random"`, `"least-used"`, `"reset-window"`. | | `stickyLimit` | Hayır | `1` | Yalnızca `round-robin` için geçerlidir; seçim başına 1 ile 100 arasında başarılı istek tam sayısı. | | `defaultEffort` | Hayır | `null` | `low`, `medium`, `high`, `xhigh`, `max` veya `ultra`; yalnızca arayan çabayı atladığında ve hedef desteği bildirdiğinde uygulanır. | +| `reasoningEffortMode` | Hayır | `"strict"` | `strict` veya `adaptive`; karma yetenek kesişimini ve hedefe özel normalizasyonu seçer. | | `alias` | Hayır | yok | İsteğe bağlı kırpılmış genel model kimliği; yukarıdaki takma ad kurallarını kullanın. Boş bir değer takma ad yok olarak saklanır. | | `nativeAlias` | Hayır | `false` | Şu anda desteklenen yalın bir yerel `alias`'ın yönlendirme ve katalog önceliği almasına açıkça izin verin. Asla takma addan çıkarılmaz. | | `displayName` | Hayır | yok | Sınırlı salt görüntüleme katalog etiketi. `nativeAlias` true olduğunda gerekli ve boş değildir. | diff --git a/docs-site/src/content/docs/tr/reference/architecture.md b/docs-site/src/content/docs/tr/reference/architecture.md index 24d2bbb7aa..5dd9b99957 100644 --- a/docs-site/src/content/docs/tr/reference/architecture.md +++ b/docs-site/src/content/docs/tr/reference/architecture.md @@ -170,6 +170,15 @@ opencodex `426 upgrade_required` döndürür; Codex daha sonra bu oturum için HTTP'ye geri döner. `"websockets": true` ayarlandığında aynı uç nokta yükseltmeyi kabul eder ve WebSocket köprüsünü kullanır. +Son gönderilen model `gpt-5.3-codex-spark` olduğunda, kanonik ChatGPT iletimi HTTP başlığında +ve yerel WS çerçevesi meta verilerinde Responses Lite'ı açıkça kapatır; Spark bir takma adla +seçildiğinde de bu geçerlidir — ancak yalnızca giden gövde boş olmayan `tools` dizisine sahip bir `additional_tools` grubu +taşımıyorsa. Bu grup Lite'ın araç teslim biçiminin kendisidir; onu kullanan bir Spark gövdesi, +çağıran veya yapılandırılmış başlık ne derse desin Lite'ı AÇIK tutar. Lite kimliği değişince eski soket kullanım dışı bırakılır; +aynı kimliğe sahip sonraki uygun istekler yeni soketi yeniden kullanabilir. Diğer modeller ve +ağ geçitleri mevcut Lite politikalarını korur. Bozuk yerel meta verilerde, istek gövdesi +değiştirilmeden HTTP'ye geri dönülmeye devam edilir. + Codex bağlam sıkıştırması yönlendirilen modeller için çalışır. `server/responses/compact.ts`, dahili bir yönlendirilen özetleme turu çalıştırarak ve sıkıştırılmış geçmişi döndürerek `POST /v1/responses/compact`'ı @@ -220,4 +229,3 @@ Dahili model `types.ts` içinde yer alır: `OcxParsedRequest`, `OcxContext`, `OcxProviderConfig`). İki yardımcı yaygın olarak kullanılır: `namespacedToolName()` ve `modelInList()` (`noVisionModels` / `noReasoningModels` için toleranslı `:size` etiketi eşleştirmesi). - diff --git a/docs-site/src/content/docs/tr/reference/configuration/routing.md b/docs-site/src/content/docs/tr/reference/configuration/routing.md index 74d239ee14..b1cbe4b484 100644 --- a/docs-site/src/content/docs/tr/reference/configuration/routing.md +++ b/docs-site/src/content/docs/tr/reference/configuration/routing.md @@ -110,7 +110,8 @@ aileleri kullanamaz. | `targets` | `{ provider: string; model: string; weight?: number }[]` | gerekli | Sıralı somut rotalar. `weight` 1–10000 arasındadır ve varsayılan olarak `1`'dir. | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | Seçim stratejisi. Hedef sırası `failover` önceliğini belirler; `weight` değerleri `round-robin` ve `random` seçimlerini biçimlendirir; `least-used` kaydedilen başarılı istekleri izler; `reset-window` en yakın kota sıfırlamasını izler. | | `stickyLimit?` | `number` | `1` | Tek bir round-robin grubunda tutulan başarılı istekler. Aralık 1–100. | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | ayarlanmamış | Yalnızca arayan çabayı atladığında ve seçilen hedef istenen basamağı bildirdiğinde uygulanır. | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | ayarlanmamış | `defaultEffort`, combo varsayılanı null değilse ve hedefin desteklenen seviye listesi bilinen ve boş olmayan bir listeyse eksik `reasoning.effort` değerini doldurur. Yapılandırılmış değer destekleniyorsa korunur; değilse bu değeri aşmayan en yüksek desteklenen seviye, böyle bir seviye yoksa en düşük desteklenen seviye kullanılır. Liste bilinmiyor veya boşsa varsayılan eklenmez. | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"`, boş listeler dahil bilinen hedef seviye listelerinin kesişimini alır; `"adaptive"` boş listeleri çıkarır. Bilinmeyen listeler iki modda da kesişimi sınırlamaz. Gönderimde açıkça boş listeler iki modda effort/thinking denetimlerini kaldırır; bilinmeyen listeler bunu yalnızca adaptive modunda yapar. `reasoning.summary` korunur. Bilinen boş olmayan hedeflerin effort çözümü ve hedef seçimi/sırası değişmez. | | `alias?` | `string` | — | Kurallı seçici slug'ı yerine isteğe bağlı genel model kimliği. | | `nativeAlias?` | `boolean` | `false` | Şu anda desteklenen bir yalın yerel kimliğin yalnızca o niteliksiz kimlik için öncelikli olmasına izin verin. Yalın `gpt-5.6-*` kimlikleri Codex Havuz/Direct kimlik bilgilerini kullanır. Hesap nitelikli rotalar ayrı kalır. `openai-apikey/gpt-5.6-*` gibi sağlayıcı nitelikli rotalar yapılandırılmış API anahtarı rotalarını kullanır ve asla yerel takma ada düşmez. | | `displayName?` | `string` | — | Yalnızca görüntüleme amaçlı katalog etiketi, yerel bir takma ad için gerekli ve boş olmamalıdır. | diff --git a/docs-site/src/content/docs/zh-cn/guides/combos.md b/docs-site/src/content/docs/zh-cn/guides/combos.md index abe32ae786..7c8c8efd63 100644 --- a/docs-site/src/content/docs/zh-cn/guides/combos.md +++ b/docs-site/src/content/docs/zh-cn/guides/combos.md @@ -165,15 +165,16 @@ combo 失败分为 **跳转** 失败和 **终止** 失败。 ## 默认推理力度 -只有在以下所有条件都满足时,`defaultEffort` 才会提供 `reasoning.effort`: +当 combo 配置了非 null 默认值且目标支持列表已知且非空时,`defaultEffort` 会填充省略的 `reasoning.effort`。目标支持配置值时保留该值,否则选择不高于配置值的最高支持档位;若不存在更低档位,则使用最低支持档位。未知或空列表不会注入默认值。 -1. combo 有一个非空默认值; -2. 调用方没有设置 effort;并且 -3. 选中的目标目录明确声明了该精确的 effort。 +默认值注入保留已有 effort 和其他 reasoning 字段。下述能力归一化可单独移除不支持的 effort/thinking 控制。默认值支持 `low`、`medium`、`high`、`xhigh`、`max`、`ultra`;省略字段或设为 `null` 可关闭注入。 -如果请求没有 `reasoning` 对象,opencodex 会创建一个。如果 `reasoning` 存在但没有 `effort` 属性,它会保留其他字段并添加默认值。调用方提供的 effort 永远不会被覆盖。 -当目标能力未知,或者不包含配置的 effort 时,opencodex 会省略默认值,并保持目标自身行为不变。支持的值是 `low`、`medium`、`high`、`xhigh`、`max` 和 `ultra`;省略该字段或将其设为 `null`,就会把 effort 完全交给调用方和目标。 +## 混合 reasoning 能力 + +`reasoningEffortMode` 默认为 `"strict"`,发布所有目标 effort 列表的交集,包括显式空列表。`"adaptive"` 在计算交集时排除空列表,让混合 combo 保留选择器。未知列表在两种模式下都不限制目录交集。 + +发送时,显式空列表在两种模式下都会移除 effort 和 thinking 控制;未知列表仅在 adaptive 下移除这些控制。`reasoning.summary` 和其他非 effort 字段保持不变,已知非空目标继续按现有规则解析 effort。strict 的未知目标及普通 native Chat 的未知声明保留调用方控制。默认值填充不会覆盖现有 effort,但能力归一化可移除不支持的控制。 ## 图片 / 多模态能力 @@ -266,6 +267,7 @@ combo 会存储在顶层的 `combos` 对象中,并以 combo id 作为键: | `cooldownMs` | 否 | 未设置 → 上游回退值(请求速率限制代码为 `1302`/`1305` 的 429 为 5 秒,否则为 60 秒) | 1 到 600000 的整数。设置后,只要没有可用的上游 `Retry-After` 或 Codex 重置信号,就会作为每个目标的冷却时间应用,包括请求速率限制 429;未设置时使用上游回退值。 | | `waitForCooldownMs` | 否 | `0` | 0 到 600000 的整数。在返回 `combo_unavailable` 前等待最早恢复资格的冷却中目标的最长时间;请求中止会取消等待。 | | `defaultEffort` | 否 | `null` | `low`、`medium`、`high`、`xhigh`、`max` 或 `ultra`;仅当调用方省略 effort 且目标声明支持时才会应用。 | +| `reasoningEffortMode` | 否 | `"strict"` | `strict` 或 `adaptive`;选择混合能力交集和目标级控制归一化。 | | `imageInput` | 否 | `"auto"` | `"auto"` 或 `"disabled"`。`"auto"` 仅在每个目标都支持图片时发布图片能力;`"disabled"` 强制仅文本(从对外能力中去掉图片,并在分发前拒绝带图请求)。 | | `alias` | 否 | 无 | 可选的、已修剪的公开模型 id;使用上面的别名规则。空值会以“无别名”形式存储。 | | `nativeAlias` | 否 | `false` | 显式允许当前受支持的裸原生 alias 接管路由和 catalog 优先级;绝不会根据 alias 自动推断。 | diff --git a/docs-site/src/content/docs/zh-cn/reference/architecture.md b/docs-site/src/content/docs/zh-cn/reference/architecture.md index 94c5eb8ea0..b1edbd1e8b 100644 --- a/docs-site/src/content/docs/zh-cn/reference/architecture.md +++ b/docs-site/src/content/docs/zh-cn/reference/architecture.md @@ -134,6 +134,13 @@ thread affinity 位于 `codex/` 下,不会出现在管理 API 响应中。请 session 中回退到 HTTP。设置 `"websockets": true` 后,同一 endpoint 会接受 upgrade 并使用 WebSocket bridge。 +当最终发送的模型为 `gpt-5.3-codex-spark` 时,canonical ChatGPT 转发会在 HTTP 请求头和 +原生 WS 帧元数据中明确关闭 Responses Lite,通过别名选择 Spark 时也一样;但这仅适用于发送正文 +不含带有非空 `tools` 数组的 `additional_tools` 分组的情况。该分组本身就是 Lite 的工具投递形态,因此仍使用它的 Spark +正文会保持 Lite 开启,无论调用方或配置的请求头如何。Lite 标识变化时, +旧 socket 会退出使用;后续标识相同且满足复用条件的请求可以复用新 socket。其他模型和网关 +保留原有 Lite 策略。原生元数据格式不合法时,仍会回退到 HTTP,并保持请求正文不变。 + Codex context compaction 同样适用于路由模型。`server/responses/compact.ts` 处理 `POST /v1/responses/compact`,运行一次内部路由 summarization turn 并返回压缩后的历史; `responses/parser.ts` 与 `bridge.ts` 则处理 remote compaction v2 的 `compaction_trigger` turn, diff --git a/docs-site/src/content/docs/zh-cn/reference/configuration/routing.md b/docs-site/src/content/docs/zh-cn/reference/configuration/routing.md index 609c9ab48f..fa7d04ffc5 100644 --- a/docs-site/src/content/docs/zh-cn/reference/configuration/routing.md +++ b/docs-site/src/content/docs/zh-cn/reference/configuration/routing.md @@ -73,7 +73,8 @@ Codex Auth 页面将此 picker 行为作为选择加入项。关闭它会隐藏 | `targets` | `{ provider: string; model: string; weight?: number }[]` | required | 有序的具体路由。`weight` 范围为 1–10000,默认值为 `1`。 | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | 选择策略。目标顺序表示 `failover` 优先级;`weight` 决定 `round-robin` 和 `random` 的抽取权重;`least-used` 根据记录的成功次数选择;`reset-window` 跟随最近的额度重置。 | | `stickyLimit?` | `number` | `1` | 在单个轮询批次中保留的成功请求数。范围 1–100。 | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | 仅在调用方省略 effort 且所选目标声明了请求的档位时应用。 | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | 当 combo 配置了非 null 默认值且目标支持列表已知且非空时,`defaultEffort` 会填充省略的 `reasoning.effort`。目标支持配置值时保留该值,否则选择不高于配置值的最高支持档位;若不存在更低档位,则使用最低支持档位。未知或空列表不会注入默认值。 | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` 对所有已知目标档位列表取交集,包括空列表;`"adaptive"` 排除空列表。未知列表在两种模式下都不限制目录交集。发送时,显式空列表在两种模式下都会移除 effort/thinking 控制;未知列表仅在 adaptive 下移除。`reasoning.summary` 保持不变。已知非空目标的 effort 解析、目标选择和顺序不变。 | | `imageInput?` | `"auto" \| "disabled"` | `"auto"` | `"auto"` 仅在每个目标都支持图片时发布图片能力;`"disabled"` 强制仅文本(从对外能力中去掉图片,并在分发前拒绝带图请求)。 | | `alias?` | `string` | — | 可选的公开 model id,用于替代规范化的选择器 slug。 | | `nativeAlias?` | `boolean` | `false` | 仅让当前受支持的裸原生 id 对该不带限定前缀的 id 优先;带账号或提供方限定的 OpenAI 路由仍是独立路由。 | diff --git a/docs-site/src/content/docs/zh-cn/reference/configuration/server.md b/docs-site/src/content/docs/zh-cn/reference/configuration/server.md index 4031f9ff19..8e8050de64 100644 --- a/docs-site/src/content/docs/zh-cn/reference/configuration/server.md +++ b/docs-site/src/content/docs/zh-cn/reference/configuration/server.md @@ -14,6 +14,7 @@ description: 监听、远程访问、准入密钥、超时、存储、侧车、 | `hostname?` | `string` | `"127.0.0.1"` | 绑定地址。非回环绑定需要 `OPENCODEX_API_AUTH_TOKEN`。 | | `proxy?` | `string` | — | 出站 HTTP(S) 代理 URL,或 `${ENV_VAR}`。仅当 `HTTP_PROXY` / `HTTPS_PROXY` 未设置时才会应用;回环地址始终保留在 `NO_PROXY` 中。 | | `emptyCompletionRetry?` | `boolean` | `false` | 显式启用:当 Responses turn 既无文本也无工具调用时,使用相同请求重试一次,包括流在终止事件之前结束的情况。重试可能产生费用。`OCX_EMPTY_COMPLETION_RETRY=0` 可在不修改配置的情况下禁用;combo 与 routed-compaction turn 不参与。 | +| `dropCodexSafetyBuffering?` | `boolean` | `false` | 从 Codex Responses 透传响应中移除 Codex safety-buffering 提示:`x-codex-safety-buffering-enabled` / `x-codex-safety-buffering-faster-model` 响应头、类型为 `safety_buffering` 的 `response.metadata` SSE 事件,以及其他 SSE 事件中的 `safety_buffering` 字段。Codex TUI 会将这些提示显示为“使用更快模型重试”的提示框,其默认操作会把会话切换到较弱的模型。其他 `x-codex-*` 响应头和其他所有 SSE 事件内容均保持不变,但会移除该字段。默认关闭。 | | `stallTimeoutSec?` | `number` | `300` | 在上游没有数据之前可等待的秒数,超过后返回 `response.incomplete`。最小值为 1。 | | `connectTimeoutMs?` | `number` | `200000` | 每次尝试的 DNS/TCP/TLS/最终响应头截止时间;它在正文生成之前结束。 | | `shutdownTimeoutMs?` | `number` | `5000` | 优雅停机截止时间,超过后会中止仍在进行中的请求。 | @@ -185,3 +186,5 @@ Anthropic OAuth 侧车会复用 opencodex 现有的 Claude Code OAuth 指纹。 ## Codex 额度网络诊断 主 Codex 账户行中的 `quotaRefresh` 描述额度查询结果,并不代表剩余额度或模型访问权限。读取缓存或未执行查询时,该字段可能省略。查询使用正在运行的代理服务的环境,而不是当前终端的环境。未设置 `proxy` 时保留现有环境;`"auto"` 只在启动时读取 Windows 静态代理设置,不自动处理 PAC/WPAD、仅 SOCKS 的设置或运行中的更改。TUN 测试成功并不能单独证明 HTTP 代理路径正常。命令和状态说明见[英文网络诊断章节](/reference/configuration/server/#codex-quota-network-diagnostics)。 + +`dropCodexSafetyBuffering`: 不会改变供应商安全策略或拒绝响应。原生 WebSocket `codex.response.metadata.headers` 和 `/responses/compact` 不在过滤范围内。 diff --git a/docs-site/src/content/docs/zh-tw/guides/combos.md b/docs-site/src/content/docs/zh-tw/guides/combos.md index ce3ad70a94..be566c77cd 100644 --- a/docs-site/src/content/docs/zh-tw/guides/combos.md +++ b/docs-site/src/content/docs/zh-tw/guides/combos.md @@ -177,15 +177,16 @@ Failover 是刻意受限的。它有助於目標特定的可用性、認證、 ## 預設推理 effort -`defaultEffort` 僅在以下全為真時提供 `reasoning.effort`: +當 combo 設定非 null 預設值且目標支援清單已知且非空時,`defaultEffort` 會補入省略的 `reasoning.effort`。目標支援設定值時保留該值,否則選擇不高於設定值的最高支援層級;若沒有更低層級,則使用最低支援層級。未知或空清單不會注入預設值。 -1. combo 有非 null 預設值; -2. 呼叫者未設定 effort;且 -3. 所選目標的目錄宣告該精確 effort。 +預設值補入會保留既有 effort 與其他 reasoning 欄位。下述能力正規化可另外移除不支援的 effort/thinking 控制。預設值支援 `low`、`medium`、`high`、`xhigh`、`max`、`ultra`;省略欄位或設為 `null` 可關閉注入。 -若請求沒有 `reasoning` 物件,opencodex 建立一個。若 `reasoning` 存在但無 `effort` 屬性,它保留其他欄位並加入預設值。呼叫者提供的 effort 永不被覆寫。 -當目標能力未知或不包含設定的 effort 時,opencodex 省略預設值並保持目標自身行為不變。支援的值為 `low`、`medium`、`high`、`xhigh`、`max` 與 `ultra`;省略欄位或設為 `null` 可將 effort 完全交給呼叫者與目標。 +## 混合 reasoning 能力 + +`reasoningEffortMode` 預設為 `"strict"`,發布所有目標 effort 清單的交集,包括明確空清單。`"adaptive"` 計算交集時排除空清單,讓混合 combo 保留選擇器。未知清單在兩種模式下都不限制目錄交集。 + +傳送時,明確空清單在兩種模式下都會移除 effort 與 thinking 控制;未知清單只在 adaptive 移除這些控制。`reasoning.summary` 與其他非 effort 欄位保持不變,已知非空目標仍按現有規則解析 effort。strict 的未知目標及一般 native Chat 的未知宣告保留呼叫者控制。預設值補入不會覆寫現有 effort,但能力正規化可移除不支援的控制。 ## 加密的 v2 子代理任務 @@ -269,6 +270,7 @@ Combo 儲存於頂層 `combos` 物件中,以 combo id 為 key: | `strategy` | 否 | `"failover"` | 可用值為 `"failover"`、`"round-robin"`、`"random"`、`"least-used"`、`"reset-window"`。 | | `stickyLimit` | 否 | `1` | 僅適用於 `round-robin`:每次選擇的成功請求數,1 到 100 的整數。 | | `defaultEffort` | 否 | `null` | `low`、`medium`、`high`、`xhigh`、`max` 或 `ultra`;僅在呼叫者省略 effort 且目標宣告支援時套用。 | +| `reasoningEffortMode` | 否 | `"strict"` | `strict` 或 `adaptive`;選擇混合能力交集及目標層級控制正規化。 | | `alias` | 否 | 無 | 可選的修剪後公開模型 id;使用上述別名規則。空值儲存為無別名。 | ## 疑難排解 diff --git a/docs-site/src/content/docs/zh-tw/reference/architecture.md b/docs-site/src/content/docs/zh-tw/reference/architecture.md index daf9bc9080..be0b2c5003 100644 --- a/docs-site/src/content/docs/zh-tw/reference/architecture.md +++ b/docs-site/src/content/docs/zh-tw/reference/architecture.md @@ -134,6 +134,14 @@ thread affinity 位於 `codex/` 下,不會出現在管理 API 回應中。請 session 中回退到 HTTP。設定 `"websockets": true` 後,同一 endpoint 會接受 upgrade 並使用 WebSocket bridge。 +當最終傳送的模型為 `gpt-5.3-codex-spark` 時,canonical ChatGPT 轉送會在 HTTP 請求標頭與 +原生 WS 訊框中繼資料中明確關閉 Responses Lite,透過別名選擇 Spark 時也一樣;但僅限於傳送本文 +不含帶有非空 `tools` 陣列的 `additional_tools` 群組的情況。該群組本身就是 Lite 的工具傳遞形態,因此仍使用它的 Spark +本文會保持 Lite 開啟,無論呼叫端或設定的標頭為何。Lite 識別值 +改變時,舊 socket 會停止使用;後續識別值相同且符合重用條件的請求可以重用新 socket。 +其他模型與閘道保留既有 Lite 政策。原生中繼資料格式不合法時,仍會退回 HTTP,並保持 +請求本文不變。 + Codex context compaction 同樣適用於路由模型。`server/responses/compact.ts` 處理 `POST /v1/responses/compact`,執行一次內部路由 summarization turn 並回傳壓縮後的歷史; `responses/parser.ts` 與 `bridge.ts` 則處理 remote compaction v2 的 `compaction_trigger` turn, diff --git a/docs-site/src/content/docs/zh-tw/reference/configuration/routing.md b/docs-site/src/content/docs/zh-tw/reference/configuration/routing.md index 13a4e2622e..2001cf14a9 100644 --- a/docs-site/src/content/docs/zh-tw/reference/configuration/routing.md +++ b/docs-site/src/content/docs/zh-tw/reference/configuration/routing.md @@ -56,7 +56,8 @@ Codex Auth 頁面將此 picker 行為作為選擇加入功能暴露。停用它 | `targets` | `{ provider: string; model: string; weight?: number }[]` | 必填 | 有序的具體路由。`weight` 為 1–10000,預設 `1`。 | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | 選擇策略。目標順序為 `failover` 優先序;`weight` 塑造 `round-robin` 與 `random` 抽選;`least-used` 依循已記錄的成功次數;`reset-window` 依循最早的配額重設。 | | `stickyLimit?` | `number` | `1` | 在一個 round-robin 批次中保留的成功請求數。範圍 1–100。 | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | 未設定 | 僅在呼叫者省略 effort 且所選目標廣告請求的階層時套用。 | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | 未設定 | 當 combo 設定非 null 預設值且目標支援清單已知且非空時,`defaultEffort` 會補入省略的 `reasoning.effort`。目標支援設定值時保留該值,否則選擇不高於設定值的最高支援層級;若沒有更低層級,則使用最低支援層級。未知或空清單不會注入預設值。 | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` 對所有已知目標層級清單取交集,包括空清單;`"adaptive"` 排除空清單。未知清單在兩種模式下都不限制目錄交集。傳送時,明確空清單在兩種模式下都會移除 effort/thinking 控制;未知清單只在 adaptive 移除。`reasoning.summary` 保持不變。已知非空目標的 effort 解析、目標選擇及順序不變。 | | `alias?` | `string` | — | 可選的公開模型 id,取代標準 picker slug。 | | `nativeAlias?` | `boolean` | `false` | 讓目前支援的裸原生 id 僅對該未限定 id 取得優先。裸 `gpt-5.6-*` id 使用 Codex 池/Direct 憑證。帳號限定路由保持獨立。供應商限定路由(如 `openai-apikey/gpt-5.6-*`)使用其設定的 API-key 路由,且永不會落到原生別名。 | | `displayName?` | `string` | — | 僅顯示的目錄標籤,對原生別名為必填且非空。 | diff --git a/gui/src/i18n/de.ts b/gui/src/i18n/de.ts index db6b8c889f..07f9c5ad05 100644 --- a/gui/src/i18n/de.ts +++ b/gui/src/i18n/de.ts @@ -796,6 +796,7 @@ export const de: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth ratenbegrenzt (429)", "logs.detail.attempt.recovery.image413": "Bildnutzlast zu groß (413)", "logs.detail.attempt.recovery.emptyCompletion": "Wiederholung nach leerer Antwort", + "logs.detail.attempt.recovery.consoleGoUpload": "Console-Upload erneut versucht", "logs.detail.attempt.recovery.unknown": "Unbekannter Wiederherstellungsgrund", "logs.detail.reason.usage_missing": "Nutzung wurde nicht gemeldet.", "logs.detail.reason.usage_unsupported": "Dieser Anbieter meldet keine Nutzung.", diff --git a/gui/src/i18n/en.ts b/gui/src/i18n/en.ts index 74e4ded4d1..d19275be73 100644 --- a/gui/src/i18n/en.ts +++ b/gui/src/i18n/en.ts @@ -845,6 +845,7 @@ export const en = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth rate-limited (429)", "logs.detail.attempt.recovery.image413": "Image payload too large (413)", "logs.detail.attempt.recovery.emptyCompletion": "Empty completion retry", + "logs.detail.attempt.recovery.consoleGoUpload": "Console upload retry", "logs.detail.attempt.recovery.unknown": "Unknown recovery reason", "logs.detail.reason.usage_missing": "Usage was not reported.", "logs.detail.reason.usage_unsupported": "This provider does not report usage.", diff --git a/gui/src/i18n/fr.ts b/gui/src/i18n/fr.ts index 2bc3d5ab51..7b245b524f 100644 --- a/gui/src/i18n/fr.ts +++ b/gui/src/i18n/fr.ts @@ -821,6 +821,7 @@ export const fr: Record = { "logs.detail.attempt.recovery.transient5xx": "Erreur 5xx temporaire", "logs.detail.attempt.recovery.connectionReset": "Réinitialisation de la connexion", "logs.detail.attempt.recovery.emptyCompletion": "Nouvelle tentative après une réponse vide", + "logs.detail.attempt.recovery.consoleGoUpload": "Nouvelle tentative d’envoi Console", "logs.detail.attempt.recovery.oauth401": "Réauthentification OAuth", "logs.detail.attempt.recovery.key429": "Clé soumise à une limitation de débit (429)", "logs.detail.attempt.recovery.rateLimit429": "Limitation de débit (429)", diff --git a/gui/src/i18n/ja.ts b/gui/src/i18n/ja.ts index a787ab735b..551fbe3e69 100644 --- a/gui/src/i18n/ja.ts +++ b/gui/src/i18n/ja.ts @@ -758,6 +758,7 @@ export const ja: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth レート制限 (429)", "logs.detail.attempt.recovery.image413": "画像ペイロードが大きすぎます (413)", "logs.detail.attempt.recovery.emptyCompletion": "空の完了を再試行", + "logs.detail.attempt.recovery.consoleGoUpload": "Console アップロード再試行", "logs.detail.attempt.recovery.unknown": "不明なリカバリ理由", "logs.detail.reason.usage_missing": "使用量が報告されませんでした。", "logs.detail.reason.usage_unsupported": "このプロバイダーは使用量を報告しません。", diff --git a/gui/src/i18n/ko.ts b/gui/src/i18n/ko.ts index ccd2b06735..883ab327a1 100644 --- a/gui/src/i18n/ko.ts +++ b/gui/src/i18n/ko.ts @@ -827,6 +827,7 @@ export const ko: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth 요청 한도 초과 (429)", "logs.detail.attempt.recovery.image413": "이미지 페이로드가 너무 큼 (413)", "logs.detail.attempt.recovery.emptyCompletion": "빈 응답 재시도", + "logs.detail.attempt.recovery.consoleGoUpload": "Console 업로드 재시도", "logs.detail.attempt.recovery.unknown": "알 수 없는 복구 사유", "logs.detail.reason.usage_missing": "usage가 보고되지 않았습니다.", "logs.detail.reason.usage_unsupported": "이 프로바이더는 usage 보고를 지원하지 않습니다.", diff --git a/gui/src/i18n/ru.ts b/gui/src/i18n/ru.ts index ceb0c18a92..5d0eb4f527 100644 --- a/gui/src/i18n/ru.ts +++ b/gui/src/i18n/ru.ts @@ -813,6 +813,7 @@ export const ru: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth ограничен (429)", "logs.detail.attempt.recovery.image413": "Слишком большой размер изображения (413)", "logs.detail.attempt.recovery.emptyCompletion": "Повтор пустого завершения", + "logs.detail.attempt.recovery.consoleGoUpload": "Повтор загрузки Console", "logs.detail.attempt.recovery.unknown": "Неизвестная причина восстановления", "logs.detail.reason.usage_missing": "Данные об использовании не были сообщены.", "logs.detail.reason.usage_unsupported": "Этот провайдер не сообщает данные об использовании.", diff --git a/gui/src/i18n/tr.ts b/gui/src/i18n/tr.ts index 26a8f93d09..1d404f7ce2 100644 --- a/gui/src/i18n/tr.ts +++ b/gui/src/i18n/tr.ts @@ -832,6 +832,7 @@ export const tr: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth kısıtlandı (429)", "logs.detail.attempt.recovery.image413": "Görsel boyutu çok büyük (413)", "logs.detail.attempt.recovery.emptyCompletion": "Boş tamamlama yeniden denemesi", + "logs.detail.attempt.recovery.consoleGoUpload": "Console yüklemesi yeniden denendi", "logs.detail.attempt.recovery.unknown": "Bilinmeyen kurtarma nedeni", "logs.detail.reason.usage_missing": "Kullanım bildirilmedi.", "logs.detail.reason.usage_unsupported": "Bu sağlayıcı kullanım bildirmeyebilir.", diff --git a/gui/src/i18n/zh-TW.ts b/gui/src/i18n/zh-TW.ts index 1fb8387f4d..943a7224a2 100644 --- a/gui/src/i18n/zh-TW.ts +++ b/gui/src/i18n/zh-TW.ts @@ -2210,6 +2210,7 @@ export const zhTW: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth 被限流 (429)", "logs.detail.attempt.recovery.image413": "圖片承載過大 (413)", "logs.detail.attempt.recovery.emptyCompletion": "空白完成重試", + "logs.detail.attempt.recovery.consoleGoUpload": "Console 上傳重試", "logs.detail.attempt.recovery.unknown": "未知的復原原因", "logs.detail.estimate.provider_cost_overlay": "已使用供應商設定的價格覆蓋。", "logs.detail.estimate.priority_lower_bound": "無法取得已確認的 Priority 價格;目前顯示的估算是已知下限。", diff --git a/gui/src/i18n/zh.ts b/gui/src/i18n/zh.ts index 4e2cd831cc..fdd0b574ba 100644 --- a/gui/src/i18n/zh.ts +++ b/gui/src/i18n/zh.ts @@ -808,6 +808,7 @@ export const zh: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth 被限流 (429)", "logs.detail.attempt.recovery.image413": "图片载荷过大 (413)", "logs.detail.attempt.recovery.emptyCompletion": "空完成重试", + "logs.detail.attempt.recovery.consoleGoUpload": "Console 上传重试", "logs.detail.attempt.recovery.unknown": "未知的恢复原因", "logs.detail.reason.usage_missing": "未上报 usage。", "logs.detail.reason.usage_unsupported": "该提供方不支持上报 usage。", diff --git a/gui/src/pages/Logs.tsx b/gui/src/pages/Logs.tsx index 774efc455a..d7fc3ab5c4 100644 --- a/gui/src/pages/Logs.tsx +++ b/gui/src/pages/Logs.tsx @@ -113,7 +113,8 @@ type AttemptRecoveryKind = | "rate-limit-429" | "anthropic-oauth-429" | "image-413" - | "empty-completion"; + | "empty-completion" + | "console-go-upload-retry"; interface LogAttempt { ordinal: number; @@ -307,6 +308,7 @@ const RECOVERY_KIND_KEYS = { "anthropic-oauth-429": "logs.detail.attempt.recovery.anthropicOauth429", "image-413": "logs.detail.attempt.recovery.image413", "empty-completion": "logs.detail.attempt.recovery.emptyCompletion", + "console-go-upload-retry": "logs.detail.attempt.recovery.consoleGoUpload", } as const satisfies Record; /** Map a metric-unavailable reason to its i18n key. */ diff --git a/scripts/test-layout/layout.json b/scripts/test-layout/layout.json index 4f4ec44705..0e9acce084 100644 --- a/scripts/test-layout/layout.json +++ b/scripts/test-layout/layout.json @@ -1090,6 +1090,7 @@ "responses-account-label.test.ts": "responses", "responses-compaction-routing.test.ts": "responses", "responses-compaction.test.ts": "responses", + "responses-console-go-upload-retry.test.ts": "responses", "responses-context-overflow.test.ts": "responses", "responses-custom-tool-guidance.test.ts": "responses", "responses-custom-tool-repair.test.ts": "responses", @@ -1113,10 +1114,10 @@ "responses-pool-401-refresh.test.ts": "responses", "responses-pool-refresh-attribution.test.ts": "responses", "responses-reasoning-summary-passthrough.test.ts": "responses", - "responses-reasoning-summary-rewrite.test.ts": "responses", "responses-routed-web-search-fields.test.ts": "responses", "responses-self-named-namespace-scrub.test.ts": "responses", "responses-shadow-intercept.test.ts": "responses", + "responses-show-thinking-summary.test.ts": "responses", "responses-snapshot-repair-server.test.ts": "responses", "responses-snapshot-repair.test.ts": "responses", "responses-state-write-amplification.test.ts": "responses", diff --git a/src/adapters/google-wire-compiler.ts b/src/adapters/google-wire-compiler.ts index 88c482ba7d..aa835e50b4 100644 --- a/src/adapters/google-wire-compiler.ts +++ b/src/adapters/google-wire-compiler.ts @@ -130,12 +130,20 @@ function compileGenerationConfig(value: unknown): JsonObject | undefined { ))].slice(0, 5); if (stopSequences.length > 0) out.stopSequences = stopSequences; } - if (isObject(value.thinkingConfig) && typeof value.thinkingConfig.thinkingLevel === "string") { - const raw = value.thinkingConfig.thinkingLevel.toLowerCase(); - const thinkingLevel = GOOGLE_THINKING_LEVELS.has(raw) - ? raw - : (["xhigh", "max", "ultra"].includes(raw) ? "high" : undefined); - if (thinkingLevel) out.thinkingConfig = { thinkingLevel }; + if (isObject(value.thinkingConfig)) { + const thinking: JsonObject = {}; + if (typeof value.thinkingConfig.thinkingLevel === "string") { + const raw = value.thinkingConfig.thinkingLevel.toLowerCase(); + const thinkingLevel = GOOGLE_THINKING_LEVELS.has(raw) + ? raw + : (["xhigh", "max", "ultra"].includes(raw) ? "high" : undefined); + if (thinkingLevel) thinking.thinkingLevel = thinkingLevel; + } + // The one key that makes Google return `thought: true` text. Cloud Code Assist serves + // thinking either way (thoughtsTokenCount stays non-zero) but withholds the text unless the + // request opts in, so dropping it here silently reinstates the missing-thinking behavior. + if (value.thinkingConfig.includeThoughts === true) thinking.includeThoughts = true; + if (Object.keys(thinking).length > 0) out.thinkingConfig = thinking; } if (Array.isArray(value.responseModalities)) { const valid = value.responseModalities.filter((m): m is string => typeof m === "string" && ["TEXT", "IMAGE", "AUDIO"].includes(m)); diff --git a/src/adapters/google.ts b/src/adapters/google.ts index 7fcc88ba59..518aca3903 100644 --- a/src/adapters/google.ts +++ b/src/adapters/google.ts @@ -568,13 +568,15 @@ function googleToolCallMetadataFromPart( * Keep that provider visibility bit authoritative here so the streaming and buffered parsers * cannot accidentally expose the same hidden reasoning through different event types. */ -function googlePartTextEvent(part: GoogleResponsePart): AdapterEvent | undefined { +function googlePartTextEvent(part: GoogleResponsePart, thoughtSummary = false): AdapterEvent | undefined { // A malformed scalar/object is not text and must not cross the AdapterEvent boundary. Dropping // only this optional field preserves the rest of the part without inventing assistant output by // coercion; an empty string keeps its existing no-event behavior. if (typeof part.text !== "string" || part.text.length === 0) return undefined; return part.thought === true - ? { type: "reasoning_raw_delta", text: part.text } + ? thoughtSummary + ? { type: "thinking_delta", thinking: part.text } + : { type: "reasoning_raw_delta", text: part.text } : { type: "text_delta", text: part.text }; } @@ -721,6 +723,7 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte // Per-request closure: resolveAdapter builds a fresh adapter per request (server.ts), so buildRequest // can stash the CCA model/session for parseStream's reasoning-replay observation. let antigravityModel: string | undefined; + let returnsThoughtSummaries = false; let antigravitySession: string | undefined; // Vertex returns the same opaque Gemini thought signatures as CCA, but its replay namespace // must stay transport-scoped: a signature minted by one Google backend must never be sent to @@ -795,6 +798,8 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte : provider.googleMode === "vertex" ? parsed.modelId : resolveDirectGeminiWireModelId(parsed.modelId, provider.directGeminiWireRenames !== false); + returnsThoughtSummaries = provider.googleMode === "cloud-code-assist" + && /^gemini-/.test(routedModelId) && !isImageCapableModel(parsed.modelId); // AI Studio's `-tiered` spelling is wire-only; CCA aliases may migrate to another generation. const identityModelId = provider.googleMode === "cloud-code-assist" ? routedModelId : parsed.modelId; const stripRejectedClaudeSdkParagraph = provider.googleMode === "cloud-code-assist" @@ -866,11 +871,20 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte ); antigravityModel = wireModelId; antigravitySession = sessionId; + // CCA Gemini exposes provider-authored thought summaries with includeThoughts. + // Other CCA model families do not share this request contract. + const includeThoughts = provider.showThinkingSummary === true + && parsed.options.hideThinkingSummary !== true + && /^gemini-/.test(wireModelId) + && !isImageCapableModel(parsed.modelId); // Effort → thinkingConfig for CCA (CLIProxyAPI proven: request.generationConfig.thinkingConfig). // Suffix/compat IDs return thinkingLevel=undefined — the suffix IS the effort, no contradiction. - if (thinkingLevel) { + if (thinkingLevel || includeThoughts) { const gc = (body.generationConfig ?? {}) as Record; - gc.thinkingConfig = { thinkingLevel }; + gc.thinkingConfig = { + ...(thinkingLevel ? { thinkingLevel } : {}), + ...(includeThoughts ? { includeThoughts: true } : {}), + }; body.generationConfig = gc; } // Reasoning continuity: Gemini models re-inject cached thoughtSignatures; Claude-on-Antigravity @@ -1139,7 +1153,7 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte if (part.thought === true && sig && isLikelyRealThoughtSignature(sig)) { pendingStreamThoughtSig = sig; } - const textEvent = googlePartTextEvent(part); + const textEvent = googlePartTextEvent(part, returnsThoughtSummaries); if (textEvent) { emittedContentEvent = true; yield textEvent; @@ -1415,7 +1429,7 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte if (part.thought === true && sig && isLikelyRealThoughtSignature(sig)) { pendingThoughtSig = sig; } - const textEvent = googlePartTextEvent(part); + const textEvent = googlePartTextEvent(part, returnsThoughtSummaries); if (textEvent) events.push(textEvent); const inline = (part as { inlineData?: { mimeType?: string; data?: string } }).inlineData; if (inline && typeof inline.data === "string") { diff --git a/src/adapters/openai-chat.ts b/src/adapters/openai-chat.ts index e338d845ad..e4136ecb3e 100644 --- a/src/adapters/openai-chat.ts +++ b/src/adapters/openai-chat.ts @@ -130,6 +130,10 @@ export function buildOpenAIChatPassthroughRequest( for (const field of CHAT_PASSTHROUGH_FIELDS) { if (rawBody[field] !== undefined) body[field] = rawBody[field]; } + const rawEfforts = modelRecordValue(provider.modelReasoningEfforts, modelId) ?? provider.reasoningEfforts; + if (modelInList(provider.noReasoningModels, modelId) || rawEfforts?.length === 0) { + delete body.reasoning_effort; + } const openRouterRouting = resolveOpenRouterRouting(provider, modelId); if (openRouterRouting) body.provider = openRouterProviderPayload(openRouterRouting); diff --git a/src/adapters/openai-responses.ts b/src/adapters/openai-responses.ts index 92de305f6d..d4d5f85e94 100644 --- a/src/adapters/openai-responses.ts +++ b/src/adapters/openai-responses.ts @@ -866,6 +866,21 @@ function promoteClientLoadedTools(body: unknown): unknown { } const MAX_RESPONSES_CALL_ID_LENGTH = 64; + +/** + * Whether the outgoing body still delivers tools through the responses-lite shape. + * + * Lite carries the client catalog as an `additional_tools` input item; the non-Lite wire shape + * expects top-level `tools`. Anything that flips the Lite advertisement has to agree with the + * shape actually being sent, or the destination silently loses the tool surface. + */ +function bodyCarriesLiteToolShape(body: Record): boolean { + if (!Array.isArray(body.input)) return false; + return body.input.some(item => + isPlainObject(item) && item.type === "additional_tools" + && Array.isArray(item.tools) && item.tools.length > 0 + ); +} const REPAIRED_CALL_ID_PREFIX = "call_ocx_"; const REPAIRED_CALL_ID_DIGEST_LENGTH = MAX_RESPONSES_CALL_ID_LENGTH - REPAIRED_CALL_ID_PREFIX.length; @@ -2516,7 +2531,7 @@ export function createResponsesPassthroughAdapter(provider: OcxProviderConfig): ), isXaiSchemaTarget(provider), ); - const finalBody = stripDisabledVerbosity( + const unnormalizedBody = stripDisabledVerbosity( stripDisabledReasoningSummaries( normalizeConfiguredReasoningSummaryDelivery(sanitizedBody, provider, parsed.modelId), provider, @@ -2525,13 +2540,32 @@ export function createResponsesPassthroughAdapter(provider: OcxProviderConfig): provider, parsed.modelId, ); + // Normalize the wire model before deriving model-dependent transport metadata. + const finalBody = + provider.modelSuffixBracketStrip + && unnormalizedBody !== null + && typeof unnormalizedBody === "object" + && !Array.isArray(unnormalizedBody) + && typeof (unnormalizedBody as { model?: unknown }).model === "string" + ? { ...(unnormalizedBody as Record), model: stripBracketedModelSuffix((unnormalizedBody as { model: string }).model) } + : unnormalizedBody; if (isCanonicalOpenAiForwardProvider(provider)) { - // Spark closes Responses Lite streams before a terminal completion. Select compatibility - // from the final wire model so aliases cannot leave the caller or a static header enabled. + // Select Spark's Lite compatibility from the final wire model, including aliases, and + // let the BODY decide it. The header also overrides native WS metadata downstream, so a + // forwarded or statically configured value must never contradict the shape being sent. + // + // The synchronized catalog keeps `use_responses_lite: true` for Spark precisely because + // it selects tool delivery (`input[].additional_tools` instead of top-level `tools`), and + // stripSparkCompatibility filters that group in place rather than promoting it. So a + // Lite-shaped body is pinned back ON — otherwise an inherited `false` advertises non-Lite + // while the tools exist only in the Lite shape, and Spark loses the tool surface. Only a + // body with no Lite tool group is downgraded, which is what the stream fix needs. if (isPlainObject(finalBody) && finalBody.model === "gpt-5.3-codex-spark") { + const liteShaped = bodyCarriesLiteToolShape(finalBody); for (const name of Object.keys(headers)) { if (name.toLowerCase() === CODEX_RESPONSES_LITE_HEADER) delete headers[name]; } + headers[CODEX_RESPONSES_LITE_HEADER] = liteShaped ? "true" : "false"; } const routingHeaders = new Headers(headers); applyCodexRoutingHint(routingHeaders, finalBody); @@ -2558,15 +2592,7 @@ export function createResponsesPassthroughAdapter(provider: OcxProviderConfig): // here, on the serialized body, not on the parsed selector. One place covers both the // HTTP and the WebSocket outbound, because the WS path transports this same request // instead of rebuilding it. - const body = JSON.stringify( - provider.modelSuffixBracketStrip - && finalBody !== null - && typeof finalBody === "object" - && !Array.isArray(finalBody) - && typeof (finalBody as { model?: unknown }).model === "string" - ? { ...(finalBody as Record), model: stripBracketedModelSuffix((finalBody as { model: string }).model) } - : finalBody, - ); + const body = JSON.stringify(finalBody); const releaseBodyObservation = translatorBudget.observeExternallyCapped( "passthrough_serialization", new TextEncoder().encode(body).byteLength, diff --git a/src/bridge.ts b/src/bridge.ts index 3f5f529d3f..612b60ba36 100644 --- a/src/bridge.ts +++ b/src/bridge.ts @@ -663,16 +663,13 @@ export function bridgeToResponsesSSE( const closeCurrentRawReasoning = () => { if (!currentRawReasoning) return; rawReasoningForNextToolCall = currentRawReasoning.text; - emit("response.reasoning_summary_text.done", { - item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, summary_index: 0, text: currentRawReasoning.text, - }); - emit("response.reasoning_summary_part.done", { - item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, summary_index: 0, - part: { type: "summary_text", text: currentRawReasoning.text }, + emit("response.reasoning_text.done", { + item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, content_index: 0, text: currentRawReasoning.text, }); const item = { type: "reasoning", id: currentRawReasoning.itemId, - summary: [{ type: "summary_text", text: currentRawReasoning.text }], + summary: [] as never[], + content: [{ type: "reasoning_text", text: currentRawReasoning.text }], }; emit("response.output_item.done", { output_index: currentRawReasoning.outputIndex, item }); retainFinishedItem(item as OutputItem, currentRawReasoning.textBytes, "reasoning"); @@ -1111,10 +1108,6 @@ export function bridgeToResponsesSSE( const itemId = `rs_${uuid()}`; const item = { type: "reasoning", id: itemId, summary: [] as { type: string; text: string }[] }; emit("response.output_item.added", { output_index: outputIndex, item }); - emit("response.reasoning_summary_part.added", { - item_id: itemId, output_index: outputIndex, summary_index: 0, - part: { type: "summary_text", text: "" }, - }); currentRawReasoning = { itemId, outputIndex, text: "", textBytes: 0 }; } ({ value: currentRawReasoning.text, bytes: currentRawReasoning.textBytes } = appendString( @@ -1123,9 +1116,12 @@ export function bridgeToResponsesSSE( event.text, "reasoning", )); - emit("response.reasoning_summary_text.delta", { + // Raw reasoning (openai-chat reasoning_content, kiro tags) rides the CONTENT + // channel. Clients control raw-reasoning display; this text is not a + // provider-authored summary. + emit("response.reasoning_text.delta", { item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, - summary_index: 0, delta: event.text, + content_index: 0, delta: event.text, }); break; } @@ -1783,7 +1779,8 @@ function buildResponseJSONWithBudget( } pushOutput({ type: "reasoning", id: `rs_${uuid()}`, - summary: [{ type: "summary_text", text: currentRawReasoning }], + summary: [], + content: [{ type: "reasoning_text", text: currentRawReasoning }], }, currentRawReasoningBytes, "reasoning"); currentRawReasoning = ""; currentRawReasoningBytes = 0; diff --git a/src/combos/request.ts b/src/combos/request.ts index abafccc525..63c5ba7fca 100644 --- a/src/combos/request.ts +++ b/src/combos/request.ts @@ -1,4 +1,4 @@ -import type { OcxComboDefaultEffort, OcxComboTarget, OcxConfig } from "../types"; +import type { OcxComboDefaultEffort, OcxComboReasoningEffortMode, OcxComboTarget, OcxConfig } from "../types"; import { resolveEffortAtOrBelow } from "../reasoning-effort"; import { resolveComboId } from "./types"; @@ -59,9 +59,14 @@ export function concreteComboRequestBody( target: Pick, defaultEffort: OcxComboDefaultEffort | null, targetReasoningEfforts: readonly string[] | undefined, + reasoningEffortMode: OcxComboReasoningEffortMode = "strict", ): Record { const clone = structuredClone(body) as Record; clone.model = `${target.provider}/${target.model}`; + if (targetReasoningEfforts?.length === 0 + || (reasoningEffortMode === "adaptive" && targetReasoningEfforts === undefined)) { + stripUnsupportedReasoningControls(clone); + } if (!defaultEffort) return clone; const reasoning = clone.reasoning; const needsDefault = reasoning === undefined || ( @@ -104,3 +109,16 @@ export function concreteComboRequestBody( } return clone; } + +function stripUnsupportedReasoningControls(body: Record): void { + const reasoning = body.reasoning; + if (reasoning && typeof reasoning === "object" && !Array.isArray(reasoning)) { + const next = { ...(reasoning as Record) }; + delete next.effort; + if (Object.keys(next).length > 0) body.reasoning = next; + else delete body.reasoning; + } + delete body.reasoning_effort; + delete body.thinking_budget; + delete body.thinking; +} diff --git a/src/config.ts b/src/config.ts index 5e81a7e5f1..a4fa6d7c02 100644 --- a/src/config.ts +++ b/src/config.ts @@ -1253,6 +1253,8 @@ const configSchema = z.object({ configRebaseProvenance: z.unknown().optional(), // A retry can be billable, so absence and malformed hand edits both stay off. emptyCompletionRetry: z.boolean().optional().catch(false), + // Header suppression changes what Codex sees, so absence and malformed edits stay off. + dropCodexSafetyBuffering: z.boolean().optional().catch(false), // A malformed hand edit must not silently stop opening the browser: fall back // to undefined, which resolves to the historical auto-open behavior. oauthOpenBrowser: z.boolean().optional().catch(undefined), @@ -2858,6 +2860,14 @@ function emptyCompletionRetryError(value: unknown): string | null { return "schema_invalid: emptyCompletionRetry: must be a boolean or omitted"; } +function dropCodexSafetyBufferingError(value: unknown): string | null { + const raw = rawConfigRecord(value); + if (!raw || !Object.hasOwn(raw, "dropCodexSafetyBuffering")) return null; + const enabled = raw.dropCodexSafetyBuffering; + if (enabled === undefined || typeof enabled === "boolean") return null; + return "schema_invalid: dropCodexSafetyBuffering: must be a boolean or omitted"; +} + function oauthOpenBrowserError(value: unknown): string | null { const raw = rawConfigRecord(value); if (!raw || !Object.hasOwn(raw, "oauthOpenBrowser")) return null; @@ -3001,6 +3011,7 @@ export function validateConfigCandidate(value: unknown): { ok: true; config: Ocx ?? codexQuotaAutoRefreshError(value) ?? codexAccountPickerEnabledError(value) ?? emptyCompletionRetryError(value) + ?? dropCodexSafetyBufferingError(value) ?? oauthOpenBrowserError(value) ?? runtimeRoleError(value) ?? remoteGuiConfigError(value) @@ -4022,6 +4033,7 @@ export function getDefaultConfig(): OcxConfig { return { port: 10100, emptyCompletionRetry: false, + dropCodexSafetyBuffering: false, fastRows: true, managementUsageMaxReadBytes: 64 * 1024 * 1024, appOwnedMemoryBudgetMb: DEFAULT_APP_OWNED_MEMORY_BUDGET_BYTES / (1024 * 1024), diff --git a/src/providers/derive.ts b/src/providers/derive.ts index 72a662aee4..3ab1d01500 100644 --- a/src/providers/derive.ts +++ b/src/providers/derive.ts @@ -44,6 +44,7 @@ export interface DerivedKeyLoginProvider { autoToolChoiceOnlyModels?: string[]; preserveReasoningContentModels?: string[]; requiresReasoningPlaceholderModels?: string[]; + showThinkingSummary?: boolean; reasoningSplitModels?: string[]; reasoningDetailsModels?: string[]; thinkingToggleModels?: string[]; @@ -271,6 +272,7 @@ export function providerConfigSeed(entry: ProviderRegistryEntry): OcxProviderCon ...(entry.autoToolChoiceOnlyModels ? { autoToolChoiceOnlyModels: [...entry.autoToolChoiceOnlyModels] } : {}), ...(entry.preserveReasoningContentModels ? { preserveReasoningContentModels: [...entry.preserveReasoningContentModels] } : {}), ...(entry.requiresReasoningPlaceholderModels ? { requiresReasoningPlaceholderModels: [...entry.requiresReasoningPlaceholderModels] } : {}), + ...(entry.showThinkingSummary !== undefined ? { showThinkingSummary: entry.showThinkingSummary } : {}), ...(entry.reasoningSplitModels ? { reasoningSplitModels: [...entry.reasoningSplitModels] } : {}), ...(entry.reasoningDetailsModels ? { reasoningDetailsModels: [...entry.reasoningDetailsModels] } : {}), ...(entry.thinkingToggleModels ? { thinkingToggleModels: [...entry.thinkingToggleModels] } : {}), @@ -320,6 +322,7 @@ export function deriveKeyLoginMap(): Record { ...(entry.autoToolChoiceOnlyModels ? { autoToolChoiceOnlyModels: [...entry.autoToolChoiceOnlyModels] } : {}), ...(entry.preserveReasoningContentModels ? { preserveReasoningContentModels: [...entry.preserveReasoningContentModels] } : {}), ...(entry.requiresReasoningPlaceholderModels ? { requiresReasoningPlaceholderModels: [...entry.requiresReasoningPlaceholderModels] } : {}), + ...(entry.showThinkingSummary !== undefined ? { showThinkingSummary: entry.showThinkingSummary } : {}), ...(entry.reasoningSplitModels ? { reasoningSplitModels: [...entry.reasoningSplitModels] } : {}), ...(entry.reasoningDetailsModels ? { reasoningDetailsModels: [...entry.reasoningDetailsModels] } : {}), ...(entry.thinkingToggleModels ? { thinkingToggleModels: [...entry.thinkingToggleModels] } : {}), @@ -574,6 +577,7 @@ export function enrichProviderFromRegistry(name: string, prov: OcxProviderConfig if (!prov.thinkingToggleModels && seed.thinkingToggleModels) prov.thinkingToggleModels = [...seed.thinkingToggleModels]; if (!prov.thinkingBudgetModels && seed.thinkingBudgetModels) prov.thinkingBudgetModels = [...seed.thinkingBudgetModels]; if (prov.escapeBuiltinToolNames === undefined && seed.escapeBuiltinToolNames !== undefined) prov.escapeBuiltinToolNames = seed.escapeBuiltinToolNames; + if (prov.showThinkingSummary === undefined && seed.showThinkingSummary !== undefined) prov.showThinkingSummary = seed.showThinkingSummary; if (prov.keyOptional === undefined && seed.keyOptional !== undefined) prov.keyOptional = seed.keyOptional; if (prov.freeTier === undefined && seed.freeTier !== undefined) prov.freeTier = seed.freeTier; if (prov.modelSuffixBracketStrip === undefined && seed.modelSuffixBracketStrip !== undefined) prov.modelSuffixBracketStrip = seed.modelSuffixBracketStrip; diff --git a/src/providers/opencode-zen-rate-limit.ts b/src/providers/opencode-zen-rate-limit.ts index c4dbb10319..383c9f70ae 100644 --- a/src/providers/opencode-zen-rate-limit.ts +++ b/src/providers/opencode-zen-rate-limit.ts @@ -175,3 +175,61 @@ export function enrichOpenCodeZenUpstreamMessage( ): string { return enrichOpenCodeZenFreeTierMessage(enrichOpenCodeZenRateLimitMessage(message, opts), opts); } + +/** The effective HTTP endpoint, never the configured row name, identifies Console. */ +export function isConsoleGoDestination(outboundUrl: string | undefined): boolean { + if (!outboundUrl) return false; + try { + const url = new URL(outboundUrl); + return url.protocol === "https:" && url.hostname === "opencode.ai" + && url.port === "" && url.username === "" && url.password === "" + && url.search === "" && url.hash === "" + && /^\/zen\/(?:go\/)?v1\/(?:responses|chat\/completions|messages)$/.test(url.pathname); + } catch { + return false; + } +} + +/** + * The canonical refusal envelope, as served on both Console routes: + * {"error":{"param":null,"type":"invalid_request_error","message":"Error from provider + * (Console Go): Upstream request failed: [invalid_request_error] Invalid upload request."}} + * + * The gateway names itself Console on the Zen key route and Console Go on the Go route, so the + * anchor is the shared product name plus the refusal sentence. The message is matched whole: a + * bare string, a suffix, or any other envelope is a different refusal and must not be replayed. + * Being stricter than necessary is the safe direction: a missed match leaves the turn failing + * exactly as it does today, while a loose match spends an extra request on unrelated 400s. + */ +const CONSOLE_UPLOAD_REFUSALS = new Set([ + "Error from provider (Console Go): Upstream request failed: [invalid_request_error] Invalid upload request.", + "Error from provider (Console): Upstream request failed: [invalid_request_error] Invalid upload request.", +]); + +/** + * True only for the canonical Console upload refusal on a canonical Console destination. + * Route-gated on purpose: the message alone would let any other upstream that happens to answer + * with this English sentence trigger a second send from an unrelated provider. + */ +export function isTransientConsoleGoUploadRejection(opts: { + status: number; + errorBody: string | undefined; + outboundUrl?: string; +}): boolean { + if (opts.status !== 400 || !opts.errorBody) return false; + if (!isConsoleGoDestination(opts.outboundUrl)) return false; + let payload: unknown; + try { + payload = JSON.parse(opts.errorBody); + } catch { + return false; + } + const error = (payload as { error?: unknown } | null)?.error; + if (!error || typeof error !== "object" || Array.isArray(error)) return false; + // The whole envelope, not just the sentence: Console always answers this refusal as + // invalid_request_error with a null param, so a partial envelope is a different error. + const envelope = error as { param?: unknown; type?: unknown; message?: unknown }; + if (envelope.type !== "invalid_request_error" || envelope.param !== null) return false; + // Matched without trimming: padding means the gateway wrapped or appended something. + return typeof envelope.message === "string" && CONSOLE_UPLOAD_REFUSALS.has(envelope.message); +} diff --git a/src/providers/registry.ts b/src/providers/registry.ts index da25a1e7d3..ff065b0561 100644 --- a/src/providers/registry.ts +++ b/src/providers/registry.ts @@ -356,6 +356,10 @@ export interface ProviderRegistryEntry { autoToolChoiceOnlyModels?: string[]; preserveReasoningContentModels?: string[]; requiresReasoningPlaceholderModels?: string[]; + /** + * Opt this provider into visible thinking summaries (see OcxProviderConfig.showThinkingSummary). + */ + showThinkingSummary?: boolean; reasoningSplitModels?: string[]; reasoningDetailsModels?: string[]; thinkingToggleModels?: string[]; @@ -380,7 +384,7 @@ export type ProviderConfigSeed = Pick< | "modelMaxInputTokens" | "defaultMaxOutputTokens" | "modelMaxOutputTokens" | "reasoningEfforts" | "modelReasoningEfforts" | "modelDefaultReasoningEfforts" | "reasoningEffortMap" | "modelReasoningEffortMap" | "reasoningWireFormat" | "noVisionModels" | "noReasoningModels" | "noTemperatureModels" | "noTopPModels" | "noPenaltyModels" - | "autoToolChoiceOnlyModels" | "preserveReasoningContentModels" | "requiresReasoningPlaceholderModels" | "reasoningSplitModels" | "reasoningDetailsModels" | "thinkingToggleModels" | "thinkingBudgetModels" | "escapeBuiltinToolNames" | "openaiChatEofTolerance" + | "autoToolChoiceOnlyModels" | "preserveReasoningContentModels" | "requiresReasoningPlaceholderModels" | "reasoningSplitModels" | "reasoningDetailsModels" | "thinkingToggleModels" | "thinkingBudgetModels" | "escapeBuiltinToolNames" | "openaiChatEofTolerance" | "showThinkingSummary" | "googleMode" | "project" | "location" | "headers" >; @@ -2135,7 +2139,7 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [ // path must stay RELATIVE: this row sets `allowBaseUrlOverride`, and an absolute `url` would // retarget a user's custom base back to Google. A leading `./` is required because a bare // `v1internal:` reads as a URL scheme and `providerModelDiscoverySpecError` rejects it. - { id: "google-antigravity", alias: "agy", label: "Google Antigravity", adapter: "google", baseUrl: "https://daily-cloudcode-pa.googleapis.com", authKind: "oauth", allowBaseUrlOverride: true, dashboardUrl: "https://antigravity.google", models: ANTIGRAVITY_MODELS, liveModels: true, defaultModel: "gemini-3.8-flash", modelContextWindows: ANTIGRAVITY_MODEL_CONTEXT_WINDOWS, modelInputModalities: ANTIGRAVITY_MODEL_INPUT_MODALITIES, modelReasoningEfforts: ANTIGRAVITY_MODEL_EFFORTS, googleMode: "cloud-code-assist", jawcodeBundle: "google", extraMetadataAliases: ["antigravity", "gemini-antigravity"], modelDiscovery: { path: "./v1internal:fetchAvailableModels" } }, + { id: "google-antigravity", alias: "agy", label: "Google Antigravity", adapter: "google", baseUrl: "https://daily-cloudcode-pa.googleapis.com", authKind: "oauth", allowBaseUrlOverride: true, dashboardUrl: "https://antigravity.google", models: ANTIGRAVITY_MODELS, liveModels: true, defaultModel: "gemini-3.8-flash", modelContextWindows: ANTIGRAVITY_MODEL_CONTEXT_WINDOWS, modelInputModalities: ANTIGRAVITY_MODEL_INPUT_MODALITIES, modelReasoningEfforts: ANTIGRAVITY_MODEL_EFFORTS, googleMode: "cloud-code-assist", showThinkingSummary: true, jawcodeBundle: "google", extraMetadataAliases: ["antigravity", "gemini-antigravity"], modelDiscovery: { path: "./v1internal:fetchAvailableModels" } }, { id: "azure-openai", label: "Azure OpenAI", adapter: "azure-openai", baseUrl: "https://{resource}.openai.azure.com/openai", authKind: "key", featured: true, dashboardUrl: "https://portal.azure.com" }, { id: "ollama", label: "Ollama (local)", adapter: "openai-chat", baseUrl: "http://localhost:11434/v1", authKind: "local", allowPrivateNetworkByDefault: true, allowBaseUrlOverride: true, featured: true, note: "Local — key usually blank" }, { id: "vllm", label: "vLLM (local)", adapter: "openai-chat", baseUrl: "http://localhost:8000/v1", authKind: "local", allowPrivateNetworkByDefault: true, allowBaseUrlOverride: true, featured: true, note: "Local — key usually blank" }, diff --git a/src/router.ts b/src/router.ts index bd8e9dc690..70e427b74b 100644 --- a/src/router.ts +++ b/src/router.ts @@ -413,6 +413,13 @@ export function routedProviderConfig(providerName: string, provider: OcxProvider ...(provider.preserveResponsesReasoningContent === undefined && registryEntry.preserveResponsesReasoningContent !== undefined ? { preserveResponsesReasoningContent: registryEntry.preserveResponsesReasoningContent } : {}), + // The request path resolves through routedProviderConfig() and never calls + // enrichProviderFromRegistry(), so a saved provider row written before the + // registry learned this flag must be backfilled here or route.provider never + // carries it and the showThinkingSummary opt-in stays dead. + ...(provider.showThinkingSummary === undefined && registryEntry.showThinkingSummary !== undefined + ? { showThinkingSummary: registryEntry.showThinkingSummary } + : {}), // Registry-only client-facing repair policy (#938): fill only when the // saved provider has no explicit policy; clone so runtime never aliases // the registry constant. diff --git a/src/server/auth-cors.ts b/src/server/auth-cors.ts index 4bf99b4f0f..e7b8d645d6 100644 --- a/src/server/auth-cors.ts +++ b/src/server/auth-cors.ts @@ -942,6 +942,7 @@ const PROVIDER_CONFIG_FIELD_POLICY = { autoToolChoiceOnlyModels: "editor", preserveReasoningContentModels: "editor", requiresReasoningPlaceholderModels: "editor", + showThinkingSummary: "editor", retryOn429: "editor", transientRetryOn5xx: "editor", reasoningSplitModels: "editor", diff --git a/src/server/index.ts b/src/server/index.ts index 3e458a3fea..150dd0f8ca 100644 --- a/src/server/index.ts +++ b/src/server/index.ts @@ -150,6 +150,7 @@ import { } from "./relay"; export { consumeForInspection, + codexSafetyBufferingFilterOptions, relaySseWithFailedTail, relaySseWithHeartbeat, relayWithAbort, diff --git a/src/server/relay-eager.ts b/src/server/relay-eager.ts index 7844c65886..8439aea85d 100644 --- a/src/server/relay-eager.ts +++ b/src/server/relay-eager.ts @@ -26,6 +26,7 @@ import { adapterEofIncompleteFrame, + type CodexSafetyBufferingFilterOptions, createSseTerminalOutputBoundary, doneFrame, failedTailFrame, @@ -84,6 +85,8 @@ export type EagerRelayOptions = { postCancelDrainBytes?: number; /** Last known upstream failure to preserve when EOF would otherwise become adapter_eof. */ upstreamError?: string; + /** Optional client-facing hint policy; inspection retains original frames. */ + terminalBoundary?: CodexSafetyBufferingFilterOptions; /** Injectable clock for tests. */ now?: () => number; }; @@ -114,7 +117,7 @@ export function relaySseEagerBounded( const terminalEncoder = new TextEncoder(); const adapterEofFrame = adapterEofIncompleteFrame(terminalEncoder); const terminalSentinel = doneFrame(terminalEncoder); - const terminalBoundary = createSseTerminalOutputBoundary(); + const terminalBoundary = createSseTerminalOutputBoundary(opts?.terminalBoundary); const activeRewrite: SseBlockRewrite | undefined = hooks.rewriteBlocks ?? (hooks.rewritePayload ? payloadRewriteAsBlockRewrite(hooks.rewritePayload) : undefined); const encodeFailedTail = (error: unknown): Uint8Array | null => { diff --git a/src/server/relay.ts b/src/server/relay.ts index a483d88f20..f480ace68a 100644 --- a/src/server/relay.ts +++ b/src/server/relay.ts @@ -183,7 +183,10 @@ export type SseTerminalOutputBoundary = { * terminal, and drops every later block/byte. A premature [DONE] is held until * a terminal arrives so clean EOF can synthesize one terminal and one sentinel. */ -export function createSseTerminalOutputBoundary(): SseTerminalOutputBoundary { +export function createSseTerminalOutputBoundary( + options?: CodexSafetyBufferingFilterOptions, +): SseTerminalOutputBoundary { + const dropSafetyBuffering = options?.dropCodexSafetyBuffering === true; const decoder = new TextDecoder(); const encoder = new TextEncoder(); const framer = new BoundedSseFrameBuffer(MAX_INSPECTION_SSE_FRAME_BYTES); @@ -207,15 +210,22 @@ export function createSseTerminalOutputBoundary(): SseTerminalOutputBoundary { // behind EOF, so its log context cannot determine the outgoing terminal. const message = boundedBareUpstreamErrorMessage(parsed); if (message !== undefined) upstreamError = message; + const safetyBuffering = dropSafetyBuffering && parsed !== undefined + ? codexSafetyBufferingBlockAction(parsed) : "keep"; + if (safetyBuffering === "drop") continue; const policyError = parsed !== undefined && isPolicyRewriteType(parsed) ? cyberPolicyTerminalError(parsed) : undefined; - const outboundBlock = policyError - ? encoder.encode(rewritePolicyTerminalBlock( - decoder.decode(frame.block), - policyFailurePayload(policyError, parsed), - )) + const policyPayload = policyError ? policyFailurePayload(policyError, parsed) : undefined; + let outboundBlock = policyPayload !== undefined + ? encoder.encode(rewritePolicyTerminalBlock(decoder.decode(frame.block), policyPayload)) : frame.block; + if (safetyBuffering === "strip") { + outboundBlock = encoder.encode(stripCodexSafetyBufferingField( + decoder.decode(outboundBlock), + policyPayload !== undefined ? parseSsePayload(policyPayload) : parsed, + )); + } if (isDone) { done = true; if (responsesTerminal) { @@ -287,11 +297,11 @@ export function relaySseWithFailedTail( body: ReadableStream, upstream: AbortController, onClientGone?: (reason?: unknown) => void, - opts?: { upstreamError?: string }, + opts?: { upstreamError?: string; terminalBoundary?: CodexSafetyBufferingFilterOptions }, ): ReadableStream { const reader = body.getReader(); const encoder = new TextEncoder(); - const terminalBoundary = createSseTerminalOutputBoundary(); + const terminalBoundary = createSseTerminalOutputBoundary(opts?.terminalBoundary); let closed = false; const relayChunk = ( controller: ReadableStreamDefaultController, @@ -468,6 +478,29 @@ function isPolicyRewriteType(parsed: unknown): boolean { return type === "response.failed" || type === "response.incomplete" || type === "error"; } +/** + * Codex emits its safety-buffering hint in the SSE body as well as in headers: + * a `response.metadata` event whose `metadata.type` is `safety_buffering`, or a + * `safety_buffering` field on another event. The metadata event is dropped whole; + * the field is stripped so the carrying event is otherwise relayed unchanged. + */ +function codexSafetyBufferingBlockAction(parsed: unknown): "keep" | "drop" | "strip" { + const root = asJsonRecord(parsed); + if (!root) return "keep"; + if (root.type === "response.metadata") { + const metadata = asJsonRecord(root.metadata); + if (metadata?.type === "safety_buffering") return "drop"; + } + return Object.hasOwn(root, "safety_buffering") ? "strip" : "keep"; +} + +function stripCodexSafetyBufferingField(block: string, parsed: unknown): string { + const root = asJsonRecord(parsed); + if (!root) return block; + const { safety_buffering: _safetyBuffering, ...rest } = root; + return replaceSseDataPayload(block, JSON.stringify(rest)); +} + function rewritePolicyTerminalBlock(block: string, payload: string): string { const newline = block.includes("\r\n") ? "\r\n" : "\n"; const rewritten = replaceSseDataPayload(block, payload); @@ -1461,7 +1494,31 @@ export function consumeForResponseLogMetadata( * body makes the caller (Codex) double-decode / truncate → "stream error" on every gpt passthrough. * Drop encoding + hop-by-hop headers; relay everything else (content-type, etc.) verbatim. */ -export function sanitizePassthroughHeaders(upstream: Headers): Headers { +export const CODEX_SAFETY_BUFFERING_HEADERS = [ + "x-codex-safety-buffering-enabled", + "x-codex-safety-buffering-faster-model", +] as const; + +const CODEX_SAFETY_BUFFERING_HEADER_SET: ReadonlySet = new Set(CODEX_SAFETY_BUFFERING_HEADERS); + +export interface CodexSafetyBufferingFilterOptions { + /** + * Drop Codex safety-buffering hints: the `x-codex-safety-buffering-*` response + * headers and the `safety_buffering` SSE metadata event / field. Absent and + * `false` relay everything unchanged. + */ + dropCodexSafetyBuffering?: boolean; +} + +/** Resolve the passthrough header policy from the loaded config (absent means "forward everything"). */ +export function codexSafetyBufferingFilterOptions( + config: { dropCodexSafetyBuffering?: boolean }, +): CodexSafetyBufferingFilterOptions { + return { dropCodexSafetyBuffering: config.dropCodexSafetyBuffering === true }; +} + +export function sanitizePassthroughHeaders(upstream: Headers, options?: CodexSafetyBufferingFilterOptions): Headers { + const dropSafetyBuffering = options?.dropCodexSafetyBuffering === true; const DROP = new Set([ "content-encoding", "content-length", @@ -1478,7 +1535,10 @@ export function sanitizePassthroughHeaders(upstream: Headers): Headers { ]); const out = new Headers(); upstream.forEach((value, key) => { - if (!DROP.has(key.toLowerCase())) out.set(key, value); + const lower = key.toLowerCase(); + if (DROP.has(lower)) return; + if (dropSafetyBuffering && CODEX_SAFETY_BUFFERING_HEADER_SET.has(lower)) return; + out.set(key, value); }); return out; } diff --git a/src/server/responses-reasoning-summary-rewrite.ts b/src/server/responses-reasoning-summary-rewrite.ts deleted file mode 100644 index 55c6d8ae7b..0000000000 --- a/src/server/responses-reasoning-summary-rewrite.ts +++ /dev/null @@ -1,178 +0,0 @@ -import type { SsePayloadRewrite } from "./sse-payload-rewrite"; - -/** - * Route content-channel reasoning from native-Responses upstreams through the - * expandable summary channel (issue #45). - * - * Codex renders the expandable reasoning trace from the Responses reasoning - * item's `summary[]` channel. DeepSeek's native `/responses` endpoint emits - * raw thinking on the content channel instead (`response.reasoning_text.delta` - * plus items with `content: [{type: "reasoning_text", text}]` and an empty - * `summary`), so routed DeepSeek turns showed the "Worked for Xs" timer with - * nothing to expand. Native OpenAI upstreams already emit summary-channel - * events; this rewrite is a no-op for them (no reasoning_text events to - * rewrite) and only engages when the upstream produces content-channel - * reasoning. - * - * Replay compatibility: Codex echoes the reasoning item it received back into - * the next request's input. DeepSeek's Responses API accepts summary-shaped - * reasoning input items (verified live), so the rewrite round-trips. - */ - -function isPlainObject(value: unknown): value is Record { - return !!value && typeof value === "object" && !Array.isArray(value); -} - -function reasoningTextOf(item: Record): string { - if (!Array.isArray(item.content)) return ""; - return item.content - .filter((part): part is Record => isPlainObject(part) && part.type === "reasoning_text") - .map(part => (typeof part.text === "string" ? part.text : "")) - .join(""); -} - -/** Move a reasoning item's content channel into the summary channel. */ -function reasoningItemToSummaryShape(item: Record): Record { - if (item.type !== "reasoning") return item; - // `encrypted_content` is opaque, state-bearing provider data, so the entire item must retain its - // upstream shape unless that backend has an explicit replay contract permitting a rewrite. This - // defensively protects content-channel backends that do issue blobs when the client replays the - // stored item. The delta rewrite can still provide the expandable trace for the live turn. - // DeepSeek — the provider this rewrite was verified against — is `statelessResponses` and issues - // no blob, so it is unaffected. - if (typeof item.encrypted_content === "string" && item.encrypted_content.length > 0) return item; - const text = reasoningTextOf(item); - // Items that already use the summary channel (or carry no content text at - // all) are left untouched: rewriting them could clear a valid summary. - if (text.length === 0) return item; - const next: Record = { ...item }; - delete next.content; - next.summary = [{ type: "summary_text", text }]; - return next; -} - -/** - * Rewrite one parsed SSE payload in place of the content channel, or return - * `null` when nothing changed (caller keeps the original payload). - */ -function rewritePayload(payload: Record): Record | null { - switch (payload.type) { - case "response.reasoning_text.delta": { - const next: Record = { - type: "response.reasoning_summary_text.delta", - item_id: payload.item_id, - output_index: payload.output_index, - summary_index: 0, - delta: payload.delta, - }; - if (payload.sequence_number !== undefined) next.sequence_number = payload.sequence_number; - return next; - } - case "response.reasoning_text.done": { - const next: Record = { - type: "response.reasoning_summary_text.done", - item_id: payload.item_id, - output_index: payload.output_index, - summary_index: 0, - text: payload.text, - }; - if (payload.sequence_number !== undefined) next.sequence_number = payload.sequence_number; - return next; - } - default: { - let changed = false; - const next: Record = { ...payload }; - if (isPlainObject(next.item) && next.item.type === "reasoning") { - const rewritten = reasoningItemToSummaryShape(next.item); - if (rewritten !== next.item) { - next.item = rewritten; - changed = true; - } - } - // SSE event shape: {type: "response.completed", response: {output}}. - const response = isPlainObject(next.response) ? { ...next.response } : null; - if (response && Array.isArray(response.output)) { - const output = response.output.map(item => { - if (!isPlainObject(item) || item.type !== "reasoning") return item; - const rewritten = reasoningItemToSummaryShape(item); - if (rewritten !== item) changed = true; - return rewritten; - }); - if (changed) { - response.output = output; - next.response = response; - } - } - // Bare response document shape (non-streaming passthrough): - // {object: "response", output: [...]}. - if (Array.isArray(next.output)) { - const output = next.output.map(item => { - if (!isPlainObject(item) || item.type !== "reasoning") return item; - const rewritten = reasoningItemToSummaryShape(item); - if (rewritten !== item) changed = true; - return rewritten; - }); - if (changed) next.output = output; - } - return changed ? next : null; - } - } -} - -/** Payload rewrite for passthrough relays whose upstream emits content-channel reasoning. */ -export function createReasoningSummaryChannelPayloadRewrite(): SsePayloadRewrite { - return (payload: string): string => { - let parsed: unknown; - try { - parsed = JSON.parse(payload); - } catch { - return payload; - } - if (!isPlainObject(parsed)) return payload; - const rewritten = rewritePayload(parsed); - return rewritten !== null ? JSON.stringify(rewritten) : payload; - }; -} - -/** - * Object-level variant for the non-streaming passthrough: the bounded-JSON - * relay bypasses the SSE payload rewrite, so reasoning items inside a full - * Responses JSON document need the same normalization before plain JSON - * serialization or forced JSON-to-SSE reframing. Returns the same reference - * when nothing changed. - */ -export function rewriteReasoningSummaryInJson(value: unknown): unknown { - if (!isPlainObject(value)) return value; - const rewritten = rewritePayload(value); - return rewritten !== null ? rewritten : value; -} - -/** String-level variant of {@link rewriteReasoningSummaryInJson}. */ -export function rewriteReasoningSummaryInJsonString(json: string): string { - let parsed: unknown; - try { - parsed = JSON.parse(json); - } catch { - return json; - } - const rewritten = rewriteReasoningSummaryInJson(parsed); - return rewritten === parsed ? json : JSON.stringify(rewritten); -} - -/** - * True when a routed native-Responses provider emits content-channel reasoning - * (raw `reasoning_text`) instead of the summary channel. DeepSeek's - * `/responses` endpoint is the current example: it ships raw thinking with an - * empty `summary` and keeps `preserveReasoningContentModels` so multi-turn - * replays round-trip. - */ -export function routeUsesContentChannelReasoning( - provider: { statelessResponses?: boolean; preserveReasoningContentModels?: string[] }, - modelId: string, -): boolean { - if (provider.statelessResponses === true) return true; - const preserved = provider.preserveReasoningContentModels; - const normalizedModelId = modelId.toLowerCase(); - return Array.isArray(preserved) - && preserved.some(id => id.toLowerCase() === normalizedModelId); -} diff --git a/src/server/responses/core.ts b/src/server/responses/core.ts index 3952db7afc..6be74bc20a 100644 --- a/src/server/responses/core.ts +++ b/src/server/responses/core.ts @@ -104,7 +104,10 @@ import { } from "../../lib/errors"; import { injectionDebugLog } from "../../lib/injection-debug-log"; import { resolveClientRetryAfter } from "../../lib/retry-after"; -import { enrichOpenCodeZenUpstreamMessage } from "../../providers/opencode-zen-rate-limit"; +import { + enrichOpenCodeZenUpstreamMessage, + isTransientConsoleGoUploadRejection, +} from "../../providers/opencode-zen-rate-limit"; import { CODE_MODE_EXEC_TOOL_NAME, modelInList, namespacedToolName } from "../../types"; import type { AdapterEvent, @@ -216,6 +219,7 @@ import { isNonReplayableResponse, isTransientUpstreamStatus, prepareSameTarget429Wait, + sleepWithAbort, } from "../../lib/upstream-retry"; import { ForwardAdmissionCredentialError, @@ -346,6 +350,7 @@ import { markEagerRelaySseResponse, markNativePassthroughSseResponse, relaySseWithFailedTail, + codexSafetyBufferingFilterOptions, relayWithAbort, sanitizePassthroughHeaders, } from "../relay"; @@ -369,12 +374,6 @@ import { hasResponsesItemIdRepair, repairResponsesJsonItemIds, } from "../responses-item-id-repair"; -import { - createReasoningSummaryChannelPayloadRewrite, - rewriteReasoningSummaryInJson, - rewriteReasoningSummaryInJsonString, - routeUsesContentChannelReasoning, -} from "../responses-reasoning-summary-rewrite"; import { createImageGenCallRestoreRewrite, imageGenToolCallAliases, @@ -859,6 +858,31 @@ async function opaqueBlobRejectionBodyForRecovery( } } +/** + * Backoff for the single exact-request replay after a canonical Console upload rejection. + */ +const CONSOLE_GO_UPLOAD_RETRY_DELAY_MS = 800; + +/** + * Peek the upstream error body for the Console Go transient-400 recovery. Only a complete, + * display-safe body may drive a retry decision (same contract as + * opaqueBlobRejectionBodyForRecovery), and reading a clone leaves the original response intact + * for the caller's own error surface when no retry is taken. + */ +async function consoleGoUploadRejectionBody( + response: Response, + alreadyAttempted: boolean, + signal: AbortSignal, +): Promise { + if (isNonReplayableResponse(response) || response.status !== 400 || alreadyAttempted) return undefined; + try { + const body = await readBoundedResponseBody(response.clone(), { signal }); + return body.displaySafe && !body.truncated ? body.text : undefined; + } catch { + return undefined; + } +} + /** * Materialize an upstream error body only when the bounded reader observed a complete, * display-safe payload. Partial timeout and over-limit prefixes are attacker-controlled, @@ -2578,6 +2602,13 @@ async function applyFinalRouteRequestNormalization(args: { route.provider = resolveOpenCodeGoTransport(route.provider, args.claudeGoAffinity ? args.claudeGoAffinity.sessionLane : getOrAllocateRequestSessionLane(req)); route.provider = resolveWireProtocolOverride(route.providerName, route.modelId, route.provider, inboundWire); + // Recompute from the original wire preference on every route, including fallback. + // A provider default never converts raw reasoning into a summary. + if (inboundWire === "responses" && parsed._rawBody) { + const summary = (parsed._rawBody as { reasoning?: { summary?: unknown } }).reasoning?.summary; + parsed.options.hideThinkingSummary = summary === "none" + || (!summary && route.provider.showThinkingSummary !== true); + } if (preserveAnthropicResponseModel) parsed._responseModelId = responseModelId; logCtx.model = route.modelId; logCtx.provider = route.providerName; @@ -2933,6 +2964,7 @@ export async function handleComboResponses( pick.target, comboDefaultEffort(config, comboId), supportedLadderFor({ provider: targetRoute.provider, modelId: targetRoute.modelId }), + combo.reasoningEffortMode, ); const childHeaders = buildComboChildHeaders(req.headers); const childRequest = new Request(req.url, { @@ -4820,6 +4852,9 @@ async function handleResponsesInner( let hostAdmissionLease = pendingHostAdmissionLease; pendingHostAdmissionLease = null; try { + const codexSafetyBufferingOptions = isCanonicalOpenAiForwardProvider(route.provider) + ? codexSafetyBufferingFilterOptions(config) + : undefined; const imageGenCallAliases = route.provider.authMode === "forward" ? new Map() : imageGenToolCallAliases(toolBridgeMaps.toolNsMap, parsed._rawBody, translatorBudget); @@ -5124,10 +5159,7 @@ async function handleResponsesInner( ? JSON.parse(normalizeFunctionCompletionJson(JSON.stringify(restored))) : restored) as { id?: unknown; output?: unknown; status?: unknown }; // Replay overlap compares the items the client echoes, including visible reasoning shape. - const replayResponse = parsed.options.hideThinkingSummary !== true - && routeUsesContentChannelReasoning(route.provider, route.modelId) - ? rewriteReasoningSummaryInJson(restoredResponse) as typeof restoredResponse - : restoredResponse; + const replayResponse = restoredResponse; if ( undeclaredToolGuardActive && undeclaredToolCallNameInResponse( @@ -5351,6 +5383,9 @@ async function handleResponsesInner( const opaqueBlobRecoveryGuard: OpaqueBlobRecoveryGuard = { attempted: false }; let oauth401ReplayAttempted = false; let codex401ReplayKind: "main" | "stored" | null = null; + // Console Go answers a transient 400 "Invalid upload request." for bodies it accepts + // moments later; at most one byte-identical replay is allowed per request. + const consoleGoUploadRetryGuard: { attempted: boolean } = { attempted: false }; const rateLimitPolicy = rateLimitRetryPolicyFor(route.provider); let rateLimitRetries = 0; const rebuildAndRefetch = async ( @@ -5365,10 +5400,12 @@ async function handleResponsesInner( return { failed: formatErrorResponse(502, "upstream_error", "Recovery changed the provider wire unexpectedly") }; } try { - request = await retryAdapter.buildRequest(parsed, { - headers: selectedForwardHeaders, - translatorBudget, - }); + if (recovery !== "console-go-upload-retry") { + request = await retryAdapter.buildRequest(parsed, { + headers: selectedForwardHeaders, + translatorBudget, + }); + } refreshRoutedNamespaceToolAliases(request); recordAdapterReasoning(logCtx, request); recordAdapterTier(logCtx, request); @@ -5902,9 +5939,39 @@ async function handleResponsesInner( logCtx.terminalIncompleteReason = preflightLog.terminalIncompleteReason; } } + // Console Go (opencode-zen / opencode-go) intermittently rejects a body it accepts seconds + // later with 400 invalid_request_error / "Invalid upload request." Replay the byte-identical + // request once after the exact gateway rejection. Single-shot guard. + // This recovery reuses the captured request; other recovery kinds still rebuild. + if (!consoleGoUploadRetryGuard.attempted) { + const uploadRejectionBody = await consoleGoUploadRejectionBody( + upstreamResponse, + consoleGoUploadRetryGuard.attempted, + upstream.signal, + ); + if (uploadRejectionBody !== undefined + && isTransientConsoleGoUploadRejection({ + status: upstreamResponse.status, + errorBody: uploadRejectionBody, + outboundUrl: request.url, + })) { + consoleGoUploadRetryGuard.attempted = true; + try { void upstreamResponse.body?.cancel().catch(() => {}); } catch { /* already consumed/closed */ } + if (!upstream.signal.aborted) { + try { + await sleepWithAbort(CONSOLE_GO_UPLOAD_RETRY_DELAY_MS, upstream.signal); + } catch { return clientCancelledResponse(); } + } + if (upstream.signal.aborted) return clientCancelledResponse(); + const result = await rebuildAndRefetch("console-go-upload-retry"); + if ("failed" in result) return result.failed; + upstreamResponse = result; + continue passthroughRecovery; + } + } break; } - const headers = sanitizePassthroughHeaders(upstreamResponse.headers); + const headers = sanitizePassthroughHeaders(upstreamResponse.headers, codexSafetyBufferingOptions); const resolvedModel = headers.get("openai-model")?.trim(); if (resolvedModel && !logCtx.preserveResolvedModelFromRoute) logCtx.resolvedModel = resolvedModel; if (isUsageDebugEnabled()) { @@ -5993,7 +6060,7 @@ async function handleResponsesInner( return new Response(upstreamResponse.body, { status: upstreamResponse.status, statusText: upstreamResponse.statusText, - headers: sanitizePassthroughHeaders(upstreamResponse.headers), + headers: sanitizePassthroughHeaders(upstreamResponse.headers, codexSafetyBufferingOptions), }); } if (!upstreamResponse.ok) { @@ -6128,10 +6195,6 @@ async function handleResponsesInner( ? createResponsesItemIdPayloadRewrite(repairConfig!, translatorBudget) : undefined, responseModelRewrite, - parsed.options.hideThinkingSummary !== true - && routeUsesContentChannelReasoning(route.provider, route.modelId) - ? createReasoningSummaryChannelPayloadRewrite() - : undefined, ].filter((rewrite): rewrite is NonNullable => rewrite !== undefined); // #893: sparse-snapshot gateways get field backfills AND lifecycle event // injection at the block level, after payload rewrites. Defaults come @@ -6262,6 +6325,7 @@ async function handleResponsesInner( onDone: () => unregisterTurn(turnAc), }, { clientGoneSignal: options.abortSignal, + terminalBoundary: codexSafetyBufferingOptions, ...(inlineEagerRewrite ? { rewriteBudget: translatorBudget } : {}), ...(logCtx.upstreamError === undefined ? {} : { upstreamError: logCtx.upstreamError }), }); @@ -6353,7 +6417,7 @@ async function handleResponsesInner( responseCompletionCancelled = true; clientGone.abort(reason); }, - { upstreamError: logCtx.upstreamError }, + { upstreamError: logCtx.upstreamError, terminalBoundary: codexSafetyBufferingOptions }, ); return markNativePassthroughSseResponse(new Response(clientBody, { status: upstreamResponse.status, @@ -6405,13 +6469,7 @@ async function handleResponsesInner( const modelRewritten = parsed._responseModelId !== undefined && parsed._responseModelId !== parsed.modelId ? rewriteResponsesModelJson(repaired, parsed._responseModelId) : repaired; - // The bounded-JSON answer bypasses the SSE payload rewrite, so content- - // channel reasoning needs the same normalization here for the plain - // JSON answer and every reframed-SSE variant built from clientJson. - return parsed.options.hideThinkingSummary !== true - && routeUsesContentChannelReasoning(route.provider, route.modelId) - ? rewriteReasoningSummaryInJsonString(modelRewritten) - : modelRewritten; + return modelRewritten; })(); // #1700: same fail-closed policy as the SSE relay above. Both the plain JSON answer and // the reframed-SSE branch below are built from this body, so one check covers them. This @@ -6487,7 +6545,7 @@ async function handleResponsesInner( } throw error; } - const sseHeaders = sanitizePassthroughHeaders(headers); + const sseHeaders = sanitizePassthroughHeaders(headers, codexSafetyBufferingOptions); sseHeaders.set("content-type", "text/event-stream"); sseHeaders.set("cache-control", "no-store"); return new Response(stream, { @@ -7397,6 +7455,9 @@ async function handleResponsesInner( // 413→429 rotation cannot silently undo the tightening. let imageRetryAttempted = false; const opaqueBlobRecoveryGuard: OpaqueBlobRecoveryGuard = { attempted: false }; + // Console Go answers a transient 400 "Invalid upload request." for bodies it accepts + // moments later; at most one byte-identical replay is allowed per request. + const consoleGoUploadRetryGuard: { attempted: boolean } = { attempted: false }; let oauth401ReplayAttempted = false; /** * Rebuild the request from the current parsed input (and any image-tier bias) and refetch @@ -7781,6 +7842,35 @@ async function handleResponsesInner( upstreamResponse = result; continue recovery; } + // Console Go (opencode-zen / opencode-go) intermittently rejects a body it accepts seconds + // later with 400 invalid_request_error / "Invalid upload request." Replay the + // byte-identical request once after the exact gateway rejection. + if (!consoleGoUploadRetryGuard.attempted) { + const uploadRejectionBody = await consoleGoUploadRejectionBody( + upstreamResponse, + consoleGoUploadRetryGuard.attempted, + upstream.signal, + ); + if (uploadRejectionBody !== undefined + && isTransientConsoleGoUploadRejection({ + status: upstreamResponse.status, + errorBody: uploadRejectionBody, + outboundUrl: sameTargetRequest?.url, + })) { + consoleGoUploadRetryGuard.attempted = true; + try { void upstreamResponse.body?.cancel().catch(() => {}); } catch { /* already consumed/closed */ } + if (!upstream.signal.aborted) { + try { + await sleepWithAbort(CONSOLE_GO_UPLOAD_RETRY_DELAY_MS, upstream.signal); + } catch { cleanupUpstreamAbort(); return clientCancelledResponse(); } + } + if (upstream.signal.aborted) { cleanupUpstreamAbort(); return clientCancelledResponse(); } + const result = await rebuildAndRefetch("console-go-upload-retry"); + if ("failed" in result) return result.failed; + upstreamResponse = result; + continue recovery; + } + } break; } if (!upstreamResponse.ok) { diff --git a/src/types/config.ts b/src/types/config.ts index 52c4ed8b0e..faaa594473 100644 --- a/src/types/config.ts +++ b/src/types/config.ts @@ -380,6 +380,8 @@ export interface OcxConfig { privacy?: OcxPrivacyConfig; /** Opt in to one identical-turn retry when a Responses completion has no text or tool call. */ emptyCompletionRetry?: boolean; + /** Suppress allowlisted client-facing Codex transport hints; provider enforcement is unchanged. */ + dropCodexSafetyBuffering?: boolean; /** * Whether a login may open a browser on the machine running the proxy. * @@ -967,8 +969,9 @@ export type OcxComboDefaultEffort = "low" | "medium" | "high" | "xhigh" | "max" * advertises no effort control (`reasoningEfforts: []`) empties the combo's picker. * `adaptive` excludes those empty ladders from the published intersection, keeping the * control usable for a mixed-capability group. Unknown (`undefined`) ladders stay - * wildcards in both modes. Dispatch is unchanged: each concrete target still resolves - * its own effort at request time. + * wildcards in both modes. An explicit empty ladder removes unsupported effort controls + * in either mode; adaptive dispatch also removes them before sending to an unknown target, + * while each known target still resolves its own effort. */ export type OcxComboReasoningEffortMode = "strict" | "adaptive"; diff --git a/src/types/provider.ts b/src/types/provider.ts index d559ffef8f..bf37ac708a 100644 --- a/src/types/provider.ts +++ b/src/types/provider.ts @@ -773,6 +773,12 @@ export interface OcxProviderConfig { * out explicitly (e.g. MiniMax, where low effort disables thinking). */ requiresReasoningPlaceholderModels?: string[]; + /** + * Default to displaying provider-authored summaries when Responses summary is omitted. + * Explicit wire summary:"none" wins; false disables a seeded provider default. + * Raw reasoning is never relabeled as a summary. + */ + showThinkingSummary?: boolean; /** * Opt-in same-target 429 retry policy. Codex itself never retries 429 (it retries 5xx only, * openai/codex#30471), and single-key pools have no failover, so the proxy waits and replays diff --git a/src/usage/log.ts b/src/usage/log.ts index 2944c22f9a..953195b7ca 100644 --- a/src/usage/log.ts +++ b/src/usage/log.ts @@ -70,6 +70,7 @@ export type AttemptRecoveryKind = | "anthropic-oauth-429" | "oauth-account-429" | "image-413" + | "console-go-upload-retry" | "opaque-blob-rejection" | "empty-completion"; @@ -309,6 +310,7 @@ const ATTEMPT_RECOVERY_KINDS = new Set([ "anthropic-oauth-429", "oauth-account-429", "image-413", + "console-go-upload-retry", "opaque-blob-rejection", "empty-completion", ]); diff --git a/structure/adapters/registry.md b/structure/adapters/registry.md index aa0bc914cb..594565ac73 100644 --- a/structure/adapters/registry.md +++ b/structure/adapters/registry.md @@ -72,6 +72,15 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). + +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Adapter events distinguish raw reasoning content from summary-channel thinking; CCA Gemini classification is request-local. See [Google provenance](../providers/google.md). + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/catalog.md b/structure/catalog.md index 49f0e353e1..73c27cbb0f 100644 --- a/structure/catalog.md +++ b/structure/catalog.md @@ -240,6 +240,11 @@ wire-clamps ultra/max to each model's real top rung (e.g. gpt-5.5 ultra → xhig (`src/server/effort-policy.ts`): they lower or preserve the requested effort rather than rejecting the request, and they never raise it. +Combo dispatch reads the final target ladder through the same `supportedLadderFor` authority. An +explicit empty ladder means that target receives no effort control; an unknown ladder receives no +parent effort controls only when the combo opts into `reasoningEffortMode: "adaptive"`. Known +non-empty ladders continue through the existing per-target resolution. + The `ocx effort` CLI accepts only the same canonical cap ladder before live probing or persistence. Its status output preserves unsupported legacy cap values and reports that those fields are ignored; the read does not normalize or migrate them, and an ignored subagent field does not disable a valid @@ -274,6 +279,11 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](transports/responses.md). + +Provider `showThinkingSummary` is a Responses request default; it does not rewrite catalog summary defaults or client configuration. See [Google summaries](providers/google.md). + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. diff --git a/structure/clients/claude-desktop.md b/structure/clients/claude-desktop.md index 54689b36a1..c8ee45ed88 100644 --- a/structure/clients/claude-desktop.md +++ b/structure/clients/claude-desktop.md @@ -83,6 +83,11 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Provider summary defaults are Responses-specific and do not rewrite connected Claude Desktop profiles. See [inbound compatibility](../data-planes/inbound-compat.md). + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. @@ -95,3 +100,5 @@ Config JSON preserves the boolean; only literal true activates the role-changing The lightweight top-level CLI help counts Cline CLI among the fifteen registered export clients; registry parity remains covered by the client help and integration tests. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/config.md b/structure/config.md index b37d315fce..908acb4630 100644 --- a/structure/config.md +++ b/structure/config.md @@ -45,7 +45,7 @@ matters for maintainers is which groups exist and who resolves them: | Group | Keys | Resolution rule | | --- | --- | --- | | Listener | `port`, `hostname` | The listener owns the port; `runtime-port.json` reports where it actually landed. | -| Routing | `defaultProvider`, `providers`, per-provider `selectedModels` | Explicit `provider/model` wins over `defaultProvider`. | +| Routing | `defaultProvider`, `providers`, per-provider `selectedModels`, `combos` | Explicit `provider/model` wins over `defaultProvider`; combo dispatch uses the selected target's existing capability ladder and does not create a second catalog authority. | | Catalog | `disabledModels`, `customModels`, `modelCacheTtlMs`, `providerContextCaps`, `contextCapValue`, per-provider `modelDisplayNames`, `codexAccountNamespaces`, `codexAccountPickerEnabled` | Catalog state is derived; config only records intent. Exact provider model display names are durable display only overlays. The picker flag is an explicit visibility override, while selector mappings remain the durable exact-routing contract. | | Retained state | `appOwnedMemoryBudgetMb` | Process-wide eviction target for app-owned logs, caches, blobs, and continuation payloads. Default 256 MiB, valid 64..4096; pinned state may temporarily exceed the target, but every pin-capable store has a finite local cap and their documented aggregate stays below `APP_OWNED_WORST_CASE_PINNED_BYTES` (512 MiB). Neither value caps RSS or native runtime memory. | | Transport | stream mode, timeouts, proxy settings, `websockets`, `emptyCompletionRetry` | `streamMode` persists in config.json; Windows services need a persisted input, and macOS uses it for explicit eager-relay opt-in. Empty-completion replay is an explicit top-level opt-in because its second upstream request may be billable. | @@ -203,6 +203,10 @@ Client connection metadata stores a stable `apiKeyId` and a non-secret rotation Codex display-cache expiry, retained main-policy evidence, and reset history follow the [quota cache contract](providers/openai-tiers.md#quota-cache-and-short-window-history). +`dropCodexSafetyBuffering` is an optional boolean, default false. Invalid API candidates reject; +malformed persisted values stay disabled. It controls only the allowlisted client-output hints +described in [Responses transport](transports/responses.md), not upstream policy or model selection. + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. diff --git a/structure/data-planes/images.md b/structure/data-planes/images.md index d1d0193048..acdde5ab97 100644 --- a/structure/data-planes/images.md +++ b/structure/data-planes/images.md @@ -77,6 +77,13 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +CCA image-capable requests do not acquire the text-summary includeThoughts opt-in. See [Google summary boundary](../providers/google.md). + Claude replay carries [Go conversation affinity](inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/data-planes/inbound-compat.md b/structure/data-planes/inbound-compat.md index 3e940b27a4..e77d788d9e 100644 --- a/structure/data-planes/inbound-compat.md +++ b/structure/data-planes/inbound-compat.md @@ -19,6 +19,9 @@ the native passthrough there is no canonical Fast injection and no wire mapping: and `fastMode` injects nothing here. Resolved-Fast-policy injection applies only to routes that take the Chat -> Responses -> Chat bridge below. `parallel_tool_calls` is emitted only for providers opted into parallel tools (or pinned false by the existing provider opt-out contract). +The native passthrough still applies the existing model capability authority to reasoning: an +explicit empty ladder removes caller `reasoning_effort`, while an unknown ladder remains +unclassified. This guard does not alter the separate raw service-tier contract. On the response side, the upstream `service_tier` echo (xAI Priority Processing, OpenAI fast tier) relays to the Chat Completions caller on every delivery shape: the non-streaming body @@ -118,6 +121,13 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). + +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +The provider summary default applies at Responses ingress; native Chat and Anthropic inbound preferences keep their existing handling. Raw content is never renamed to a summary. See [bridge contract](../providers/chat-compat.md). + ## Claude affinity at final Go dispatch `src/server/claude-messages.ts` carries validated conversation affinity privately through diff --git a/structure/gui-and-management-api.md b/structure/gui-and-management-api.md index 9fde77fc9e..dc2eccd5bf 100644 --- a/structure/gui-and-management-api.md +++ b/structure/gui-and-management-api.md @@ -531,6 +531,11 @@ expiry, including a deadline crossed before effects run, rechecks activation and refreshes quota with Combo data while preserving drafts. Each successful quota snapshot also advances the observation clock, so a retained older row cannot defer evaluation of a fresh row. +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](transports/responses.md). + +The provider editor field policy exposes `showThinkingSummary` as a boolean provider option; it controls Responses summary defaults without a dashboard rendering change. See [Google provider](providers/google.md). + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. @@ -544,3 +549,5 @@ integration IO adapter. Its snapshot fingerprint cannot be checked against provi The existing dashboard file-client maps include Cline CLI and reuse its committed color mark. The export panel labels its download as a settings/catalog bundle; all locales explain that Undo restores both original files. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/ops/docs-and-release.md b/structure/ops/docs-and-release.md index 9570086e21..d795b00063 100644 --- a/structure/ops/docs-and-release.md +++ b/structure/ops/docs-and-release.md @@ -309,6 +309,13 @@ The Combo guides describe the distinction between display quota and single-crede The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). + +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Provider configuration documents distinguish actual summaries from raw reasoning content. The test layout registers the summary-default contract cases and removes the obsolete content-rewrite test with its implementation. + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](../codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. @@ -318,3 +325,5 @@ The integrations guide documents Cline CLI as a two-file, loopback-only integrat The lightweight top-level CLI help counts Cline CLI among the fifteen registered export clients; registry parity remains covered by the client help and integration tests. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/ops/service-and-sidecars.md b/structure/ops/service-and-sidecars.md index c8ef069110..f142505ce8 100644 --- a/structure/ops/service-and-sidecars.md +++ b/structure/ops/service-and-sidecars.md @@ -140,6 +140,13 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Provider summary defaults are evaluated per routed Responses request without changing service lifecycle or sidecar activation. See [runtime](../runtime.md). + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/overview.md b/structure/overview.md index d5d1a2207a..68a2a29de3 100644 --- a/structure/overview.md +++ b/structure/overview.md @@ -107,4 +107,9 @@ would pass while the rule was violated. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](transports/responses.md). + +Raw reasoning content and provider-authored summaries remain distinct on the Responses wire. See [reasoning presentation](providers/chat-compat.md). + Cline CLI is a managed file integration: its provider settings and catalog share one recoverable journal operation. The [paired-file contract](clients/integrations.md#cline-paired-files) defines its stop/restart requirement. diff --git a/structure/providers/chat-compat.md b/structure/providers/chat-compat.md index 7cf84ef7ec..b6a3573fc5 100644 --- a/structure/providers/chat-compat.md +++ b/structure/providers/chat-compat.md @@ -22,7 +22,9 @@ content converter after validation and imports no optional subsystem. Stateful developer-guidance injection reuses that validator for its raw insertion boundary, so parsed messages and stored raw history retain the same task/guidance order. -Native OpenAI passthrough sanitizes routed reasoning history so `reasoning` input items do not send +Native OpenAI passthrough consults the existing configured capability ladder before forwarding +`reasoning_effort`; an explicitly empty ladder removes that unsupported control while an unknown +ladder remains unclassified. It also sanitizes routed reasoning history so `reasoning` input items do not send non-empty `content` arrays to upstream models that reject them. Chat Completions bridging repairs orphan `toolResult` messages by inserting a synthetic assistant `tool_call` before tool messages. It also repairs the opposite direction (260718): an assistant `tool_calls` round left dangling — @@ -220,18 +222,13 @@ honored by BOTH reasoning paths: anthropic `thinking_delta` AND raw `reasoning_r item (`summary: []`, txt-only `ocxr1:` `encrypted_content`, no text deltas) — invisible in the Codex app, so tool cells group like native models — while the text still round-trips for `preserveReasoningContentModels` replay. Visible mode (summary "auto") keeps the raw -`content[reasoning_text]` shape. Diagnosis and codex-rs grouping evidence: -`devlog/_fin/260709_native_response_pattern/`. - -The content-to-summary channel rewrite skips any reasoning item that carries a native -`encrypted_content` blob. The blob is opaque, state-bearing provider data, so the item must -round-trip unchanged unless that backend has an explicit replay contract permitting a rewrite. -This defensively protects providers that issue blobs and later join the route through -`preserveReasoningContentModels`. The rewrite's round trip was verified against DeepSeek, which is -`statelessResponses` and issues no blob. Grok is unaffected in practice because it natively emits -summary-channel reasoning and no `reasoning_text` events, so this content-to-summary item rewrite -does not engage on its route. Only the stored item is exempt — `reasoning_text` delta events carry -no blob and still route to the summary channel, so the live expandable trace is unchanged. +`content[reasoning_text]` shape: raw deltas stream as `response.reasoning_text.delta` and the final +item carries `content: [{type: "reasoning_text", text}]`, so Codex applies its own display policy — +the desktop thinking band shows the "Thinking…" placeholder, and raw text appears only when +`show_raw_agent_reasoning` is enabled. Routing raw CoT through the summary channel instead (the +#45 display intent, intentionally reverted 260911) put unsummarized thinking in the desktop band, +which only fits native OpenAI providers that author real summaries. Diagnosis and codex-rs +grouping evidence: `devlog/_fin/260709_native_response_pattern/`. The process-local raw-reasoning fallback is fail-closed unless a request has an explicit client thread plus an exact provider destination, wire adapter, final model, and physical credential @@ -269,3 +266,5 @@ fragments are not guessed onto pending ID-only calls. parallel/colliding identities, distinct unsafe raw JSON index literals, the maximum safe-integer boundary, invalid index types, missing/null continuations and UTF-8 byte-limit boundaries. + +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). diff --git a/structure/providers/cursor.md b/structure/providers/cursor.md index be42793e0b..a3a4f2f144 100644 --- a/structure/providers/cursor.md +++ b/structure/providers/cursor.md @@ -82,3 +82,9 @@ constraints cannot widen the canonical shape. Bare shell bridge names are reject on the freeform path. Namespaced tools do not acquire bare-shell behavior. Regression coverage lives in `tests/providers/cursor/cursor-tool-definitions.test.ts`. + +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). + +Shared raw-reasoning events retain content-channel presentation; provider-authored thinking keeps its existing summary path. See [bridge contract](chat-compat.md). + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/providers/google.md b/structure/providers/google.md index 26187aeebf..1c1d5648f8 100644 --- a/structure/providers/google.md +++ b/structure/providers/google.md @@ -2,12 +2,14 @@ ## Google thought-text visibility boundary -Google-family responses may represent model-internal reasoning as a text-bearing part with -`thought: true`. The Google adapter maps that text to the internal `reasoning_raw_delta` event; -only text without the marker becomes visible `text_delta`. Streaming SSE and buffered JSON share -one classifier so transport selection cannot change whether provider-declared reasoning is shown -as assistant output. Thought-signature observation still runs on the original parts before text -classification, preserving the opaque continuation state independently of display semantics. +Google-family parts with `thought: true` stay separate from assistant output. After a CCA +Gemini request is built, the shared streaming/buffered classifier emits `thinking_delta` for +these provider-authored summaries. Other Google wires, non-Gemini CCA models and uninitialized +adapters retain `reasoning_raw_delta`. Model provenance is refreshed on every build. +`showThinkingSummary` defaults on only for the Antigravity preset; explicit provider false and +explicit wire summary none win. Eligible CCA Gemini requests use `includeThoughts: true` only +when provider opt-in and per-request display both allow it. Thought signatures remain attached +to their tool calls independently; they never become Anthropic thinking signatures. > Decision record: [ADR-0055](../decisions/ADR-0055-google-thought-text-visibility-boundary.md) diff --git a/structure/providers/openai-tiers.md b/structure/providers/openai-tiers.md index 50c7d28ce3..384426909e 100644 --- a/structure/providers/openai-tiers.md +++ b/structure/providers/openai-tiers.md @@ -417,6 +417,9 @@ model settings, and noncanonical `openai` rows never receive that recovery path. terminal 403 codes as `needsReauth`; generic permission failures remain non-terminal, and a successful main usage refresh clears the runtime mark. +Canonical forwarding alone can apply the optional client-output safety-buffering hint filter; +API-key and custom forward destinations preserve their metadata. See [Responses transport](../transports/responses.md). + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](../codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. diff --git a/structure/providers/xai-grok.md b/structure/providers/xai-grok.md index fc6cf7c63f..2bac4b5ccc 100644 --- a/structure/providers/xai-grok.md +++ b/structure/providers/xai-grok.md @@ -63,12 +63,20 @@ Account-scoped OAuth quota remains display evidence for provider-level Combo sel The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Grok chat raw reasoning uses content-channel output with an empty summary; hidden replay envelopes retain continuation text. Native Responses content is not promoted to summaries. See [chat compatibility](chat-compat.md). + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Devin CLI credential path composition in `src/oauth/devin-cli.ts` follows the selected platform: Windows uses Win32 APPDATA paths, other platforms use POSIX XDG-data paths. The explicit absolute override remains verbatim; credential parsing and login behavior are unchanged. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. + ### OAuth Fast Tier (Priority Processing) xAI's Priority Processing (`service_tier: "priority"` on Chat Completions and Responses, diff --git a/structure/runtime.md b/structure/runtime.md index 7a0231f097..997d3ea7db 100644 --- a/structure/runtime.md +++ b/structure/runtime.md @@ -146,6 +146,7 @@ The server exposes `POST /api/stop` which restores native Codex config, stops an | `src/providers/registry.ts` | Canonical provider presets for CLI, dashboard, OAuth, key providers, and metadata. | | `src/providers/derive.ts` | Enrichment from provider presets into user config. | | `src/oauth/` | OAuth providers, token storage, refresh, and auth-token resolution. The login callback listener binds a per-provider FIXED loopback port, so consecutive logins reuse the same number; every response it sends ends its connection (`Connection: close`, including non-callback paths such as a stray `/favicon.ico` 404). Stopping the listener does not close an established socket, so without that a pooled client would deliver the next login's callback to the retired flow, which rejects the unknown state as a CSRF mismatch while the live flow waits. | +| `src/combos/request.ts` | Clones each selected combo target request and applies the existing target capability ladder: adaptive unknown targets and explicit empty ladders receive no unsupported reasoning/thinking controls, while known ladders retain per-target resolution. | | `src/adapters/openai-responses.ts` | Native OpenAI/ChatGPT Responses passthrough. | | `src/responses/muse-tool-name-alias.ts` | Host-gated Meta Muse 64-char tool-name alias/restore used by the Responses passthrough. | | `src/adapters/openai-chat.ts` | OpenAI-compatible Chat Completions bridge. Its client delivery shapes in `src/chat/outbound.ts` and `src/server/chat-native-sse.ts` relay the upstream `service_tier` echo on non-stream, folded-stream, and synthesized-SSE bodies, never inventing the key when the upstream omits it. | @@ -220,6 +221,13 @@ cooldowns and response-driven retry remain authoritative. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](gui-and-management-api.md#combo-editor-routing-quota). +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](transports/responses.md). + +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](transports/responses.md). + +Responses route normalization resolves provider summary defaults from the original wire preference on every final route. See [reasoning presentation](providers/chat-compat.md) and [CCA summary provenance](providers/google.md). + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. diff --git a/structure/subagents.md b/structure/subagents.md index 24d37af883..f17bd5fb41 100644 --- a/structure/subagents.md +++ b/structure/subagents.md @@ -208,6 +208,11 @@ Provider-level Combo eligibility uses explicit inference evidence for the curren The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](transports/responses.md). + +Final-route summary visibility is recomputed after fallback from the original Responses preference; an earlier provider opt-in does not carry into a later provider. See [reasoning presentation](providers/chat-compat.md). + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. @@ -215,3 +220,5 @@ see [Combo editor routing quota](gui-and-management-api.md#combo-editor-routing- Claude replay carries [Go conversation affinity](data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/transports/inventory.md b/structure/transports/inventory.md index 623abe10a8..145bc7d86b 100644 --- a/structure/transports/inventory.md +++ b/structure/transports/inventory.md @@ -15,7 +15,7 @@ surface is listed here so a maintainer can find the owner without grepping: | Adapter execution support | `src/adapters/run-turn-queue.ts`, `src/adapters/tool-catalog-nudge.ts`, `src/adapters/identity.ts`, `src/adapters/image.ts`, `src/adapters/upstream-http-error.ts` | Shared machinery: turn ordering, tool-catalog nudging, client fingerprinting, image conversion, upstream error normalization. | | Cursor (beyond the sections above) | `src/adapters/cursor/live-transport.ts`, `src/adapters/cursor/http1-bidi.ts`, `src/adapters/cursor/live-models.ts`, `src/adapters/cursor/transport-retry.ts`, `src/adapters/cursor/mcp-manager.ts`, `src/adapters/cursor/thread-continuity.ts`, `src/adapters/cursor/checkpoint-store.ts` | Thread continuity is the point: a retry must not start a new Cursor thread, and a validated checkpoint must not rebuild the full root history. HTTP/2 remains the default; an explicit `http1.1`/`h1` pin maps the bidi run onto Cursor's `RunSSE` receive stream plus sequenced `BidiAppend` sends, and applies to live discovery too. | | Claude Messages | `src/server/claude-messages.ts` | Routed translation, a native Anthropic passthrough branch, and `count_tokens`. | -| Chat Completions inbound | `src/server/chat-completions.ts`, `src/chat/` | Inbound translation onto the same routing pipeline. The content mapper preserves image URLs and supported detail, including screenshot-bearing tool results; target adapters own image placement on their wire. Image-free tool results stay strings. On the response side, the upstream `service_tier` echo relays on every delivery shape (`src/chat/outbound.ts` projections, `src/server/chat-native-sse.ts` chunks); an upstream without the field gets no injected key. | +| Chat Completions inbound | `src/server/chat-completions.ts`, `src/server/chat-native.ts`, `src/chat/`, `src/adapters/openai-chat.ts` | Inbound translation onto the same routing pipeline. The content mapper preserves image URLs and supported detail, including screenshot-bearing tool results; target adapters own image placement on their wire. Image-free tool results stay strings. The native handler owns pin/cap normalization; the adapter wire builder removes effort only for explicit empty declarations or no-reasoning models, preserving unknown raw declarations. On the response side, the upstream `service_tier` echo relays on every delivery shape (`src/chat/outbound.ts` projections, `src/server/chat-native-sse.ts` chunks); an upstream without the field gets no injected key. | | Hosted search relay | `src/server/search.ts` | Direct relay; distinct from the web-search sidecar loop below. | | Image/video generation loop | `src/images/loop.ts`, `src/images/plan.ts`, `src/images/fulfill.ts`, `src/images/xai-client.ts`, `src/images/xai-video-client.ts`, `src/images/artifacts.ts` | A provider-returned image URL is downloaded into a local artifact once, then served locally; warnings stay URL-free because provider CDN URLs may embed credentials. | | GitHub Copilot | `src/providers/xai-transport.ts` (`resolveProviderTransport`), `src/providers/github-copilot-transport.ts` | `resolveProviderTransport` selects the Copilot transport when the routed provider name is `github-copilot`; the Copilot module then resolves its headers and base URL, and the registry seeds the provider row and model fallback. | @@ -69,6 +69,13 @@ Quota publication distinguishes display reports from explicitly supplied inferen The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). + +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +CCA Gemini summary provenance and request opt-in are specified in [Google provider](../providers/google.md); raw Responses content retains its wire channel. + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. diff --git a/structure/transports/responses.md b/structure/transports/responses.md index 9b2c0f883e..92de0b039a 100644 --- a/structure/transports/responses.md +++ b/structure/transports/responses.md @@ -316,14 +316,11 @@ custom result has no local call, because its original wire type cannot be establ would send an unmatched result upstream. The check resolves the selected wire protocol and the request's own tool declarations after final route selection, so stateful destinations keep their upstream-owned native function and native-only custom continuations. Explicit input still receives -orphan repair; this path asks the client to replay rather than reconstructing history. This flag also enables the existing -visible content-to-summary rewrite for SSE and JSON; summary-channel items and opaque reasoning -blobs keep their existing response handling. The shared recording callback applies the same -reasoning rewrite under the exact client-visible predicate before caching output, after tool -restoration and function normalization. This keeps full-content replay fingerprints comparable -for both full-history-plus-ID and delta continuations without weakening identity checks. Hidden -summaries and opaque blobs keep their existing cache representation. It does not change streaming selection or Chat -model routes. Go fixtures cover Luna, Grok and Muse against both response formats. +orphan repair; this path asks the client to replay rather than reconstructing history. Content-channel reasoning stays content in SSE, JSON and stored replay output; native +summary items and opaque blobs retain their upstream representation. Full-content replay +fingerprints compare the same client-visible items without content-to-summary conversion. +It does not change streaming selection or Chat model routes. Go fixtures cover Luna, Grok +and Muse against both response formats. The canonical OpenCode Go transport also derives `x-opencode-session` from the existing hashed session lane before per-model wire selection. One conversation keeps one opaque affinity value @@ -424,6 +421,16 @@ final outgoing model/tier. No caller identity is synthesized. Noncanonical opt-in gateways keep their own metadata policy. Oversized/unsupported-runtime HTTP fallback preserves the original HTTP body and Lite header. +For the final wire model `gpt-5.3-codex-spark`, the canonical forward adapter normalizes the +Lite header from the BODY, overriding caller/configured headers and stale native WS Lite +metadata in both directions. A body carrying a nonempty `additional_tools` input item is +pinned to `true`: that item IS the Lite tool-delivery format and the non-Lite wire shape +expects top-level `tools`, so an inherited `false` would advertise non-Lite while the tools +exist only in the Lite shape and hide the client tool surface. Any other Spark body is set to +`false`, selecting the non-Lite framing policy. A changed Lite identity retires the previous socket; +subsequent eligible Spark requests with the same identity can reuse the new socket. Malformed +native metadata retains HTTP fallback eligibility without rewriting its body. + Canonical WS quota and response metadata preceding the first Responses event are projected into bounded, allowlisted HTTP headers before the response is committed. Later quota observations update only the captured serving account; @@ -499,6 +506,10 @@ retried. Guarded paths: the ChatGPT passthrough and generic adapter fetch in fallback. Adapters with their own `fetchResponse` (kiro, cursor, google) keep their own retry policies; kiro imports the shared abort/sleep helpers from this module. +## Console upload rejection recovery + +`src/providers/opencode-zen-rate-limit.ts` recognizes the complete Console upload-rejection envelope only at the effective HTTPS opencode.ai Zen/Go generation endpoint. A provider row name cannot authorize another destination. The two recovery loops in `src/server/responses/core.ts` wait 800 ms and replay the captured serialized request once; cancellation, nonreplayable responses, other errors and a second upload rejection keep their failure semantics. The recovery kind is persisted as `console-go-upload-retry` and has a localized Logs label. + ## Same-provider combo quota fallback For a failover combo with multiple models on the same Codex-login OpenAI provider, a pre-stream @@ -511,6 +522,15 @@ combo whose remaining eligible targets use other providers. > Decision record: [ADR-0070](../decisions/ADR-0070-same-provider-combo-quota-fallback.md) +## Combo per-target reasoning controls + +`src/server/responses/core.ts` passes the combo's `reasoningEffortMode` and the final target's +`supportedLadderFor` result to `src/combos/request.ts` before adapter parsing. Explicit empty +capability ladders remove effort and thinking controls in every combo mode; adaptive mode also +removes those controls for unknown ladders and preserves `reasoning.summary`. Known non-empty +ladders retain the existing per-target effort resolution. This request normalization does not +change target order, attempt accounting, or the existing provider-400 failover classification. + ## Combo streaming commit boundary An HTTP 200 does not by itself commit a streaming combo child. The combo parent runs the child's @@ -536,6 +556,8 @@ not retried. > Decision record: [ADR-0071](../decisions/ADR-0071-combo-streaming-commit-boundary.md) +Console upload-rejection recovery excludes query-bearing and fragment-bearing destinations even when their host and generation path match the canonical endpoint. + Chat helper admission in `src/server/responses/core.ts` follows the [deferred stored-main contract](../providers/openai-tiers.md): only a needed Direct OpenAI helper claims stored main, after terminal vision, routed vision and search exclusions. @@ -543,6 +565,17 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Spark Lite and routing metadata use the same suffix-normalized model object as serialization, including configured bracket-suffix removal. + +## Optional client transport hints + +`dropCodexSafetyBuffering` defaults to false. Canonical OpenAI forward Responses can remove only +the two safety-buffering response headers, matching response.metadata events and top-level +safety_buffering fields. Pull/eager client output boundaries compose this with policy failure +normalization; refusal/error semantics, retryability, cancellation and captured EOF errors remain +intact. Internal inspection observes original upstream frames. Native codex.response.metadata.headers +WebSocket metadata and compact are excluded. This does not disable upstream safety enforcement. + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. diff --git a/structure/transports/streaming-health.md b/structure/transports/streaming-health.md index 6a5574c6d2..c6627f47bd 100644 --- a/structure/transports/streaming-health.md +++ b/structure/transports/streaming-health.md @@ -197,6 +197,13 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Raw reasoning and provider-authored summary deltas both remain real upstream activity; visibility does not change heartbeat or terminal ownership. See [reasoning presentation](../providers/chat-compat.md). + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/tests/adapters/bridge-raw-reasoning-hidden.test.ts b/tests/adapters/bridge-raw-reasoning-hidden.test.ts index 4acc66ce62..162d98b206 100644 --- a/tests/adapters/bridge-raw-reasoning-hidden.test.ts +++ b/tests/adapters/bridge-raw-reasoning-hidden.test.ts @@ -77,18 +77,19 @@ describe("hidden raw reasoning (hideThinkingSummary parity for reasoning_raw_del expect(fc).toMatchObject({ call_id: "call_1", name: "read_file" }); }); - test("streamed visible (flag off): raw reasoning rides the expandable summary channel (#2007)", async () => { + test("streamed visible (flag off): raw reasoning rides the content channel (#2007)", async () => { const frames = await collectSse(bridgeToResponsesSSE(replay([ { type: "reasoning_raw_delta", text: "visible raw" }, { type: "done" }, ]), "routed/model")); - expect(frames.some(f => f.event === "response.reasoning_summary_text.delta")).toBe(true); - expect(frames.some(f => f.event === "response.reasoning_text.delta")).toBe(false); + expect(frames.some(f => f.event === "response.reasoning_text.delta")).toBe(true); + expect(frames.some(f => f.event === "response.reasoning_summary_text.delta")).toBe(false); const completed = frames.find(f => f.event === "response.completed")?.data.response as Record; const output = completed.output as Record[]; expect(output[0]).toMatchObject({ type: "reasoning", - summary: [{ type: "summary_text", text: "visible raw" }], + summary: [], + content: [{ type: "reasoning_text", text: "visible raw" }], }); }); @@ -118,14 +119,15 @@ describe("hidden raw reasoning (hideThinkingSummary parity for reasoning_raw_del expect(decodeReasoningEnvelope(reasoning.encrypted_content as string)?.txt).toBe("quiet"); }); - test("non-streaming visible: raw reasoning lands in the summary channel (#2007)", () => { + test("non-streaming visible: raw reasoning lands on the content channel (#2007)", () => { const json = buildResponseJSON([ { type: "reasoning_raw_delta", text: "loud" }, { type: "done" }, ], "routed/model", {}); const output = (json as { output: Record[] }).output; expect(output.find(o => o.type === "reasoning")).toMatchObject({ - summary: [{ type: "summary_text", text: "loud" }], + summary: [], + content: [{ type: "reasoning_text", text: "loud" }], }); }); diff --git a/tests/adapters/bridge.test.ts b/tests/adapters/bridge.test.ts index e9c2b050b6..f283b77015 100644 --- a/tests/adapters/bridge.test.ts +++ b/tests/adapters/bridge.test.ts @@ -87,27 +87,26 @@ describe("Responses bridge reasoning and usage parity", () => { expect(firstOutputs).toBe(1); }); - test("streaming raw reasoning is routed through the expandable summary channel", async () => { + test("streaming raw reasoning rides the content channel like native gpt-oss", async () => { const frames = await collectSse(bridgeToResponsesSSE(replay([ { type: "reasoning_raw_delta", text: "raw detail" }, { type: "done", usage: { inputTokens: 10, outputTokens: 5, cachedInputTokens: 3, reasoningOutputTokens: 2 } }, ]), "routed/model")); - // Chat-completions providers (DeepSeek-style) deliver thinking as raw - // reasoning_content. Codex renders the expandable reasoning trace from the - // Responses summary channel only, so raw reasoning is routed through the - // summary channel (issue #45) instead of the content channel. - expect(frames.find(f => f.event === "response.reasoning_summary_text.delta")?.data) - .toMatchObject({ summary_index: 0, delta: "raw detail" }); - expect(frames.some(f => f.event === "response.reasoning_text.delta")).toBe(false); + // Raw reasoning_content rides the content channel so Codex applies its own display + // policy: the desktop band shows the "Thinking…" placeholder, and raw text appears + // only when show_raw_agent_reasoning is enabled — never as a fake summary. + expect(frames.find(f => f.event === "response.reasoning_text.delta")?.data) + .toMatchObject({ content_index: 0, delta: "raw detail" }); + expect(frames.some(f => f.event === "response.reasoning_summary_text.delta")).toBe(false); const completed = frames.find(f => f.event === "response.completed")?.data.response as Record; const output = completed.output as Record[]; expect(output[0]).toMatchObject({ type: "reasoning", - summary: [{ type: "summary_text", text: "raw detail" }], + summary: [], + content: [{ type: "reasoning_text", text: "raw detail" }], }); - expect((output[0] as { content?: unknown }).content).toBeUndefined(); expect(completed.usage).toMatchObject({ input_tokens: 10, input_tokens_details: { cached_tokens: 3 }, @@ -502,9 +501,9 @@ describe("Responses bridge reasoning and usage parity", () => { const output = json.output as Record[]; expect(output.map(item => item.type)).toEqual(["reasoning", "message"]); expect(output[0]).toMatchObject({ - summary: [{ type: "summary_text", text: "raw json" }], + summary: [], + content: [{ type: "reasoning_text", text: "raw json" }], }); - expect((output[0] as { content?: unknown }).content).toBeUndefined(); expect(json.usage).toMatchObject({ input_tokens: 6, input_tokens_details: { cached_tokens: 1, cache_write_tokens: 2 }, diff --git a/tests/adapters/google/google-adapter.test.ts b/tests/adapters/google/google-adapter.test.ts index 61e1d15d4b..1dd5a7c97a 100644 --- a/tests/adapters/google/google-adapter.test.ts +++ b/tests/adapters/google/google-adapter.test.ts @@ -1,8 +1,10 @@ import { describe, expect, test } from "bun:test"; import { createGoogleAdapter } from "../../../src/adapters/google"; import { chatCompletionsToResponsesBody } from "../../../src/chat/inbound"; +import { bridgeToResponsesSSE, buildResponseJSON } from "../../../src/bridge"; +import { withTestTranslatorBudget } from "../../helpers/translator-budget"; import { parseRequest } from "../../../src/responses/parser"; -import type { OcxParsedRequest } from "../../../src/types"; +import type { AdapterEvent, OcxParsedRequest } from "../../../src/types"; const provider = { adapter: "google", baseUrl: "https://generativelanguage.googleapis.com", apiKey: "key" }; @@ -583,3 +585,134 @@ describe("google adapter — direct -tiered wire renames", () => { } }); }); + +describe("google adapter — Antigravity thought-text opt-in", () => { + // CCA keeps generating thinking either way (thoughtsTokenCount stays non-zero) but returns + // NO `thought` text unless the request sets generationConfig.thinkingConfig.includeThoughts. + // Probed 2026-09-12: gemini-3.8-flash-high answered with 0 thought parts and 321 thoughts + // tokens, then 358-652 chars of reasoning once the key was present. + const ccaProvider = { + adapter: "google", + googleMode: "cloud-code-assist", + baseUrl: "https://daily-cloudcode-pa.googleapis.com", + apiKey: "key", + project: "proj-123", + } as const; + const optedIn = { ...ccaProvider, showThinkingSummary: true } as const; + + function thoughtParsed(modelId: string, effort?: string, hideThinkingSummary?: boolean): OcxParsedRequest { + return { + modelId, + stream: false, + options: { ...(effort ? { reasoning: effort } : {}), ...(hideThinkingSummary ? { hideThinkingSummary } : {}) }, + context: { messages: [{ role: "user", content: "hi" }], tools: [] }, + } as unknown as OcxParsedRequest; + } + + async function thinkingConfig( + providerConfig: Record, + modelId: string, + effort?: string, + hideThinkingSummary?: boolean, + ): Promise | undefined> { + const { body } = await createGoogleAdapter(providerConfig as never) + .buildRequest(thoughtParsed(modelId, effort, hideThinkingSummary)); + const envelope = JSON.parse(body) as { + request: { generationConfig?: { thinkingConfig?: Record } }; + }; + return envelope.request.generationConfig?.thinkingConfig; + } + + test("asks CCA for thought text on the Gemini wire families", async () => { + // Suffix tier ids deliberately state no level — the suffix IS the effort — so the opt-in + // has to stand on its own for those. + expect(await thinkingConfig(optedIn, "gemini-3.8-flash", "high")).toEqual({ includeThoughts: true }); + expect(await thinkingConfig(optedIn, "gemini-3.8-flash-medium")).toEqual({ includeThoughts: true }); + expect(await thinkingConfig(optedIn, "gemini-3.7-flash", "high")) + .toEqual({ thinkingLevel: "high", includeThoughts: true }); + expect(await thinkingConfig(optedIn, "gemini-3.1-pro", "high")) + .toEqual({ thinkingLevel: "high", includeThoughts: true }); + }); + + test("never sends the flag to models that reject or ignore it", async () => { + // gpt-oss answers 400 INVALID_ARGUMENT with the key present, so it would break the turn. + expect(await thinkingConfig(optedIn, "gpt-oss-120b-medium")).toBeUndefined(); + // Claude-on-CCA accepts the key but returns no thought parts, so it stays off that wire. + expect(await thinkingConfig(optedIn, "claude-sonnet-4-6", "high")).toEqual({ thinkingLevel: "high" }); + }); + + test("a provider without the opt-in keeps the CCA wire unchanged", async () => { + expect(await thinkingConfig(ccaProvider, "gemini-3.8-flash", "high")).toBeUndefined(); + expect(await thinkingConfig(ccaProvider, "gemini-3.7-flash", "high")).toEqual({ thinkingLevel: "high" }); + }); + + test("an explicit client opt-out stops the thought text at the source", async () => { + // Same per-request gate the response path uses: hideThinkingSummary is set for an explicit + // reasoning.summary "none", and paying upstream for text the client refused is waste. + expect(await thinkingConfig(optedIn, "gemini-3.8-flash", "high", true)).toBeUndefined(); + expect(await thinkingConfig(optedIn, "gemini-3.7-flash", "high", true)).toEqual({ thinkingLevel: "high" }); + }); +}); + + +describe("CCA thought summary provenance and replay", () => { + const cca = { adapter: "google", googleMode: "cloud-code-assist", baseUrl: "https://daily-cloudcode-pa.googleapis.com", + apiKey: "fixture-key", project: "fixture-project", showThinkingSummary: true } as const; + const signature = "CiQAx-summary-tool-signature-0123456789abcdef"; + for (const stream of [false, true]) for (const hideThinkingSummary of [false, true]) test(`Gemini signature stream=${stream} hidden=${hideThinkingSummary}`, async () => { + const adapter = withTestTranslatorBudget(createGoogleAdapter(cca)); + const parsed = parsedWith([{ role: "user", content: "lookup" }], [ + { name: "lookup", description: "look up", parameters: { type: "object", properties: {} } }, + ]); + parsed.modelId = "gemini-3.8-flash"; + await adapter.buildRequest(parsed); + const payload = { response: { candidates: [{ content: { parts: [ + { thought: true, text: "Provider summary", thoughtSignature: signature }, + { functionCall: { name: "lookup", args: {} } }, + ] }, finishReason: "STOP" }], usageMetadata: { promptTokenCount: 5, candidatesTokenCount: 2 } } }; + const events: AdapterEvent[] = []; + if (stream) { + for await (const event of adapter.parseStream(new Response(`data: ${JSON.stringify(payload)}\n\n`, + { headers: { "content-type": "text/event-stream" } }))) events.push(event); + } else events.push(...await adapter.parseResponse!(Response.json(payload))); + expect(events[0]).toEqual({ type: "thinking_delta", thinking: "Provider summary" }); + expect(events.some(event => event.type === "thinking_signature")).toBe(false); + const call = events.find(event => event.type === "tool_call_start"); + expect(call?.type === "tool_call_start" && call.providerMetadata?.google?.thoughtSignature).toBe(signature); + expect(events.at(-1)?.type).toBe("done"); + let output: Record; + if (stream) { + async function* replay() { yield* events; } + const text = await new Response(bridgeToResponsesSSE(replay(), parsed.modelId, + undefined, undefined, undefined, undefined, undefined, { hideThinkingSummary })).text(); + const payloads = text.split("\n").filter(line => line.startsWith("data: {")).map(line => JSON.parse(line.slice(6))); + output = payloads.find(frame => frame.type === "response.completed").response; + expect(text.includes("response.reasoning_summary_text.delta")).toBe(!hideThinkingSummary); + } else output = buildResponseJSON(events, parsed.modelId, { hideThinkingSummary }); + expect(JSON.stringify(output).includes("Provider summary")).toBe(!hideThinkingSummary); + if (!Array.isArray(output.output)) throw new Error("missing Responses output"); + const continuation = parseRequest({ model: parsed.modelId, input: [ + ...output.output, { type: "function_call_output", call_id: call && "id" in call ? call.id : "", output: "result" }, + ] }); + const next = JSON.parse((await withTestTranslatorBudget(createGoogleAdapter(cca)).buildRequest(continuation)).body); + const parts = next.request.contents.flatMap((turn: { parts: unknown[] }) => turn.parts); + expect(parts).toContainEqual(expect.objectContaining({ functionCall: expect.objectContaining({ name: "lookup" }), thoughtSignature: signature })); + expect(parts).toContainEqual(expect.objectContaining({ functionResponse: expect.objectContaining({ name: "lookup", response: { result: "result" } }) })); + expect(JSON.stringify(next)).not.toContain("no tool result"); + }); + + test("reused adapter resets Gemini summary provenance for a CCA non-Gemini model", async () => { + const adapter = withTestTranslatorBudget(createGoogleAdapter(cca)); + for (const modelId of ["gemini-3.8-flash", "gpt-oss-120b-medium"]) { + const request = parsedWith([{ role: "user", content: "hi" }]); + request.modelId = modelId; + await adapter.buildRequest(request); + const events = await adapter.parseResponse!(Response.json({ response: { candidates: [{ + content: { parts: [{ thought: true, text: "thinking" }] }, finishReason: "STOP", + }] } })); + expect(events[0]).toEqual(modelId.startsWith("gemini-") + ? { type: "thinking_delta", thinking: "thinking" } + : { type: "reasoning_raw_delta", text: "thinking" }); + } + }); +}); diff --git a/tests/adapters/google/google-wire-compiler.test.ts b/tests/adapters/google/google-wire-compiler.test.ts index 1523d11d85..482aa809e6 100644 --- a/tests/adapters/google/google-wire-compiler.test.ts +++ b/tests/adapters/google/google-wire-compiler.test.ts @@ -132,4 +132,28 @@ describe("Google wire compiler", () => { const repaired = JSON.parse(repairGoogleInvalidRequestBody(body, error)!); expect(repaired.request.generationConfig).toEqual({ maxOutputTokens: 4096 }); }); + + test("keeps the includeThoughts opt-in while still dropping unknown thinking keys", () => { + const withOptIn = compileGoogleWireBody({ + generationConfig: { + thinkingConfig: { includeThoughts: true, thinkingLevel: "max", futureThinkingField: true }, + }, + }); + expect(withOptIn.body.generationConfig).toEqual({ + thinkingConfig: { thinkingLevel: "high", includeThoughts: true }, + }); + + // The flag has to survive on its own too: suffix tier ids deliberately carry no + // thinkingLevel, so an includeThoughts-only config is the whole request. + const optInOnly = compileGoogleWireBody({ + generationConfig: { thinkingConfig: { includeThoughts: true } }, + }); + expect(optInOnly.body.generationConfig).toEqual({ thinkingConfig: { includeThoughts: true } }); + + // Non-boolean / absent values must not invent the key. + const notRequested = compileGoogleWireBody({ + generationConfig: { thinkingConfig: { includeThoughts: "yes", thinkingLevel: "high" } }, + }); + expect(notRequested.body.generationConfig).toEqual({ thinkingConfig: { thinkingLevel: "high" } }); + }); }); diff --git a/tests/adapters/openai/openai-chat-hardening.test.ts b/tests/adapters/openai/openai-chat-hardening.test.ts index b051817469..3a8fe98f03 100644 --- a/tests/adapters/openai/openai-chat-hardening.test.ts +++ b/tests/adapters/openai/openai-chat-hardening.test.ts @@ -140,6 +140,23 @@ describe("AgentRouter openai-chat compatibility", () => { expect(body.messages[0]?.content.map(part => part.text)).toEqual([preamble, "responda somente: OK"]); expect(rawBody.messages[0]?.content).toBe("responda somente: OK"); }); + + test("passthrough chat drops reasoning_effort for an explicitly empty capability ladder", () => { + const rawBody = { + messages: [{ role: "user", content: "hi" }], + reasoning_effort: "xhigh", + }; + const request = buildOpenAIChatPassthroughRequest( + provider({ reasoningEfforts: [] }), + rawBody, + "test-model", + false, + ); + const body = JSON.parse(request.body as string) as Record; + + expect(body).not.toHaveProperty("reasoning_effort"); + expect(rawBody.reasoning_effort).toBe("xhigh"); + }); }); function parsed(): OcxParsedRequest { @@ -1455,3 +1472,25 @@ test("tool-call deltas emit heartbeats so a long buffering phase is not read as expect(visible.at(-1)).toMatchObject({ type: "done" }); }); }); + + +describe("native Chat raw reasoning declarations", () => { + const raw = { messages: [{ role: "user", content: "hello" }], reasoning_effort: "enabled" }; + function wire(overrides: Partial) { + const request = buildOpenAIChatPassthroughRequest({ + adapter: "openai-chat", baseUrl: "https://example.test/v1", ...overrides, + }, raw, "target", false); + return JSON.parse(request.body as string) as Record; + } + test("preserves a nonempty wire-only ladder as unknown", () => { + expect(wire({ reasoningEfforts: ["enabled"] }).reasoning_effort).toBe("enabled"); + expect(wire({ reasoningEfforts: [], modelReasoningEfforts: { target: ["enabled"] } }).reasoning_effort).toBe("enabled"); + }); + test("honors explicit empty model overrides over a provider ladder", () => { + expect(wire({ reasoningEfforts: ["high"], modelReasoningEfforts: { target: [] } }).reasoning_effort).toBeUndefined(); + }); + test("honors noReasoningModels over a nonempty model ladder without mutating input", () => { + expect(wire({ noReasoningModels: ["target"], modelReasoningEfforts: { target: ["high"] } }).reasoning_effort).toBeUndefined(); + expect(raw.reasoning_effort).toBe("enabled"); + }); +}); diff --git a/tests/codex-integration/codex-metadata-integrity.test.ts b/tests/codex-integration/codex-metadata-integrity.test.ts index ee70d4829e..329e90a973 100644 --- a/tests/codex-integration/codex-metadata-integrity.test.ts +++ b/tests/codex-integration/codex-metadata-integrity.test.ts @@ -208,33 +208,115 @@ describe("Codex request transport metadata", () => { expect(new Headers(dropped.headers).get(hintHeader)).toBe("model=gpt-5.6-sol"); }); - test("canonical adapter drops Lite only for the Spark wire model", async () => { + test("canonical adapter disables Spark Lite in HTTP headers and WS metadata without mutating input", async () => { + const { prepareCodexWsRequest } = await import("../../src/server/responses/codex-ws-request"); const adapter = createResponsesPassthroughAdapter({ adapter: "openai-responses", authMode: "forward", baseUrl: "https://chatgpt.com/backend-api/codex", headers: { "X-OpenAI-Internal-Codex-Responses-Lite": "true" }, }); for (const [model, incomingLite, expectedLite] of [ - ["gpt-5.3-codex-spark", "true", null], - ["gpt-5.3-codex-spark", undefined, null], + ["gpt-5.3-codex-spark", "true", "false"], + ["gpt-5.3-codex-spark", "false", "false"], + ["gpt-5.3-codex-spark", undefined, "false"], ["gpt-5.6-sol", "true", "true"], + ["gpt-5.6-sol", "false", "false"], + ["gpt-5.6-sol", undefined, "true"], ] as const) { const parsed = minimalParsed(); parsed.modelId = model; - parsed._rawBody = { model, input: [], stream: true }; + parsed._rawBody = { model, input: [], stream: true, + client_metadata: { [liteKey]: "true", other: "preserved" } }; + const before = JSON.stringify(parsed._rawBody); const incoming = new Headers(); if (incomingLite !== undefined) incoming.set(liteHeader, incomingLite); const request = await adapter.buildRequest(parsed, { headers: incoming, }); expect(new Headers(request.headers).get(liteHeader)).toBe(expectedLite); + const prepared = prepareCodexWsRequest(url, { body: request.body, headers: request.headers })!; + expect(JSON.parse(prepared.frameText).client_metadata).toEqual({ + [liteKey]: expectedLite, other: "preserved", + }); + expect(prepared.httpInit.body).toBe(request.body); + expect(JSON.stringify(parsed._rawBody)).toBe(before); + expect(incoming.get(liteHeader)).toBe(incomingLite ?? null); } const routed = minimalParsed(); routed.modelId = "spark-alias"; routed._rawBody = { model: "gpt-5.3-codex-spark", input: [], stream: true }; const request = await adapter.buildRequest(routed, { headers: new Headers({ [liteHeader]: "true" }) }); - expect(new Headers(request.headers).get(liteHeader)).toBeNull(); + expect(new Headers(request.headers).get(liteHeader)).toBe("false"); + const prepared = prepareCodexWsRequest(url, { body: request.body, headers: request.headers })!; + expect(JSON.parse(prepared.frameText).client_metadata[liteKey]).toBe("false"); + + routed.modelId = "gpt-5.3-codex-spark"; + routed._rawBody = { model: "gpt-5.6-sol", input: [], stream: true }; + const otherWireModel = await adapter.buildRequest(routed, { headers: new Headers({ [liteHeader]: "true" }) }); + expect(new Headers(otherWireModel.headers).get(liteHeader)).toBe("true"); + }); + + test("a Lite-shaped Spark body pins Lite back on, whatever the inherited header said", async () => { + const { prepareCodexWsRequest } = await import("../../src/server/responses/codex-ws-request"); + // The catalog keeps use_responses_lite: true for Spark because it selects tool DELIVERY: + // the client catalog rides `input[].additional_tools`, not top-level `tools`. A forwarded or + // configured `false` must not survive on such a body, or the frame advertises non-Lite while + // the tools exist only in the Lite shape and Spark loses them. + const adapter = createResponsesPassthroughAdapter({ + adapter: "openai-responses", authMode: "forward", baseUrl: "https://chatgpt.com/backend-api/codex", + headers: { "X-OpenAI-Internal-Codex-Responses-Lite": "false" }, + }); + const liteShapedInput = [ + { type: "message", role: "user", content: [{ type: "input_text", text: "hi" }] }, + { type: "additional_tools", tools: [{ type: "function", name: "shell", parameters: {} }] }, + ]; + + for (const incomingLite of ["false", "true", undefined] as const) { + const parsed = minimalParsed(); + parsed.modelId = "gpt-5.3-codex-spark"; + parsed._rawBody = { model: "gpt-5.3-codex-spark", input: liteShapedInput, stream: true, + client_metadata: { [liteKey]: "false", other: "preserved" } }; + const incoming = new Headers(); + if (incomingLite !== undefined) incoming.set(liteHeader, incomingLite); + const request = await adapter.buildRequest(parsed, { headers: incoming }); + expect(new Headers(request.headers).get(liteHeader)).toBe("true"); + const prepared = prepareCodexWsRequest(url, { body: request.body, headers: request.headers })!; + expect(JSON.parse(prepared.frameText).client_metadata).toEqual({ + [liteKey]: "true", other: "preserved", + }); + } + + // An empty group is not a Lite tool surface, so the stream fix still applies. + const toolless = minimalParsed(); + toolless.modelId = "gpt-5.3-codex-spark"; + toolless._rawBody = { model: "gpt-5.3-codex-spark", stream: true, + input: [{ type: "additional_tools", tools: [] }] }; + const downgraded = await adapter.buildRequest(toolless, { headers: new Headers() }); + expect(new Headers(downgraded.headers).get(liteHeader)).toBe("false"); + }); + + test("Spark disables Lite without configured headers and retains malformed-metadata HTTP fallback", async () => { + const { prepareCodexWsRequest } = await import("../../src/server/responses/codex-ws-request"); + const adapter = createResponsesPassthroughAdapter({ + adapter: "openai-responses", authMode: "forward", baseUrl: "https://chatgpt.com/backend-api/codex", + }); + for (const client_metadata of [undefined, {}, null, [], { [liteKey]: true }]) { + const parsed = minimalParsed(); + parsed._rawBody = { model: "gpt-5.3-codex-spark", input: [], stream: true, + ...(client_metadata === undefined ? {} : { client_metadata }) }; + const before = JSON.stringify(parsed._rawBody); + const request = await adapter.buildRequest(parsed, { headers: new Headers() }); + expect(new Headers(request.headers).get(liteHeader)).toBe("false"); + const prepared = prepareCodexWsRequest(url, { body: request.body, headers: request.headers }); + if (client_metadata === undefined || JSON.stringify(client_metadata) === "{}") { + expect(JSON.parse(prepared!.frameText).client_metadata).toEqual({ [liteKey]: "false" }); + } else { + expect(prepared).toBeNull(); + expect(JSON.parse(request.body).client_metadata).toEqual(client_metadata); + } + expect(JSON.stringify(parsed._rawBody)).toBe(before); + } }); test("noncanonical adapters neither forward caller Lite nor synthesize a routing hint", async () => { @@ -243,7 +325,10 @@ describe("Codex request transport metadata", () => { adapter: "openai-responses", authMode, baseUrl: "https://gateway.example/v1", headers: { [hintHeader]: "operator-owned" }, }); - const request = await adapter.buildRequest(minimalParsed(), { + const parsed = minimalParsed(); + parsed.modelId = "gpt-5.3-codex-spark"; + parsed._rawBody = { model: parsed.modelId, input: [] }; + const request = await adapter.buildRequest(parsed, { headers: new Headers({ [liteHeader]: "true", [hintHeader]: "caller-owned" }), }); expect(new Headers(request.headers).has(liteHeader)).toBe(false); @@ -363,3 +448,60 @@ describe("Codex request transport metadata", () => { expect(headers.get(hintHeader)).toBe("model=gpt-5.6-luna"); }); }); + + +describe("Spark Lite follows serialized model and surviving tool shape", () => { + const liteHeader = "x-openai-internal-codex-responses-lite"; + const liteKey = "ws_request_header_x_openai_internal_codex_responses_lite"; + test("bracket normalization and both alias directions use the wire model", async () => { + const adapter = createResponsesPassthroughAdapter({ adapter: "openai-responses", authMode: "forward", + baseUrl: "https://chatgpt.com/backend-api/codex", modelSuffixBracketStrip: true }); + for (const [selector, model, expectedModel, expectedLite] of [ + ["alias", "gpt-5.3-codex-spark[1m]", "gpt-5.3-codex-spark", "false"], + ["gpt-5.3-codex-spark", "gpt-5.6-sol", "gpt-5.6-sol", "true"], + ]) { + const parsed = minimalParsed(); + parsed.modelId = selector; + parsed._rawBody = { model, input: [] }; + const before = JSON.stringify(parsed._rawBody); + const built = await adapter.buildRequest(parsed, { headers: new Headers({ [liteHeader]: "true" }) }); + expect(JSON.parse(built.body).model).toBe(expectedModel); + expect(new Headers(built.headers).get(liteHeader)).toBe(expectedLite); + expect(new Headers(built.headers).get("x-codex-routing-hint")).toContain(`model=${expectedModel}`); + expect(JSON.stringify(parsed._rawBody)).toBe(before); + } + }); + + for (const [name, inputTools, topTools, expectedTools, expectedLite] of [ + ["filtered empty", [{ type: "tool_search" }], undefined, [], "false"], + ["reserved functions", [{ type: "namespace", name: "functions", tools: [{ type: "function", name: "lookup", parameters: { type: "object" } }] }], undefined, + [{ type: "namespace", name: "functions", tools: [{ type: "function", name: "lookup", parameters: { type: "object" } }] }], "true"], + ["top-level only", undefined, [{ type: "function", name: "lookup", parameters: { type: "object" } }], undefined, "false"], + ] as const) test(`post-transform body: ${name}`, async () => { + const { prepareCodexWsRequest } = await import("../../src/server/responses/codex-ws-request"); + const adapter = createResponsesPassthroughAdapter({ adapter: "openai-responses", authMode: "forward", + baseUrl: "https://chatgpt.com/backend-api/codex", headers: { [liteHeader]: expectedLite === "true" ? "false" : "true" } }); + const parsed = minimalParsed(); + parsed.modelId = "gpt-5.3-codex-spark"; + parsed._rawBody = { model: parsed.modelId, input: inputTools ? [{ type: "additional_tools", tools: inputTools }] : [], + ...(topTools ? { tools: topTools } : {}) }; + const built = await adapter.buildRequest(parsed); + const body = JSON.parse(built.body); + expect(body.input.find((item: { type: string }) => item.type === "additional_tools")?.tools).toEqual(expectedTools); + if (topTools) expect(body.tools).toEqual(topTools); + expect(new Headers(built.headers).get(liteHeader)).toBe(expectedLite); + const prepared = prepareCodexWsRequest("https://chatgpt.com/backend-api/codex/responses", { body: built.body, headers: built.headers }); + expect(JSON.parse(prepared!.frameText).client_metadata[liteKey]).toBe(expectedLite); + }); + + test("noncanonical static Lite remains operator-owned", async () => { + for (const authMode of ["key", "forward"] as const) { + const adapter = createResponsesPassthroughAdapter({ adapter: "openai-responses", authMode, + baseUrl: "https://gateway.example/v1", headers: { [liteHeader]: "operator-owned" } }); + const parsed = minimalParsed(); + parsed._rawBody = { model: "gpt-5.3-codex-spark", input: [] }; + const built = await adapter.buildRequest(parsed, { headers: new Headers({ [liteHeader]: "true" }) }); + expect(new Headers(built.headers).get(liteHeader)).toBe("operator-owned"); + } + }); +}); diff --git a/tests/codex-integration/combos.test.ts b/tests/codex-integration/combos.test.ts index 76e4f26ca9..0bc1230e0f 100644 --- a/tests/codex-integration/combos.test.ts +++ b/tests/codex-integration/combos.test.ts @@ -297,22 +297,93 @@ describe("combo request cloning", () => { expect(concrete.input).not.toBe(raw.input); }); - test("combo default respects client-owned ignored reasoning values", () => { + test("combo target capability strips unsupported client reasoning controls", () => { expect(concreteComboRequestBody({ model: "combo/x", reasoning: null }, target, "high", []).reasoning).toBeNull(); expect(concreteComboRequestBody( { model: "combo/x", reasoning: { effort: "" } }, target, "high", [], - ).reasoning).toEqual({ effort: "" }); + ).reasoning).toBeUndefined(); expect(concreteComboRequestBody( { model: "combo/x", reasoning: { effort: "banana" } }, target, "high", [], - ).reasoning).toEqual({ effort: "banana" }); + ).reasoning).toBeUndefined(); expect(concreteComboRequestBody( { model: "combo/x", reasoning: { effort: null } }, target, "high", [], - ).reasoning).toEqual({ effort: null }); + ).reasoning).toBeUndefined(); expect(concreteComboRequestBody( { model: "combo/x", reasoning: { summary: "concise" } }, target, "high", ["high"], ).reasoning).toEqual({ summary: "concise", effort: "high" }); }); + test("adaptive normalization strips unsupported controls for an unknown target while preserving summary", () => { + const raw = { + model: "combo/x", + input: "hi", + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }; + const concrete = concreteComboRequestBody(raw, target, null, undefined, "adaptive"); + + expect(concrete).toEqual({ + model: "a/m1", + input: "hi", + reasoning: { summary: "concise" }, + }); + expect(raw).toEqual({ + model: "combo/x", + input: "hi", + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }); + }); + + test("strict normalization preserves reasoning controls for an unknown target", () => { + const raw = { + model: "combo/x", + input: "hi", + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }; + const concrete = concreteComboRequestBody(raw, target, null, undefined, "strict"); + + expect(concrete).toEqual({ + model: "a/m1", + input: "hi", + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }); + }); + + test("explicit empty ladder strips unsupported controls while preserving reasoning summary", () => { + const concrete = concreteComboRequestBody({ + model: "combo/x", + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }, target, "high", []); + + expect(concrete).toEqual({ + model: "a/m1", + reasoning: { summary: "concise" }, + }); + }); + + test("adaptive normalization preserves xhigh for a known reasoning ladder", () => { + const concrete = concreteComboRequestBody({ + model: "combo/x", + reasoning: { effort: "xhigh", summary: "concise" }, + }, target, null, ["low", "medium", "high", "xhigh"], "adaptive"); + + expect(concrete.reasoning).toEqual({ effort: "xhigh", summary: "concise" }); + }); + test("omits combo defaults for unset, no-reasoning, and unknown target capabilities", () => { expect(concreteComboRequestBody({ model: "combo/x" }, target, null, ["high"]).reasoning).toBeUndefined(); // An explicitly empty ladder is how a no-reasoning model is expressed. diff --git a/tests/fixtures/test-layout-expected.json b/tests/fixtures/test-layout-expected.json index 0118a80600..ba54e84bf2 100644 --- a/tests/fixtures/test-layout-expected.json +++ b/tests/fixtures/test-layout-expected.json @@ -922,6 +922,7 @@ "responses-account-label.test.ts": "responses", "responses-compaction-routing.test.ts": "responses", "responses-compaction.test.ts": "responses", + "responses-console-go-upload-retry.test.ts": "responses", "responses-context-overflow.test.ts": "responses", "responses-custom-tool-guidance.test.ts": "responses", "responses-custom-tool-repair.test.ts": "responses", @@ -945,10 +946,10 @@ "responses-pool-401-refresh.test.ts": "responses", "responses-pool-refresh-attribution.test.ts": "responses", "responses-reasoning-summary-passthrough.test.ts": "responses", - "responses-reasoning-summary-rewrite.test.ts": "responses", "responses-routed-web-search-fields.test.ts": "responses", "responses-self-named-namespace-scrub.test.ts": "responses", "responses-shadow-intercept.test.ts": "responses", + "responses-show-thinking-summary.test.ts": "responses", "responses-snapshot-repair-server.test.ts": "responses", "responses-snapshot-repair.test.ts": "responses", "responses-state-write-amplification.test.ts": "responses", diff --git a/tests/providers/opencode-go-luna-wire.test.ts b/tests/providers/opencode-go-luna-wire.test.ts index c0afb15553..7866ec9415 100644 --- a/tests/providers/opencode-go-luna-wire.test.ts +++ b/tests/providers/opencode-go-luna-wire.test.ts @@ -163,19 +163,17 @@ describe("OpenCode Go stateless reasoning and continuation routes", () => { const initial = { type: "message", role: "user", content: [{ type: "input_text", text: "Run probe" }] }; const first = await drive({ input: [initial] }); expect(first.document.output[0]).toEqual(reasoning[0]); - expect(first.document.output[1]).toEqual(continuation.summary === "auto" ? { - type: "reasoning", id: `rs_${prefix}_content`, status: "completed", summary: [{ type: "summary_text", text: "Visible thinking" }], - } : reasoning[1]); + // The passthrough keeps native content-channel reasoning in both display modes. + expect(first.document.output[1]).toEqual(reasoning[1]); expect(first.document.output[2]).toEqual(reasoning[2]); expect(first.document.output[3]).toMatchObject(call); expect(first.document.output[4]).toEqual(priorMessage); if (streaming) { - const channel = continuation.summary === "auto" ? "reasoning_summary_text" : "reasoning_text"; - expect(first.text).toContain(`"type":"response.${channel}.delta"`); + expect(first.text).toContain('"type":"response.reasoning_text.delta"'); } const result = { type: "function_call_output", call_id: call.call_id, output: "probe succeeded" }; // Echo exactly the client-visible history through handleResponses. An upstream-shape - // cache would prepend it again after the content-to-summary rewrite (F1). + // cache would prepend it again (F1). const nextBody = { input: continuation.fullHistory ? [initial, ...first.document.output, result] : [result], previous_response_id: first.document.id, store: true, @@ -207,9 +205,8 @@ describe("OpenCode Go stateless reasoning and continuation routes", () => { expect(replay.filter(item => item.type === "reasoning")).toHaveLength(3); expect(replay).toContainEqual(expect.objectContaining({ type: "reasoning", encrypted_content: blob })); expect(JSON.stringify(replay)).toContain("Already summarized"); - if (continuation.summary === "auto") expect(replay).toContainEqual(expect.objectContaining({ - type: "reasoning", summary: [{ type: "summary_text", text: "Visible thinking" }], - })); + // Replay sanitation strips reasoning content in both display modes (F1), so the + // visible "Visible thinking" trace does not re-enter the upstream history. expect(JSON.stringify(replay)).not.toContain("no tool result was recorded"); }); } diff --git a/tests/providers/opencode-zen-rate-limit.test.ts b/tests/providers/opencode-zen-rate-limit.test.ts index 8751082c2a..e8e90d9416 100644 --- a/tests/providers/opencode-zen-rate-limit.test.ts +++ b/tests/providers/opencode-zen-rate-limit.test.ts @@ -6,6 +6,7 @@ import { enrichOpenCodeZenFreeTierMessage, enrichOpenCodeZenRateLimitMessage, enrichOpenCodeZenUpstreamMessage, + isTransientConsoleGoUploadRejection, isOpenCodeZenFreeTierLockIn, isOpenCodeZenRateLimitProvider, } from "../../src/providers/opencode-zen-rate-limit"; @@ -230,3 +231,79 @@ describe("opencode-free keyless tier lock-in (#4121)", () => { expect(lockedOut).not.toContain(OPENCODE_ZEN_OBSERVED_RPM_HINT); }); }); + +describe("Console Go transient upload refusal", () => { + // The observed Go-route refusal, byte-for-byte as Console serves it. + const GO_MESSAGE = "Error from provider (Console Go): Upstream request failed: [invalid_request_error] Invalid upload request."; + // The Zen key route names the same gateway without the Go suffix. + const ZEN_MESSAGE = "Error from provider (Console): Upstream request failed: [invalid_request_error] Invalid upload request."; + const envelope = (message: string) => JSON.stringify({ model: "muse-spark-1.3-contributor", error: { param: null, type: "invalid_request_error", message } }); + const GO_ROUTE = { outboundUrl: "https://opencode.ai/zen/go/v1/responses" }; + const ZEN_ROUTE = { outboundUrl: "https://opencode.ai/zen/v1/responses" }; + + test("accepts the canonical refusal on both canonical Console routes", () => { + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), ...GO_ROUTE })).toBe(true); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(ZEN_MESSAGE), ...ZEN_ROUTE })).toBe(true); + // A custom row pointed at the same destination is still Console: the base URL decides. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), outboundUrl: "https://opencode.ai/zen/go/v1/responses" })).toBe(true); + }); + + test("rejects the refusal text from a non-Console route", () => { + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), outboundUrl: "https://api.deepseek.com/v1/responses" })).toBe(false); + // opencode.ai without the /zen segment is not the Console gateway. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), outboundUrl: "https://opencode.ai/v1/responses" })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE) })).toBe(false); + }); + + test("requires the effective canonical HTTPS endpoint", () => { + const userInfoUrl = new URL("https://opencode.ai/zen/go/v1/responses"); + userInfoUrl.username = "fixture-user"; + userInfoUrl.password = "fixture-password"; + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), outboundUrl: userInfoUrl.href })).toBe(false); + for (const outboundUrl of [ + "http://opencode.ai/zen/go/v1/responses", + "https://opencode.ai:8443/zen/go/v1/responses", + "https://opencode.ai/zen/go/v1/responses?tenant=fixture", + "https://opencode.ai/zen/go/v1/responses#fragment", + "https://opencode.ai.evil.test/zen/go/v1/responses", + "https://opencode.ai/zen-other/v1/responses", + "https://opencode.ai/zen/go/v1/models", + ]) expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), outboundUrl })).toBe(false); + }); + + test("rejects noncanonical envelopes, suffixes, and other statuses", () => { + // A bare string is not the structured envelope. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: JSON.stringify({ error: "Invalid upload request." }), ...GO_ROUTE })).toBe(false); + // A suffix means the gateway said something else; do not guess. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE + " Please retry."), ...GO_ROUTE })).toBe(false); + // Different status: only the gateway 400 is the flap. + expect(isTransientConsoleGoUploadRejection({ status: 500, errorBody: envelope(GO_MESSAGE), ...GO_ROUTE })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: undefined, ...GO_ROUTE })).toBe(false); + // Other 400s on the same wire are verdicts on the request, not flaps. + expect(isTransientConsoleGoUploadRejection({ + status: 400, + errorBody: envelope("Error from provider (Console Go): Upstream request failed: [invalid_request_error] reasoning_effort max requires an active Muse Code subscription for model muse-spark-1.3-contributor."), + ...GO_ROUTE, + })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ + status: 400, + errorBody: JSON.stringify({ type: "error", error: { type: "MissingSessionID", message: "Request is missing x-opencode-session" } }), + ...GO_ROUTE, + })).toBe(false); + }); + + test("rejects partial envelopes and padded messages", () => { + const withError = (error: unknown) => JSON.stringify({ model: "muse-spark-1.3-contributor", error }); + // type carries the refusal identity; a partial envelope is a different error. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ param: null, message: GO_MESSAGE }), ...GO_ROUTE })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ param: null, type: "server_error", message: GO_MESSAGE }), ...GO_ROUTE })).toBe(false); + // param must be present and null. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ type: "invalid_request_error", message: GO_MESSAGE }), ...GO_ROUTE })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ param: "input", type: "invalid_request_error", message: GO_MESSAGE }), ...GO_ROUTE })).toBe(false); + // Padding means the gateway wrapped or appended something. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ param: null, type: "invalid_request_error", message: " " + GO_MESSAGE }), ...GO_ROUTE })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ param: null, type: "invalid_request_error", message: GO_MESSAGE + "\n" }), ...GO_ROUTE })).toBe(false); + // The exact canonical envelope still matches. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), ...GO_ROUTE })).toBe(true); + }); +}); diff --git a/tests/responses/openai-responses-passthrough.test.ts b/tests/responses/openai-responses-passthrough.test.ts index d5a49fd39e..f6b411ec4e 100644 --- a/tests/responses/openai-responses-passthrough.test.ts +++ b/tests/responses/openai-responses-passthrough.test.ts @@ -335,6 +335,55 @@ test("canonical forward providers normalize trailing slashes and let the pool ov expect(request.headers["chatgpt-account-id"]).toBe("runtime-account"); }); +test("noncanonical Responses preserves provider-owned safety-buffering hints", async () => { + const upstream = [ + 'event: response.created\ndata: {"type":"response.created","response":{"id":"resp_custom"},"safety_buffering":{"provider_owned":true}}\n\n', + 'event: response.metadata\ndata: {"type":"response.metadata","metadata":{"type":"safety_buffering","provider_owned":true}}\n\n', + 'event: response.completed\ndata: {"type":"response.completed","response":{"id":"resp_custom","status":"completed","output":[]}}\n\n', + "data: [DONE]\n\n", + ].join(""); + const savedFetch = globalThis.fetch; + globalThis.fetch = (async () => new Response(upstream, { headers: { + "content-type": "text/event-stream", + "x-codex-safety-buffering-enabled": "provider-owned", + "x-codex-safety-buffering-faster-model": "provider-model", + } })) as typeof fetch; + try { + for (const providerConfig of [ + { + adapter: "openai-responses", + baseUrl: "https://fixture.test/v1", + authMode: "key" as const, + apiKey: "fixture-key", + }, + { + adapter: "openai-responses", + baseUrl: "https://fixture.test/v1", + authMode: "forward" as const, + headers: { authorization: "Bearer provider-static" }, + }, + ]) { + const config = { + port: 0, + defaultProvider: "fixture", + dropCodexSafetyBuffering: true, + providers: { fixture: providerConfig }, + } as OcxConfig; + const response = await handleResponses(new Request("http://localhost/v1/responses", { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ model: "fixture/model", stream: true, input: "ping" }), + }), config, { model: "", provider: "" }); + + expect(response.headers.get("x-codex-safety-buffering-enabled")).toBe("provider-owned"); + expect(response.headers.get("x-codex-safety-buffering-faster-model")).toBe("provider-model"); + expect(await response.text()).toBe(upstream); + } + } finally { + globalThis.fetch = savedFetch; + } +}); + test("noncanonical pool-required providers use only their configured static credentials", () => { const adapter = createResponsesPassthroughAdapter({ adapter: "openai-responses", @@ -4533,3 +4582,31 @@ describe("raw usage passthrough on the forward path (#41980 parity, #37138 adjac } }); }); + + +test("canonical Responses hint suppression is opt-in at the request boundary", async () => { + const savedFetch = globalThis.fetch; + globalThis.fetch = (async () => new Response([ + 'data: {"type":"response.created","response":{"id":"resp_hint"},"safety_buffering":true}\n\n', + 'data: {"type":"response.metadata","metadata":{"type":"safety_buffering"}}\n\n', + 'data: {"type":"response.completed","response":{"id":"resp_hint","status":"completed","output":[]}}\n\n', + ].join(""), { headers: { "content-type": "text/event-stream", + "x-codex-safety-buffering-enabled": "true", "x-codex-safety-buffering-faster-model": "fixture-model", "x-codex-turn-id": "fixture-turn" } })) as typeof fetch; + try { + for (const dropCodexSafetyBuffering of [undefined, false, true]) { + const config = { port: 0, dropCodexSafetyBuffering, providers: { openai: { + ...provider, codexAccountMode: "direct", upstreamWebsocket: false, + } } } as OcxConfig; + const response = await handleResponses(new Request("http://localhost/v1/responses", { + method: "POST", headers: { "content-type": "application/json", authorization: "Bearer fixture-forward-token" }, + body: JSON.stringify({ model: "openai/gpt-5.6-sol", input: "ping", stream: true }), + }), config, { model: "", provider: "" }); + expect(response.status).toBe(200); + expect(response.headers.has("x-codex-safety-buffering-enabled")).toBe(dropCodexSafetyBuffering !== true); + expect(response.headers.get("x-codex-turn-id")).toBe("fixture-turn"); + const text = await response.text(); + expect(text.includes("safety_buffering")).toBe(dropCodexSafetyBuffering !== true); + expect(text).toContain("response.completed"); + } + } finally { globalThis.fetch = savedFetch; } +}); diff --git a/tests/responses/passthrough-abort.test.ts b/tests/responses/passthrough-abort.test.ts index 46100c6902..913da3eadd 100644 --- a/tests/responses/passthrough-abort.test.ts +++ b/tests/responses/passthrough-abort.test.ts @@ -79,7 +79,7 @@ describe("passthrough relayWithAbort (RC2, passthrough path)", () => { expect(sseBranch).toContain("rewriteBlocks: clientBlockRewrite"); // Elsewhere the failed-tail relay converts mid-stream resets into a clean response.failed. expect(sseBranch).toMatch( - /relaySseWithFailedTail\(\s*rewrittenBody,\s*upstream,\s*reason\s*=>\s*\{\s*responseCompletionCancelled\s*=\s*true;\s*clientGone\.abort\(reason\);\s*\},\s*\{\s*upstreamError:\s*logCtx\.upstreamError\s*\},\s*\)/, + /relaySseWithFailedTail\(\s*rewrittenBody,\s*upstream,\s*reason\s*=>\s*\{\s*responseCompletionCancelled\s*=\s*true;\s*clientGone\.abort\(reason\);\s*\},\s*\{\s*upstreamError:\s*logCtx\.upstreamError,\s*terminalBoundary:\s*codexSafetyBufferingOptions\s*\},\s*\)/, ); expect(sseBranch).toContain("new Response(clientBody"); expect(sseBranch).toContain("markNativePassthroughSseResponse"); diff --git a/tests/responses/passthrough-headers.test.ts b/tests/responses/passthrough-headers.test.ts index 8018e77906..2cd5261609 100644 --- a/tests/responses/passthrough-headers.test.ts +++ b/tests/responses/passthrough-headers.test.ts @@ -1,5 +1,6 @@ import { describe, expect, test } from "bun:test"; -import { sanitizePassthroughHeaders } from "../../src/server"; +import { codexSafetyBufferingFilterOptions, sanitizePassthroughHeaders } from "../../src/server"; +import { createSseTerminalOutputBoundary } from "../../src/server/relay"; describe("passthrough header sanitization (RC5 / F4)", () => { test("content-type: text/event-stream survives sanitization", () => { @@ -43,3 +44,75 @@ describe("passthrough header sanitization (RC5 / F4)", () => { expect(sanitized.get("content-type")).toBe("text/event-stream"); }); }); + +describe("codex safety-buffering hint headers", () => { + const upstream = () => new Headers({ + "content-type": "text/event-stream", + "x-codex-safety-buffering-enabled": "true", + "X-Codex-Safety-Buffering-Faster-Model": "gpt-5.6-luna", + "x-codex-primary-used-percent": "12", + "openai-model": "gpt-6-astra", + }); + + test("forwarded verbatim by default and when the option is off", () => { + for (const options of [undefined, {}, { dropCodexSafetyBuffering: false }]) { + const sanitized = sanitizePassthroughHeaders(upstream(), options); + expect(sanitized.get("x-codex-safety-buffering-enabled")).toBe("true"); + expect(sanitized.get("x-codex-safety-buffering-faster-model")).toBe("gpt-5.6-luna"); + } + }); + + test("dropped case-insensitively when opted in, other x-codex headers survive", () => { + const sanitized = sanitizePassthroughHeaders(upstream(), { dropCodexSafetyBuffering: true }); + expect(sanitized.has("x-codex-safety-buffering-enabled")).toBe(false); + expect(sanitized.has("x-codex-safety-buffering-faster-model")).toBe(false); + expect(sanitized.get("x-codex-primary-used-percent")).toBe("12"); + expect(sanitized.get("openai-model")).toBe("gpt-6-astra"); + expect(sanitized.get("content-type")).toBe("text/event-stream"); + }); + + test("codexSafetyBufferingFilterOptions only enables the drop on an explicit true", () => { + expect(codexSafetyBufferingFilterOptions({})).toEqual({ dropCodexSafetyBuffering: false }); + expect(codexSafetyBufferingFilterOptions({ dropCodexSafetyBuffering: false })) + .toEqual({ dropCodexSafetyBuffering: false }); + expect(codexSafetyBufferingFilterOptions({ dropCodexSafetyBuffering: true })) + .toEqual({ dropCodexSafetyBuffering: true }); + }); +}); + +describe("Codex safety-buffering SSE hints at the client output boundary", () => { + const encoder = new TextEncoder(); + const decoder = new TextDecoder(); + const frames = [ + 'event: response.created\ndata: {"type":"response.created","response":{"id":"resp_1"},"safety_buffering":{"retry_model":"gpt-5.6-luna"}}\n\n', + 'event: response.metadata\ndata: {"type":"response.metadata","metadata":{"type":"safety_buffering","retry_model":"gpt-5.6-luna"}}\n\n', + 'event: response.metadata\ndata: {"type":"response.metadata","metadata":{"type":"other","turn":1}}\n\n', + 'event: response.output_text.delta\ndata: {"type":"response.output_text.delta","delta":"hi"}\n\n', + 'event: response.completed\ndata: {"type":"response.completed","response":{"id":"resp_1","status":"completed"}}\n\n', + ]; + const relay = (options?: { dropCodexSafetyBuffering?: boolean }): string => { + const boundary = createSseTerminalOutputBoundary(options); + let out = ""; + for (const frame of frames) out += decoder.decode(boundary.feed(encoder.encode(frame))); + out += decoder.decode(boundary.finish()); + boundary.dispose(); + return out; + }; + + test("relayed verbatim by default and when the option is off", () => { + for (const options of [undefined, {}, { dropCodexSafetyBuffering: false }]) { + expect(relay(options)).toBe(frames.join("")); + } + }); + + test("metadata event dropped and field stripped when opted in, other events untouched", () => { + const out = relay({ dropCodexSafetyBuffering: true }); + expect(out).not.toContain("safety_buffering"); + expect(out).not.toContain("gpt-5.6-luna"); + expect(out).toContain('data: {"type":"response.created","response":{"id":"resp_1"}}'); + expect(out).toContain(frames[2]); + expect(out).toContain(frames[3]); + expect(out).toContain(frames[4]); + expect(out.match(/^event: /gm)).toHaveLength(4); + }); +}); diff --git a/tests/responses/responses-console-go-upload-retry.test.ts b/tests/responses/responses-console-go-upload-retry.test.ts new file mode 100644 index 0000000000..4421848bf2 --- /dev/null +++ b/tests/responses/responses-console-go-upload-retry.test.ts @@ -0,0 +1,263 @@ +import { afterEach, beforeEach, describe, expect, spyOn, test } from "bun:test"; +import { mkdtempSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { markResponseNonReplayable } from "../../src/lib/upstream-retry"; +import { handleResponses } from "../../src/server/responses/core"; +import type { RequestLogContext } from "../../src/server/request-log"; +import type { OcxConfig } from "../../src/types"; +import { removeTreeWithRetry } from "../helpers/remove-tree"; + +const originalFetch = globalThis.fetch; +const originalOpenCodexHome = process.env.OPENCODEX_HOME; + +/** The exact Console Go rejection for a payload it accepts moments later. */ +const UPLOAD_REFUSAL = JSON.stringify({ + model: "muse-spark-1.3-contributor", + error: { + param: null, + type: "invalid_request_error", + message: "Error from provider (Console Go): Upstream request failed: [invalid_request_error] Invalid upload request.", + }, +}); + +/** A deterministic 400 on the same wire: a verdict on the request, never a flap. */ +const EFFORT_REFUSAL = JSON.stringify({ + model: "muse-spark-1.3-contributor", + error: { + param: "reasoning.effort", + type: "invalid_request_error", + message: "Error from provider (Console Go): Upstream request failed: [invalid_request_error] reasoning_effort max requires an active Muse Code subscription for model muse-spark-1.3-contributor.", + }, +}); + +let testDir = ""; + +beforeEach(() => { + testDir = mkdtempSync(join(tmpdir(), "ocx-console-go-upload-retry-")); + process.env.OPENCODEX_HOME = testDir; +}); + +afterEach(() => { + globalThis.fetch = originalFetch; + if (originalOpenCodexHome === undefined) delete process.env.OPENCODEX_HOME; + else process.env.OPENCODEX_HOME = originalOpenCodexHome; + removeTreeWithRetry(testDir); +}); + +function config(): OcxConfig { + return { + defaultProvider: "go", + providers: { + go: { + adapter: "openai-responses", + baseUrl: "https://opencode.ai/zen/go/v1", + authMode: "key", + apiKey: "go-test-key", + }, + other: { + adapter: "openai-responses", + baseUrl: "https://other.example.test/v1", + authMode: "key", + apiKey: "other-test-key", + }, + }, + } as OcxConfig; +} + +function request(stream = false, provider = "go"): Request { + return new Request("http://localhost/v1/responses", { + method: "POST", + headers: { + "content-type": "application/json", + "session_id": "thread-console-go-upload-retry", + }, + body: JSON.stringify({ + model: provider + "/muse-spark-1.3-contributor", + stream, + store: false, + input: [ + { type: "message", role: "user", content: [{ type: "input_text", text: "hi" }] }, + ], + }), + }); +} + +function refusal(status = 400, body = UPLOAD_REFUSAL): Response { + return new Response(body, { status, headers: { "content-type": "application/json" } }); +} + +function success(id: string): Response { + return Response.json({ id, object: "response", status: "completed", model: "muse-spark-1.3-contributor", output: [] }); +} + +describe("Console Go transient upload refusal recovery", () => { + test("replays the refusal once and serves the retry with a byte-identical body", async () => { + const outbound: string[] = []; + globalThis.fetch = (async (_input: RequestInfo | URL, init?: RequestInit) => { + outbound.push(String(init?.body)); + return outbound.length === 1 ? refusal() : success("resp-upload-retry-recovered"); + }) as typeof fetch; + const logCtx: RequestLogContext = { model: "", provider: "" }; + + const response = await handleResponses(request(), config(), logCtx); + + expect(response.status).toBe(200); + expect(outbound).toHaveLength(2); + // The replay must preserve the exact serialized request. + expect(outbound[1]).toBe(outbound[0]); + expect(logCtx.activeAttempt?.recoveryKinds).toEqual(["console-go-upload-retry"]); + }); + + test("does not replay a different 400 from the same wire", async () => { + let sends = 0; + globalThis.fetch = (async () => { + sends += 1; + return refusal(400, EFFORT_REFUSAL); + }) as typeof fetch; + + const response = await handleResponses(request(), config(), { model: "", provider: "" }); + + expect(response.status).toBe(400); + expect(sends).toBe(1); + }); + + test("keeps a repeated refusal visible after the single bounded replay", async () => { + let sends = 0; + globalThis.fetch = (async () => { + sends += 1; + return refusal(); + }) as typeof fetch; + const logCtx: RequestLogContext = { model: "", provider: "" }; + + const response = await handleResponses(request(), config(), logCtx); + + expect(response.status).toBe(400); + expect(sends).toBe(2); + expect(logCtx.activeAttempt?.recoveryKinds).toEqual(["console-go-upload-retry"]); + }); + + test("does not replay the same refusal text from a non-Console provider", async () => { + let sends = 0; + globalThis.fetch = (async () => { + sends += 1; + return refusal(); + }) as typeof fetch; + + const response = await handleResponses(request(false, "other"), config(), { model: "", provider: "" }); + + expect(response.status).toBe(400); + expect(sends).toBe(1); + }); +}); + + +describe("Console destination and translated recovery controls", () => { + test("a query-bearing Console destination never authorizes another POST", async () => { + const cfg = config(); + cfg.providers.go!.baseUrl = "https://opencode.ai/zen/go/v1?tenant=fixture"; + const outbound: string[] = []; + globalThis.fetch = (async input => { outbound.push(String(input)); return refusal(); }) as typeof fetch; + const response = await handleResponses(request(), cfg, { model: "", provider: "" }); + expect(response.status).toBe(400); + expect(outbound).toHaveLength(1); + expect(new URL(outbound[0]!).search).toBe("?tenant=fixture"); + }); + + test("a canonical row name cannot authorize a noncanonical generation path", async () => { + const cfg = config(); + cfg.providers["opencode-go"] = { + ...cfg.providers.go!, + responsesPath: "/unrelated", + chatCompletionsPath: "/unrelated", + }; + const outbound: string[] = []; + globalThis.fetch = (async input => { outbound.push(String(input)); return refusal(); }) as typeof fetch; + const response = await handleResponses(request(false, "opencode-go"), cfg, { model: "", provider: "" }); + expect(response.status).toBe(400); + // Canonical row names normalize their base URL. A configured send path survives + // that normalization and reaches the effective-destination recovery gate. + expect(outbound).toEqual(["https://opencode.ai/zen/go/v1/unrelated"]); + expect(await response.text()).toContain("Invalid upload request."); + }); + + test("normalization to the canonical endpoint keeps its bounded recovery", async () => { + const cfg = config(); + cfg.providers["opencode-go"] = { ...cfg.providers.go!, baseUrl: "https://other.example.test/v1" }; + const outbound: string[] = []; + globalThis.fetch = (async input => { outbound.push(String(input)); return refusal(); }) as typeof fetch; + const response = await handleResponses(request(false, "opencode-go"), cfg, { model: "", provider: "" }); + expect(response.status).toBe(400); + expect(outbound).toEqual([ + "https://opencode.ai/zen/go/v1/responses", + "https://opencode.ai/zen/go/v1/responses", + ]); + expect(await response.text()).toContain("Invalid upload request."); + }); + + for (const adapter of ["openai-responses", "openai-chat"] as const) { + for (const stream of [false, true]) { + test(`${adapter} stream=${stream} replays identical bytes once`, async () => { + const cfg = config(); + cfg.providers.go!.adapter = adapter; + const outbound: string[] = []; + globalThis.fetch = (async (_url: RequestInfo | URL, init?: RequestInit) => { + outbound.push(String(init?.body)); + if (outbound.length === 1) return refusal(); + if (adapter === "openai-responses") { + const completed = { id: "resp_fixture", object: "response", status: "completed", output: [] }; + return stream ? new Response(`event: response.completed\ndata: ${JSON.stringify({ type: "response.completed", response: completed })}\n\n`, { headers: { "content-type": "text/event-stream" } }) : Response.json(completed); + } + if (!stream) return Response.json({ id: "chat_fixture", object: "chat.completion", choices: [{ index: 0, message: { role: "assistant", content: "answer" }, finish_reason: "stop" }] }); + const chunk = { id: "chat_fixture", object: "chat.completion.chunk", choices: [{ index: 0, delta: { content: "answer" }, finish_reason: "stop" }] }; + return new Response(`data: ${JSON.stringify(chunk)}\n\ndata: [DONE]\n\n`, { headers: { "content-type": "text/event-stream" } }); + }) as typeof fetch; + const logCtx: RequestLogContext = { model: "", provider: "" }; + const response = await handleResponses(request(stream), cfg, logCtx); + const body = await response.text(); + expect(response.status).toBe(200); + expect(outbound).toHaveLength(2); + expect(outbound[1]).toBe(outbound[0]); + expect(logCtx.activeAttempt?.recoveryKinds).toEqual(["console-go-upload-retry"]); + expect(body).toContain("completed"); + }); + } + test(`${adapter} abort during backoff sends no replay`, async () => { + const cfg = config(); cfg.providers.go!.adapter = adapter; + const controller = new AbortController(); + let sends = 0; + globalThis.fetch = (async () => { sends++; return refusal(); }) as typeof fetch; + const originalTimeout = globalThis.setTimeout; + const spy = spyOn(globalThis, "setTimeout").mockImplementation(((handler: TimerHandler, ms?: number, ...args: unknown[]) => { + if (ms === 800) queueMicrotask(() => controller.abort()); + return originalTimeout(handler, ms, ...args); + }) as typeof setTimeout); + try { + const response = await handleResponses(request(), cfg, { model: "", provider: "" }, { abortSignal: controller.signal }); + expect(controller.signal.aborted).toBe(true); + expect(response.status).toBe(499); + expect(sends).toBe(1); + } finally { spy.mockRestore(); } + }); + } +}); + + +describe("Console nonreplayable response boundary", () => { + for (const adapter of ["openai-responses", "openai-chat"] as const) { + test(`${adapter} does not replay a marked response`, async () => { + const cfg = config(); cfg.providers.go!.adapter = adapter; + let sends = 0; + globalThis.fetch = (async () => { + sends++; + const response = refusal(); + markResponseNonReplayable(response); + return response; + }) as typeof fetch; + const response = await handleResponses(request(), cfg, { model: "", provider: "" }); + expect(response.status).toBe(400); + expect(sends).toBe(1); + expect(await response.text()).toContain("Invalid upload request."); + }); + } +}); diff --git a/tests/responses/responses-reasoning-summary-passthrough.test.ts b/tests/responses/responses-reasoning-summary-passthrough.test.ts index 5912221354..ab6e099bcc 100644 --- a/tests/responses/responses-reasoning-summary-passthrough.test.ts +++ b/tests/responses/responses-reasoning-summary-passthrough.test.ts @@ -6,10 +6,12 @@ import type { OcxConfig } from "../../src/types"; /** * The passthrough relay for DeepSeek's native /responses endpoint emits - * content-channel reasoning (reasoning_text.delta + content items). The - * summary-channel rewrite must engage only when the client did NOT ask for - * hidden thinking (hideThinkingSummary) - otherwise a client that asked to - * hide reasoning would get it surfaced as visible summary output. + * content-channel reasoning (reasoning_text.delta + content items) in BOTH + * display modes: Codex applies its own raw-reasoning display policy, so a + * requested summary must not rewrite the native passthrough shape either. + * Hidden thinking (hideThinkingSummary) and visible summary get the same + * content-channel passthrough; the hidden variant additionally arrives as an + * envelope-only item upstream when the adapter layer handles suppression. */ function deepseekSeed() { @@ -86,15 +88,16 @@ describe("passthrough reasoning summary rewrite honors hideThinkingSummary", () expect(text).toContain('"content":[{"type":"reasoning_text","text":"think"}]'); }); - test("SSE: requested summary routes raw reasoning through the summary channel", async () => { + test("SSE: requested summary keeps the native content-channel passthrough", async () => { const response = await runHandleResponses( { model: "deepseek-v4-flash", input: "ping", stream: true, reasoning: { effort: "max", summary: "detailed" } }, SSE_UPSTREAM_FRAMES.join(""), "text/event-stream", ); const text = await response.text(); - expect(text).toContain("response.reasoning_summary_text.delta"); - expect(text).toContain('"summary":[{"type":"summary_text","text":"think"}]'); + expect(text).toContain("response.reasoning_text.delta"); + expect(text).not.toContain("response.reasoning_summary_text.delta"); + expect(text).toContain('"content":[{"type":"reasoning_text","text":"think"}]'); }); test("bounded JSON: hidden thinking keeps the content shape", async () => { @@ -108,14 +111,14 @@ describe("passthrough reasoning summary rewrite honors hideThinkingSummary", () expect(text).not.toContain('"summary":[{"type":"summary_text"'); }); - test("bounded JSON: requested summary moves item content into summary", async () => { + test("bounded JSON: requested summary keeps the content shape", async () => { const response = await runHandleResponses( { model: "deepseek-v4-flash", input: "ping", stream: false, reasoning: { effort: "max", summary: "detailed" } }, JSON_UPSTREAM, "application/json", ); const text = await response.text(); - expect(text).toContain('"summary":[{"type":"summary_text","text":"think"}]'); - expect(text).not.toContain('"content":[{"type":"reasoning_text","text":"think"}]'); + expect(text).toContain('"content":[{"type":"reasoning_text","text":"think"}]'); + expect(text).not.toContain('"summary":[{"type":"summary_text"'); }); }); diff --git a/tests/responses/responses-reasoning-summary-rewrite.test.ts b/tests/responses/responses-reasoning-summary-rewrite.test.ts deleted file mode 100644 index 32940a24e7..0000000000 --- a/tests/responses/responses-reasoning-summary-rewrite.test.ts +++ /dev/null @@ -1,286 +0,0 @@ -import { describe, expect, test } from "bun:test"; -import { - createReasoningSummaryChannelPayloadRewrite, - routeUsesContentChannelReasoning, - rewriteReasoningSummaryInJson, - rewriteReasoningSummaryInJsonString, -} from "../../src/server/responses-reasoning-summary-rewrite"; - -const rewrite = createReasoningSummaryChannelPayloadRewrite(); - -function apply(payload: unknown): unknown { - return JSON.parse(rewrite(JSON.stringify(payload))); -} - -describe("responses reasoning summary channel rewrite", () => { - test("routes reasoning_text.delta through the summary channel", () => { - expect(apply({ - type: "response.reasoning_text.delta", - content_index: 0, - delta: "think", - item_id: "rs_1", - output_index: 0, - sequence_number: 4, - })).toEqual({ - type: "response.reasoning_summary_text.delta", - summary_index: 0, - delta: "think", - item_id: "rs_1", - output_index: 0, - sequence_number: 4, - }); - }); - - test("routes reasoning_text.done through the summary channel", () => { - expect(apply({ - type: "response.reasoning_text.done", - content_index: 0, - text: "full thinking", - item_id: "rs_1", - output_index: 0, - })).toEqual({ - type: "response.reasoning_summary_text.done", - summary_index: 0, - text: "full thinking", - item_id: "rs_1", - output_index: 0, - }); - }); - - test("moves reasoning item content into summary on output_item.done", () => { - expect(apply({ - type: "response.output_item.done", - output_index: 0, - item: { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [{ type: "reasoning_text", text: "thinking" }], - summary: [], - }, - })).toEqual({ - type: "response.output_item.done", - output_index: 0, - item: { - type: "reasoning", - id: "rs_1", - status: "completed", - summary: [{ type: "summary_text", text: "thinking" }], - }, - }); - }); - - test("moves reasoning item content into summary inside response.completed", () => { - const payload = { - type: "response.completed", - response: { - id: "resp_1", - status: "completed", - output: [ - { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [{ type: "reasoning_text", text: "thinking" }], - summary: [], - }, - { type: "message", id: "msg_1", status: "completed", content: [{ type: "output_text", text: "OK" }] }, - ], - }, - }; - const result = apply(payload) as { response: { output: Record[] } }; - expect(result.response.output[0]).toEqual({ - type: "reasoning", - id: "rs_1", - status: "completed", - summary: [{ type: "summary_text", text: "thinking" }], - }); - expect(result.response.output[1]).toEqual(payload.response.output[1]); - }); - - test("leaves summary-channel and message events untouched", () => { - const untouched = [ - { type: "response.reasoning_summary_text.delta", summary_index: 0, delta: "s", item_id: "rs_1", output_index: 0 }, - { type: "response.output_text.delta", content_index: 0, delta: "OK", item_id: "msg_1", output_index: 1 }, - { type: "response.output_item.added", output_index: 1, item: { type: "message", id: "msg_1", status: "in_progress", content: [] } }, - ]; - for (const payload of untouched) { - expect(apply(payload)).toEqual(payload); - } - }); - - test("leaves a reasoning item without content text untouched", () => { - expect(apply({ - type: "response.output_item.done", - output_index: 0, - item: { type: "reasoning", id: "rs_1", status: "completed", content: [], summary: [] }, - })).toEqual({ - type: "response.output_item.done", - output_index: 0, - item: { type: "reasoning", id: "rs_1", status: "completed", content: [], summary: [] }, - }); - }); - - test("preserves a summary-channel reasoning item as-is", () => { - expect(apply({ - type: "response.output_item.done", - output_index: 0, - item: { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [], - summary: [{ type: "summary_text", text: "already summarized" }], - }, - })).toEqual({ - type: "response.output_item.done", - output_index: 0, - item: { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [], - summary: [{ type: "summary_text", text: "already summarized" }], - }, - }); - }); - - test("rewrites reasoning items inside a bare completed response document", () => { - const doc = { - id: "resp_1", - object: "response", - status: "completed", - output: [ - { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [{ type: "reasoning_text", text: "thinking" }], - summary: [], - }, - { type: "message", id: "msg_1", status: "completed", content: [{ type: "output_text", text: "OK" }] }, - ], - }; - const result = rewriteReasoningSummaryInJson(doc) as { output: Record[] }; - expect(result.output[0]).toEqual({ - type: "reasoning", - id: "rs_1", - status: "completed", - summary: [{ type: "summary_text", text: "thinking" }], - }); - expect(result.output[1]).toEqual(doc.output[1]); - }); - - test("rewrites reasoning items inside an SSE completed event document", () => { - const doc = { - type: "response.completed", - response: { - id: "resp_1", - status: "completed", - output: [ - { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [{ type: "reasoning_text", text: "thinking" }], - summary: [], - }, - ], - }, - }; - const result = rewriteReasoningSummaryInJson(doc) as { response: { output: Record[] } }; - expect(result.response.output[0]).toEqual({ - type: "reasoning", - id: "rs_1", - status: "completed", - summary: [{ type: "summary_text", text: "thinking" }], - }); - }); - - test("string-level rewrite leaves summary-channel documents untouched", () => { - const doc = JSON.stringify({ - id: "resp_1", - output: [{ type: "reasoning", id: "rs_1", summary: [{ type: "summary_text", text: "already summarized" }] }], - }); - expect(rewriteReasoningSummaryInJsonString(doc)).toBe(doc); - }); - - test("malformed payloads pass through unchanged", () => { - expect(rewrite("not json")).toBe("not json"); - expect(rewrite("[1,2]")).toBe("[1,2]"); - }); - - // `encrypted_content` is opaque, state-bearing provider data, so preserve the complete item - // shape defensively when the client replays it. This rewrite's round-trip was verified against - // DeepSeek, which is stateless and issues no blob; providers that do issue one joined later - // through `preserveReasoningContentModels`. - describe("items carrying encrypted_content", () => { - const blobItem = { - type: "reasoning", - id: "rs_1", - status: "completed", - encrypted_content: "gAAAAAB-upstream-issued-blob", - content: [{ type: "reasoning_text", text: "thinking" }], - summary: [], - }; - - test("are returned byte-for-byte on output_item.done", () => { - const payload = { type: "response.output_item.done", output_index: 0, item: blobItem }; - expect(apply(payload)).toEqual(payload); - }); - - test("are returned byte-for-byte inside response.completed output", () => { - const payload = { - type: "response.completed", - response: { id: "resp_1", output: [blobItem] }, - }; - expect(apply(payload)).toEqual(payload); - }); - - test("are returned byte-for-byte through the non-streaming document rewrite", () => { - const doc = { id: "resp_1", object: "response", output: [blobItem] }; - expect(rewriteReasoningSummaryInJson(doc)).toBe(doc); - const json = JSON.stringify(doc); - expect(rewriteReasoningSummaryInJsonString(json)).toBe(json); - }); - - // Only the stored item is protected: the live trace Codex renders comes from the delta events, - // which carry no blob and are still routed to the summary channel. - test("do not disable the delta rewrite that renders the live trace", () => { - expect(apply({ - type: "response.reasoning_text.delta", - delta: "think", - item_id: "rs_1", - output_index: 0, - })).toMatchObject({ type: "response.reasoning_summary_text.delta", delta: "think" }); - }); - }); -}); - -describe("routeUsesContentChannelReasoning", () => { - test("statelessResponses providers use the content channel", () => { - expect(routeUsesContentChannelReasoning({ statelessResponses: true }, "deepseek-v4-flash")).toBe(true); - }); - - test("preserveReasoningContentModels lists qualify", () => { - expect(routeUsesContentChannelReasoning( - { preserveReasoningContentModels: ["deepseek-v4-flash"] }, - "deepseek-v4-flash", - )).toBe(true); - }); - - test("model matching is case-insensitive on both sides", () => { - expect(routeUsesContentChannelReasoning( - { preserveReasoningContentModels: ["DeepSeek-V4-Flash"] }, - "deepseek-v4-flash", - )).toBe(true); - expect(routeUsesContentChannelReasoning( - { preserveReasoningContentModels: ["deepseek-v4-flash"] }, - "DeepSeek-V4-Flash", - )).toBe(true); - }); - - test("other providers do not", () => { - expect(routeUsesContentChannelReasoning({}, "gpt-5.5")).toBe(false); - }); -}); diff --git a/tests/responses/responses-show-thinking-summary.test.ts b/tests/responses/responses-show-thinking-summary.test.ts new file mode 100644 index 0000000000..a749518b63 --- /dev/null +++ b/tests/responses/responses-show-thinking-summary.test.ts @@ -0,0 +1,173 @@ +import { afterEach, describe, expect, test } from "bun:test"; +import { mkdtempSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { providerConfigSeed } from "../../src/providers/derive"; +import { getProviderRegistryEntry } from "../../src/providers/registry"; +import { handleResponses } from "../../src/server/responses/core"; +import type { OcxConfig, OcxProviderConfig } from "../../src/types"; + +// Provider-opted visible thinking (showThinkingSummary): a provider that serves +// genuine user-facing reasoning surfaces it on the summary channel even when the +// client omits reasoning.summary (the Codex default, which otherwise hides all +// thinking in replay-only envelopes). An explicit client summary of "none" still +// wins and keeps thinking hidden. + +function shownSeed() { + const seed = providerConfigSeed(getProviderRegistryEntry("deepseek")!); + return { ...seed, apiKey: "sk-test", showThinkingSummary: true } as OcxProviderConfig; +} + +function sseFrame(payload: unknown): string { + return "data: " + JSON.stringify(payload) + "\n\n"; +} + +const SSE_UPSTREAM = [ + sseFrame({ type: "response.created", response: { id: "resp_1", status: "in_progress", output: [] } }), + sseFrame({ type: "response.output_item.added", output_index: 0, item: { type: "reasoning", id: "rs_1", status: "in_progress", content: [], summary: [] } }), + sseFrame({ type: "response.reasoning_text.delta", content_index: 0, delta: "think", item_id: "rs_1", output_index: 0 }), + sseFrame({ type: "response.reasoning_text.done", content_index: 0, text: "think", item_id: "rs_1", output_index: 0 }), + sseFrame({ type: "response.output_item.done", output_index: 0, item: { type: "reasoning", id: "rs_1", status: "completed", content: [{ type: "reasoning_text", text: "think" }], summary: [] } }), + sseFrame({ type: "response.completed", response: { id: "resp_1", status: "completed", output: [{ type: "reasoning", id: "rs_1", status: "completed", content: [{ type: "reasoning_text", text: "think" }], summary: [] }] } }), +].join(""); + +async function runHandleResponses(body: Record, seed: OcxProviderConfig) { + const encoder = new TextEncoder(); + globalThis.fetch = (async () => new Response( + new ReadableStream({ + start(controller) { + controller.enqueue(encoder.encode(SSE_UPSTREAM)); + controller.close(); + }, + }), + { status: 200, headers: { "content-type": "text/event-stream" } }, + )) as typeof fetch; + const config = { providers: { deepseek: seed } } as unknown as OcxConfig; + return handleResponses( + new Request("http://localhost/v1/responses", { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify(body), + }), + config, + { model: "", provider: "" }, + { abortSignal: AbortSignal.timeout(5_000) }, + ); +} + +describe("showThinkingSummary provider option", () => { + const originalFetch = globalThis.fetch; + afterEach(() => { globalThis.fetch = originalFetch; }); + + test("provider opt-in never relabels raw content as a summary", async () => { + const response = await runHandleResponses( + { model: "deepseek-v4-flash", input: "ping", stream: true }, + shownSeed(), + ); + const text = await response.text(); + expect(text).not.toContain("response.reasoning_summary_text.delta"); + expect(text).toContain('"content":[{"type":"reasoning_text","text":"think"}]'); + }); + + test("explicit client summary none keeps thinking hidden", async () => { + const response = await runHandleResponses( + { model: "deepseek-v4-flash", input: "ping", stream: true, reasoning: { summary: "none" } }, + shownSeed(), + ); + const text = await response.text(); + expect(text).not.toContain("response.reasoning_summary_text.delta"); + expect(text).toContain("response.reasoning_text.delta"); + }); + + test("without the provider option, omitted summary stays hidden", async () => { + const seed = { ...providerConfigSeed(getProviderRegistryEntry("deepseek")!), apiKey: "sk-test" } as OcxProviderConfig; + const response = await runHandleResponses( + { model: "deepseek-v4-flash", input: "ping", stream: true }, + seed, + ); + const text = await response.text(); + expect(text).not.toContain("response.reasoning_summary_text.delta"); + expect(text).toContain("response.reasoning_text.delta"); + }); + + test("google-antigravity preset opts in", () => { + expect(providerConfigSeed(getProviderRegistryEntry("google-antigravity")!).showThinkingSummary).toBe(true); + expect(providerConfigSeed(getProviderRegistryEntry("deepseek")!).showThinkingSummary).toBeUndefined(); + }); + + for (const stream of [false, true]) for (const [summary, providerFlag, visible] of [ + [undefined, undefined, true], ["none", true, false], [undefined, false, false], ["auto", false, false], ["auto", true, true], + ] as const) test(`CCA summary=${summary} provider=${providerFlag} stream=${stream}`, async () => { + const home = mkdtempSync(join(tmpdir(), "ocx-show-thinking-")); + const prevHome = process.env.OPENCODEX_HOME; + process.env.OPENCODEX_HOME = home; + writeFileSync(join(home, "auth.json"), JSON.stringify({ + "google-antigravity": { + activeAccountId: "active", + accounts: [{ + id: "active", + credential: { + access: "access-token", + refresh: "refresh-token", + expires: Date.now() + 3_600_000, + projectId: "project-id", + }, + }], + }, + })); + const seen: string[] = []; + const requests: Array<{ request: { generationConfig?: { thinkingConfig?: { includeThoughts?: boolean } } } }> = []; + globalThis.fetch = (async (input: RequestInfo | URL, init?: RequestInit) => { + seen.push(String(input)); + requests.push(JSON.parse(String(init?.body))); + const requestedThoughts = requests.at(-1)?.request.generationConfig?.thinkingConfig?.includeThoughts === true; + const payload = { + response: { + candidates: [{ + content: { parts: [...(requestedThoughts ? [{ thought: true, text: "cca-think" }] : []), { text: "OK" }] }, + finishReason: "STOP", + }], + usageMetadata: { promptTokenCount: 10, candidatesTokenCount: 5, totalTokenCount: 15, thoughtsTokenCount: 3 }, + }, + }; + return stream + ? new Response(sseFrame(payload), { headers: { "content-type": "text/event-stream" } }) + : Response.json(payload); + }) as typeof fetch; + try { + const seed = { + ...providerConfigSeed(getProviderRegistryEntry("google-antigravity")!), + liveModels: false, + models: ["gemini-3.8-flash"], + } as OcxProviderConfig; + // Simulate a saved provider row written before the registry learned the flag: + // the request path must backfill it from the registry entry (routedProviderConfig), + // enrichProviderFromRegistry never runs there. + delete seed.showThinkingSummary; + if (providerFlag !== undefined) seed.showThinkingSummary = providerFlag; + const config = { providers: { "google-antigravity": seed } } as unknown as OcxConfig; + const response = await handleResponses( + new Request("http://localhost/v1/responses", { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ model: "google-antigravity/gemini-3.8-flash", input: "ping", stream, reasoning: { effort: "low", ...(summary ? { summary } : {}) } }), + }), + config, + { model: "", provider: "" }, + { abortSignal: AbortSignal.timeout(10_000) }, + ); + const text = await response.text(); + expect(seen).toHaveLength(1); + expect(seen[0]).toContain(stream ? "v1internal:streamGenerateContent" : "v1internal:generateContent"); + expect(text.includes('"summary":[{"type":"summary_text","text":"cca-think"}]')).toBe(visible); + expect(requests[0].request.generationConfig?.thinkingConfig?.includeThoughts === true) + .toBe(visible && providerFlag !== false); + if (stream) expect(text).toContain("response.completed"); + expect(text).toContain("OK"); + } finally { + if (prevHome === undefined) delete process.env.OPENCODEX_HOME; + else process.env.OPENCODEX_HOME = prevHome; + rmSync(home, { recursive: true, force: true }); + } + }); +}); diff --git a/tests/responses/sse-failed-tail.test.ts b/tests/responses/sse-failed-tail.test.ts index f21301ffb7..0a38424f42 100644 --- a/tests/responses/sse-failed-tail.test.ts +++ b/tests/responses/sse-failed-tail.test.ts @@ -422,3 +422,43 @@ describe("relaySseWithFailedTail", () => { expect(out).not.toContain("event: response.failed"); }); }); + + +describe("optional Codex hint filtering preserves relay semantics", () => { + for (const eager of [false, true]) { + const relay = (chunks: string[], upstreamError?: string) => eager + ? relaySseEagerBounded(sourceStream(chunks), new AbortController(), parityHooks, + { upstreamError, terminalBoundary: { dropCodexSafetyBuffering: true } }) + : relaySseWithFailedTail(sourceStream(chunks), new AbortController(), undefined, + { upstreamError, terminalBoundary: { dropCodexSafetyBuffering: true } }); + test(`policy failure plus hint is composed, eager=${eager}`, async () => { + const frame = `event: error\r\ndata: ${JSON.stringify({ type: "error", safety_buffering: { enabled: true }, + error: { code: "cyber_policy", message: "blocked by upstream policy", type: "invalid_request_error" } })}\r\n\r\n`; + const text = await drain(relay([frame.slice(0, 19), frame.slice(19)])); + expect(text).not.toContain("safety_buffering"); + expect(text).toContain('"type":"response.failed"'); + expect(text).toContain('"code":"cyber_policy"'); + expect(text).toContain('"retryable":false'); + expect(text.match(/data: \[DONE\]/g)).toHaveLength(1); + }); + test(`metadata removal retains other frames and one terminal, eager=${eager}`, async () => { + const metadata = 'data: {"type":"response.metadata","metadata":{"type":"safety_buffering"}}\n\n'; + const other = 'data: {"type":"codex.response.metadata","headers":{"x-codex-safety-buffering-enabled":"true"}}\n\n'; + const malformed = 'data: {malformed}\n\n'; + const terminal = 'data: {"type":"response.completed","response":{"status":"completed"},"safety_buffering":true}\n\ndata: [DONE]\n\n'; + const text = await drain(relay([metadata.slice(0, 7), metadata.slice(7), other, malformed, terminal])); + expect(text).not.toContain('"type":"safety_buffering"'); + expect(text).not.toContain('"safety_buffering":true'); + expect(text).toContain(other); + expect(text).toContain(malformed); + expect(text).toContain('"type":"response.completed"'); + expect(text.match(/data: \[DONE\]/g)).toHaveLength(1); + }); + test(`hint-only EOF preserves captured error fallback, eager=${eager}`, async () => { + const text = await drain(relay(['data: {"type":"response.metadata","metadata":{"type":"safety_buffering"}}\n\n'], "provider unavailable")); + expect(text).toContain("provider unavailable"); + expect(text).not.toContain("adapter_eof"); + expect(text).not.toContain("safety_buffering"); + }); + } +}); diff --git a/tests/responses/ws-upstream-reuse.test.ts b/tests/responses/ws-upstream-reuse.test.ts index b957fdb317..48b720d2ea 100644 --- a/tests/responses/ws-upstream-reuse.test.ts +++ b/tests/responses/ws-upstream-reuse.test.ts @@ -3,6 +3,8 @@ import { codexWsUpstreamFetch } from "../../src/server/responses/ws-upstream"; import { runOptionalShutdownHooks } from "../../src/lib/optional-shutdown-hooks"; import { CodexWsPool, codexWsPool } from "../../src/server/responses/codex-ws-pool"; import { prepareCodexWsRequest } from "../../src/server/responses/codex-ws-request"; +import { createResponsesPassthroughAdapter } from "../../src/adapters/openai-responses"; +import { withTestTranslatorBudget } from "../helpers/translator-budget"; const URL = "https://chatgpt.com/backend-api/codex/responses"; const realWebSocket = globalThis.WebSocket; @@ -307,3 +309,31 @@ test("a Lite mode change retires the old handshake", async () => { expect(Socket.all).toHaveLength(2); expect(Socket.all[0]!.readyState).toBe(3); }); + +test("adapter Spark Lite override retires a legacy socket and reuses the disabled identity", async () => { + const liteHeader = "x-openai-internal-codex-responses-lite"; + const liteKey = "ws_request_header_x_openai_internal_codex_responses_lite"; + const options = init(); + const rawBody = { ...JSON.parse(options.body as string), model: "gpt-5.3-codex-spark", + client_metadata: { thread_id: "fixture-thread", turn_id: "fixture-turn", [liteKey]: "true" } }; + const before = JSON.stringify(rawBody); + const adapter = withTestTranslatorBudget(createResponsesPassthroughAdapter({ + adapter: "openai-responses", authMode: "forward", baseUrl: "https://chatgpt.com/backend-api/codex", + })); + const built = await adapter.buildRequest({ modelId: "spark-alias", context: { messages: [] }, + stream: true, options: {}, _rawBody: rawBody, + }, { headers: new Headers(options.headers) }); + const current = { ...options, body: built.body, headers: built.headers }; + // Keep the exact same Spark model/scope/headers; only the old delete-only Lite policy differs. + const legacyHeaders = new Headers(current.headers); + legacyHeaders.delete(liteHeader); + await drain({ ...current, headers: legacyHeaders }); + await drain(current); + await drain(current); + expect(Socket.all).toHaveLength(2); + expect(Socket.all.map(socket => socket.readyState)).toEqual([3, 1]); + expect(Socket.all.map(socket => socket.frames.map(frame => + (frame.client_metadata as Record)[liteKey]))).toEqual([["true"], ["false", "false"]]); + expect(Socket.all.flatMap(socket => socket.frames).every(frame => frame.model === rawBody.model)).toBe(true); + expect(JSON.stringify(rawBody)).toBe(before); +}); diff --git a/tests/server/config.test.ts b/tests/server/config.test.ts index b096b5857c..13cc92f708 100644 --- a/tests/server/config.test.ts +++ b/tests/server/config.test.ts @@ -777,6 +777,19 @@ describe("opencodex config defaults", () => { }); }); + test("codex safety-buffering header drop is an explicit top-level opt-in", () => { + const defaults = getDefaultConfig(); + expect(defaults.dropCodexSafetyBuffering).toBe(false); + expect(validateConfigCandidate({ ...defaults, dropCodexSafetyBuffering: true })).toMatchObject({ + ok: true, + config: { dropCodexSafetyBuffering: true }, + }); + expect(validateConfigCandidate({ ...defaults, dropCodexSafetyBuffering: "yes" })).toMatchObject({ + ok: false, + error: expect.stringContaining("dropCodexSafetyBuffering"), + }); + }); + test("usage and MCP config overrides change the effective bound while defaults remain compatible", () => { const defaults = getDefaultConfig(); expect(defaults.managementUsageMaxReadBytes).toBe(64 * 1024 * 1024); diff --git a/tests/server/server-combo-failover-e2e.test.ts b/tests/server/server-combo-failover-e2e.test.ts index 6f649c52d3..436f62a0ff 100644 --- a/tests/server/server-combo-failover-e2e.test.ts +++ b/tests/server/server-combo-failover-e2e.test.ts @@ -1182,6 +1182,88 @@ describe("server combo failover 030 activation matrix", () => { expectMappedReceipt(hydrated[0]!); }); + test("adaptive combo normalizes unknown and empty target capability before the upstream wire", async () => { + const bodies: Array> = []; + const upstream = serve(async request => { + bodies.push(await request.json() as Record); + return chatSuccess("normalized", "m1"); + }); + const request = { + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }; + + const unknownResponse = await post( + comboConfig( + { a: provider("openai-chat", baseUrl(upstream), "key-a") }, + undefined, + { reasoningEffortMode: "adaptive" }, + ), + request, + ); + expect(unknownResponse.status).toBe(200); + + const emptyResponse = await post( + comboConfig( + { a: provider("openai-chat", baseUrl(upstream), "key-a", { reasoningEfforts: [] }) }, + undefined, + { reasoningEffortMode: "adaptive" }, + ), + request, + ); + expect(emptyResponse.status).toBe(200); + + const knownResponse = await post( + comboConfig( + { a: provider("openai-chat", baseUrl(upstream), "key-a", { + reasoningEfforts: ["low", "medium", "high", "xhigh"], + }) }, + undefined, + { reasoningEffortMode: "adaptive" }, + ), + request, + ); + expect(knownResponse.status).toBe(200); + + expect(bodies).toHaveLength(3); + for (const body of bodies.slice(0, 2)) { + expect(body).not.toHaveProperty("reasoning_effort"); + expect(body).not.toHaveProperty("thinking_budget"); + expect(body).not.toHaveProperty("thinking"); + } + expect(bodies[2]!.reasoning_effort).toBe("xhigh"); + }); + + test("adaptive Responses combo preserves summary after removing unsupported controls", async () => { + const bodies: Array> = []; + const upstream = serve(async request => { + bodies.push(await request.json() as Record); + return Response.json(responsesSuccess("normalized", "m1")); + }); + for (const reasoningEfforts of [undefined, [], ["high", "xhigh"]]) { + const response = await post(comboConfig({ + a: provider("openai-responses", baseUrl(upstream), "key-a", { + ...(reasoningEfforts === undefined ? {} : { reasoningEfforts }), + modelSupportsReasoningSummaries: { m1: true }, + }), + }, undefined, { reasoningEffortMode: "adaptive" }), { + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", thinking_budget: 8192, thinking: { type: "enabled" }, + }); + expect(response.status).toBe(200); + } + expect(bodies).toHaveLength(3); + for (const body of bodies.slice(0, 2)) { + expect(body.reasoning).toEqual({ summary: "concise" }); + expect(body).not.toHaveProperty("reasoning_effort"); + expect(body).not.toHaveProperty("thinking_budget"); + expect(body).not.toHaveProperty("thinking"); + } + expect(bodies[2]!.reasoning).toMatchObject({ effort: "xhigh", summary: "concise" }); + }); + test("all-target exhaustion promotes the final attempt reasoning wire to the logical row", async () => { const a = serve(() => Response.json({ error: { message: "first overloaded" } }, { status: 503 })); const b = serve(() => Response.json({ error: { message: "last overloaded" } }, { status: 503 })); @@ -3937,3 +4019,35 @@ describe("combo compact failover", () => { expect(await response.text()).toContain("empty summary"); }); }); + + +describe("thinking-summary defaults follow the serving combo route", () => { + for (const firstVisible of [true, false]) for (const summary of [undefined, "none", "auto"]) { + test(`fallback from ${firstVisible} with summary=${summary}`, async () => { + const observed: Array<[string, boolean | undefined]> = []; + customRunTurn = async (parsed, _incoming, emit) => { + observed.push([parsed.modelId, parsed.options.hideThinkingSummary]); + if (parsed.modelId === "m1") { + emit({ type: "error", message: "provider unavailable", status: 503, retryable: true }); + return; + } + emit({ type: "thinking_delta", thinking: "Actual provider summary" }); + emit({ type: "text_delta", text: "Final fallback answer" }); + emit({ type: "done" }); + }; + const config = comboConfig({ + a: provider("test-run-turn", "https://a.test/v1", "key-a", { showThinkingSummary: firstVisible }), + b: provider("test-run-turn", "https://b.test/v1", "key-b", { showThinkingSummary: !firstVisible }), + }); + const response = await post(config, { ...(summary ? { reasoning: { summary } } : {}) }); + const output = JSON.stringify(await response.json()); + expect(response.status).toBe(200); + expect(observed).toEqual([ + ["m1", summary === "none" || (!summary && !firstVisible)], + ["m2", summary === "none" || (!summary && firstVisible)], + ]); + expect(output.includes("Actual provider summary")).toBe(summary === "auto" || (!summary && !firstVisible)); + expect(output).toContain("Final fallback answer"); + }); + } +}); diff --git a/tests/server/server-xai-chat-reasoning-streaming.test.ts b/tests/server/server-xai-chat-reasoning-streaming.test.ts index 23da211fd0..0b1db50d0d 100644 --- a/tests/server/server-xai-chat-reasoning-streaming.test.ts +++ b/tests/server/server-xai-chat-reasoning-streaming.test.ts @@ -151,7 +151,7 @@ describe("xAI OAuth Chat reasoning streaming", () => { let received = ""; await Promise.race([ (async () => { - while (!received.includes("response.reasoning_summary_text.delta")) { + while (!received.includes("response.reasoning_text.delta")) { const chunk = await reader!.read(); if (chunk.done) throw new Error("stream ended before the first xAI reasoning delta"); received += decoder.decode(chunk.value, { stream: true }); @@ -187,7 +187,7 @@ describe("xAI OAuth Chat reasoning streaming", () => { if (chunk.done) break; received += decoder.decode(chunk.value, { stream: true }); } - const reasoningIndex = received.indexOf("response.reasoning_summary_text.delta"); + const reasoningIndex = received.indexOf("response.reasoning_text.delta"); const contentIndex = received.indexOf("response.output_text.delta"); const completedIndex = received.indexOf("response.completed"); expect(reasoningIndex).toBeGreaterThanOrEqual(0);