diff --git a/devlog/_plan/260912_stream_recovery/092_resume.md b/devlog/_plan/260912_stream_recovery/092_resume.md new file mode 100644 index 0000000000..a6463ddfe4 --- /dev/null +++ b/devlog/_plan/260912_stream_recovery/092_resume.md @@ -0,0 +1,20 @@ +# Stream delivery verification update + +Five source fixes were delivered as independent dev-based PRs. PR4341 (terminal integrity) is merged; PR4354 (Console), PR4356 (search), PR4363 (Cursor) and PR4367 (live sideband) remain open at the 2026-09-12 verification update. This task does not merge integration branches. Runtime verification remains incomplete; shared Cline, history and Windows failures are not treated as passing baseline evidence. + +## Console label capture + +![Console recovery label with synthetic log data](093_console_recovery.png) + +The capture uses the dashboard-preview artifact from GitHub Actions run34675214816, artifact10294058486, recorded build commit f9dcd6449298821e6ba02d026037a677effe6aaa and GUI tree e0d22385336080717ad29a14d65a270a207a41e7. The artifact was built in hosted CI. No local product build, install, typecheck or test suite ran. + +The artifact GUI differs from Console source head2c63a5d4283936aa9d0525e39090afb5e492c0ec only in Combo workspace files and associated tests; Logs.tsx, its locale strings and styles are unchanged. The Logs page was served locally from the existing bundle with a synthetic API fixture. The screenshot shows the actual Console upload retry label in the request detail dialog. Values, request ID and model are synthetic, not live request or billing evidence. This is rendered label evidence, not a backend retry test. + +## Remaining acceptance + +- #3389 remains HOLD: zero observed bytes cannot establish upstream nonexecution. +- #4191 retains its native failing-stage evidence requirement; existing diagnostics are not a reproduced fix. +- #4312 has its open-tool status sub-defect addressed; the actual client nonretryable refusal contract remains unresolved. +- #3506 requires a redacted translation-fidelity exchange. No semantic-progress cutoff is added. + +Fresh source/security reviews and exact CI heads/runs are retained in the task-local handoff. Skipped, cancelled, failed and superseded runs do not certify completion. Local suites/build/typecheck/install: NOT RUN. diff --git a/devlog/_plan/260912_stream_recovery/093_console_recovery.png b/devlog/_plan/260912_stream_recovery/093_console_recovery.png new file mode 100644 index 0000000000..43e6fdf359 Binary files /dev/null and b/devlog/_plan/260912_stream_recovery/093_console_recovery.png differ diff --git a/devlog/_plan/260912_stream_recovery/094_console_fixture.md b/devlog/_plan/260912_stream_recovery/094_console_fixture.md new file mode 100644 index 0000000000..31e4e77c5b --- /dev/null +++ b/devlog/_plan/260912_stream_recovery/094_console_fixture.md @@ -0,0 +1,7 @@ +# Effective Console destination fixture correction + +Hosted macOS control run34693025423/job103551974484 failed the canonical-row other-host fixture. The log records that routing discarded the configured other-host URL and selected the canonical Go endpoint. The test therefore never reached its intended negative condition. This is not evidence of a noncanonical replay. + +The negative case now configures an unsupported generation path on the same canonical row. Both supported adapter path fields name `/unrelated`; the assertion requires exactly one outgoing URL at that path and preserves the original refusal. Existing custom-row other-host coverage remains. A separate control records the two exact `/responses` sends produced after canonical base normalization; the exact Muse model wire default selects Responses. No production code, policy, retry count or assertion skip changes. + +Local tests/build/typecheck/install are NOT RUN. Independent source review and updated exact-head hosted CI are required. Windows pnpm/Devin failures in the same run remain separate shared repair items; no baseline runtime pass is claimed. diff --git a/devlog/_plan/260912_thinking_contract/000_plan.md b/devlog/_plan/260912_thinking_contract/000_plan.md new file mode 100644 index 0000000000..86a87f9d36 --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/000_plan.md @@ -0,0 +1,31 @@ +# Preserve reasoning provenance and transport intent + +Readers: maintainers choosing whether to integrate the thinking lane. Raw reasoning must remain content, while provider-authored summaries can be displayed under an explicit provider default. The plan reconciles #4301 and #4287, separately reviews #3652 hint suppression, and carries #4130 Spark compatibility without retirement. + +Loop: satisfy-spec HOTL, triggered by authorized thinking-lane delivery. Goal: reviewable carry PRs and final-head hosted CI. Non-goals: merges, closure of source PRs, retirement #4334, releases, user service/config changes, other worktrees. All local product suites/build/typecheck/install are NOT RUN by instruction. Only available existing credentials/tools are used; no user token/time/agent ceiling was set. Stop: every source PR has a justified disposition and every delivered branch has exact-head hosted CI evidence. Outcomes: DONE on evidence, HOLD/NEEDS_HUMAN on explicit unresolved acceptance, never fake green. Escalation: real tool denial or requirement beyond scope; main reclaims after two distinct reviewer failures. Native architect selector is unavailable; inherited independent design review and reflection follow the user instruction, with a separate A audit. + +## Dependency map + +| Cycle | Artifact | Result | +| --- | --- | --- | +| roadmap | this file and all decade docs | docs-only plan lock | +| presentation | 010_presentation.md | raw/summary contract and provider opt-in | +| hint | 020_transport_hint.md | independent transport-hint disposition/carry | +| spark | 030_spark.md | independent Spark Lite carry | +| delivery | 040_delivery.md | final heads, review closure and hosted CI | + +Presentation combines two conflicting source proposals into one contract. Hint and Spark are independent and receive ordinary dev-based PRs, not artificial stack dependencies. Final review consumes all branches. No GitHub native stacks are requested. + +## Evidence and owner map + +Baseline origin/dev: 69e3dcda755a52feb1327edad6c8ea6cefd6e871. Source PR heads: #4301 5d6d1862a11da6e4d0c04eb7f35f9f48ae1285fd; #4287 fe13bdb7bf8403a2a2cdb10f258a68b649177953; #3652 13fb263778e9036e66ae86d41e29f9f47bbbed92; #4130 5d56f5461ea3d18668b85f6bb0d8a523920f2536. All open when inspected. Original authors: Robin Bially, yxr1995-maker, itismyfield, luvs01; exact Git trailers will be read from original commits before carrying. + +Current owners: src/bridge.ts:663 raw-reasoning finalization; src/adapters/google.ts:571 shared part classifier; src/server/responses/core.ts:2490 final-route normalization; src/responses/parser.ts:543 summary omission policy; src/types/request.ts:310 AdapterEvent. Reuse these boundaries; no new event enum or generic service layer. Structure INDEX maps shared areas to topical documents; main contracts are providers/chat-compat.md, providers/google.md, transports/responses.md and config.md with references from affected area owners. + +Verification: git diff --check was run at baseline and exited 0, checking diff whitespace only. GitHub ci.yml workflow_dispatch lane=all reads checkout source, typechecks, runs product suites and cross-platform jobs; NOT RUN locally. Every conditional scenario is named in decade docs and must be asserted in committed regression tests. Source inspection is not runtime proof. + +## Cycle records + +Roadmap P: requirements/source inspection and independent design review in progress. No product patch applied. + +Roadmap B: locked amended contract after independent A PASS and both design reflections ALIGNED. Product implementation starts in the next cycle. diff --git a/devlog/_plan/260912_thinking_contract/001_sources.md b/devlog/_plan/260912_thinking_contract/001_sources.md new file mode 100644 index 0000000000..f05813eae8 --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/001_sources.md @@ -0,0 +1,7 @@ +# Source decisions + +Public PR diffs and latest comments are the source proposal evidence. #4130's September 11 corrections pin Lite on for nonempty additional_tools bodies and off otherwise; adopting the earlier unconditional false version loses tools. #4334 is an explicit retirement HOLD and is not carried. + +#4301 removes automatic content-to-summary conversion. #4287 tests raw DeepSeek content as a visible summary; that expectation conflicts with provenance and will be replaced, not adopted. Google thought-summary API documentation distinguishes summaries from opaque thought signatures: https://ai.google.dev/gemini-api/docs/generate-content/thinking (opened 2026-09-12). CCA generationConfig/includeThoughts behavior is contributor probe evidence, not a newly performed live-service probe. + +Searches used: reasoning_raw_delta, thinking_delta, hideThinkingSummary, googlePartTextEvent, preserveReasoningContent, and the four PR numbers. Existing bridge event types can distinguish raw content and summary without a new enum. No-code/config-only options cannot repair the existing mislabeled content; raw rewrite deletion plus existing boundaries is the smallest change. diff --git a/devlog/_plan/260912_thinking_contract/010_presentation.md b/devlog/_plan/260912_thinking_contract/010_presentation.md new file mode 100644 index 0000000000..b1503a5f34 --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/010_presentation.md @@ -0,0 +1,322 @@ +# Presentation contract + +Class C4 public protocol contract. Depends on roadmap lock. MODIFY src/bridge.ts: closeCurrentRawReasoning and reasoning_raw_delta emit response.reasoning_text.delta/done with content_index:0; final items use summary:[] and content:[{type:reasoning_text,text}]. buildResponseJSONWithBudget mirrors this. Keep hidden txt-only replay envelopes intact. DELETE src/server/responses-reasoning-summary-rewrite.ts and its obsolete unit test; MODIFY core.ts to remove imports and SSE/JSON content-to-summary rewrites. MODIFY both layout manifests to remove that test. Adopt the exact #4301 hunks below except reporter video and historical verification record. + +MODIFY provider.ts, registry.ts, derive.ts, router.ts and auth-cors.ts to carry showThinkingSummary boolean (preserve explicit false). Seed only google-antigravity true. Creation: provider config/registry; serialization: providerConfigSeed and deriveKeyLoginMap; deserialization: config provider passthrough and management field policy; consumers: routedProviderConfig, final-route normalization, Google request builder. No new enum. + +MODIFY core.ts final-route normalization: apply provider default only when original reasoning.summary is omitted, never explicit none; recompute on each final route so fallback cannot inherit another provider default. Provider opt-in authorizes summary display, not raw-to-summary conversion. + +MODIFY google.ts shared part classifier to use existing thinking_delta only for Gemini thought summaries under verified Gemini model provenance; CCA Claude/gpt-oss thought text remains reasoning_raw_delta. Persist request-local Gemini identity using existing adapter state, used by both stream and buffered classifier calls. includeThoughts stays provider-opted, Gemini-only, non-image and explicit-hide aware. MODIFY google-wire-compiler.ts to retain only boolean true includeThoughts, independently of thinkingLevel. Do not claim raw text is an actual summary. + +MODIFY the #4287 end-to-end fixture: raw DeepSeek content remains content with empty summary even under provider opt-in; actual CCA Gemini thought parts use summary; omitted vs none vs auto, explicit provider false, saved-row enrichment, fallback reset, streaming/buffered paths. Extend existing Google tests and bridge raw tests; both layout manifests register responses-show-thinking-summary.test.ts. Update English providers docs and structure owners, keeping locale statements consistent. Source tests are authored but run only by hosted CI. + +Acceptance: raw event fixture => content delta and no summary delta; actual Gemini summary fixture => summary only when requested/provider-opted; explicit none => no synthesized summary and no includeThoughts request; false/unknown provider => no opt-in; fallback to unopted route => hidden behavior reset; replay envelope decodes same raw text and tool continuation remains valid; native Responses mixed content/summary remains byte-semantically native. No model prose synthesizer is introduced. + +## Source patch blueprint + +```diff +diff --git a/src/bridge.ts b/src/bridge.ts +index 20e7c3fe09..bc90f35b94 100644 +--- a/src/bridge.ts ++++ b/src/bridge.ts +@@ -663,16 +663,13 @@ export function bridgeToResponsesSSE( + const closeCurrentRawReasoning = () => { + if (!currentRawReasoning) return; + rawReasoningForNextToolCall = currentRawReasoning.text; +- emit("response.reasoning_summary_text.done", { +- item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, summary_index: 0, text: currentRawReasoning.text, +- }); +- emit("response.reasoning_summary_part.done", { +- item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, summary_index: 0, +- part: { type: "summary_text", text: currentRawReasoning.text }, ++ emit("response.reasoning_text.done", { ++ item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, content_index: 0, text: currentRawReasoning.text, + }); + const item = { + type: "reasoning", id: currentRawReasoning.itemId, +- summary: [{ type: "summary_text", text: currentRawReasoning.text }], ++ summary: [] as never[], ++ content: [{ type: "reasoning_text", text: currentRawReasoning.text }], + }; + emit("response.output_item.done", { output_index: currentRawReasoning.outputIndex, item }); + retainFinishedItem(item as OutputItem, currentRawReasoning.textBytes, "reasoning"); +@@ -1111,10 +1108,6 @@ export function bridgeToResponsesSSE( + const itemId = `rs_${uuid()}`; + const item = { type: "reasoning", id: itemId, summary: [] as { type: string; text: string }[] }; + emit("response.output_item.added", { output_index: outputIndex, item }); +- emit("response.reasoning_summary_part.added", { +- item_id: itemId, output_index: outputIndex, summary_index: 0, +- part: { type: "summary_text", text: "" }, +- }); + currentRawReasoning = { itemId, outputIndex, text: "", textBytes: 0 }; + } + ({ value: currentRawReasoning.text, bytes: currentRawReasoning.textBytes } = appendString( +@@ -1123,9 +1116,13 @@ export function bridgeToResponsesSSE( + event.text, + "reasoning", + )); +- emit("response.reasoning_summary_text.delta", { ++ // Raw reasoning (openai-chat reasoning_content, kiro tags) rides the CONTENT ++ // channel, matching native gpt-oss passthrough: Codex applies its own display ++ // policy, so the desktop band shows the "Thinking…" placeholder instead of the ++ // raw CoT (the #45 summary-channel display intent is intentionally reverted). ++ emit("response.reasoning_text.delta", { + item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, +- summary_index: 0, delta: event.text, ++ content_index: 0, delta: event.text, + }); + break; + } +@@ -1780,7 +1777,8 @@ function buildResponseJSONWithBudget( + } + pushOutput({ + type: "reasoning", id: `rs_${uuid()}`, +- summary: [{ type: "summary_text", text: currentRawReasoning }], ++ summary: [], ++ content: [{ type: "reasoning_text", text: currentRawReasoning }], + }, currentRawReasoningBytes, "reasoning"); + currentRawReasoning = ""; + currentRawReasoningBytes = 0; + +``` + +## Pinned source hunks (apply with corrections above) + +```diff +diff --git a/src/adapters/google-wire-compiler.ts b/src/adapters/google-wire-compiler.ts +index 88c482ba7d..aa835e50b4 100644 +--- a/src/adapters/google-wire-compiler.ts ++++ b/src/adapters/google-wire-compiler.ts +@@ -130,12 +130,20 @@ function compileGenerationConfig(value: unknown): JsonObject | undefined { + ))].slice(0, 5); + if (stopSequences.length > 0) out.stopSequences = stopSequences; + } +- if (isObject(value.thinkingConfig) && typeof value.thinkingConfig.thinkingLevel === "string") { +- const raw = value.thinkingConfig.thinkingLevel.toLowerCase(); +- const thinkingLevel = GOOGLE_THINKING_LEVELS.has(raw) +- ? raw +- : (["xhigh", "max", "ultra"].includes(raw) ? "high" : undefined); +- if (thinkingLevel) out.thinkingConfig = { thinkingLevel }; ++ if (isObject(value.thinkingConfig)) { ++ const thinking: JsonObject = {}; ++ if (typeof value.thinkingConfig.thinkingLevel === "string") { ++ const raw = value.thinkingConfig.thinkingLevel.toLowerCase(); ++ const thinkingLevel = GOOGLE_THINKING_LEVELS.has(raw) ++ ? raw ++ : (["xhigh", "max", "ultra"].includes(raw) ? "high" : undefined); ++ if (thinkingLevel) thinking.thinkingLevel = thinkingLevel; ++ } ++ // The one key that makes Google return `thought: true` text. Cloud Code Assist serves ++ // thinking either way (thoughtsTokenCount stays non-zero) but withholds the text unless the ++ // request opts in, so dropping it here silently reinstates the missing-thinking behavior. ++ if (value.thinkingConfig.includeThoughts === true) thinking.includeThoughts = true; ++ if (Object.keys(thinking).length > 0) out.thinkingConfig = thinking; + } + if (Array.isArray(value.responseModalities)) { + const valid = value.responseModalities.filter((m): m is string => typeof m === "string" && ["TEXT", "IMAGE", "AUDIO"].includes(m)); +diff --git a/src/adapters/google.ts b/src/adapters/google.ts +index 7fcc88ba59..5a6675f6e4 100644 +--- a/src/adapters/google.ts ++++ b/src/adapters/google.ts +@@ -866,11 +866,27 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte + ); + antigravityModel = wireModelId; + antigravitySession = sessionId; ++ // Gemini returns no chain-of-thought TEXT unless the request opts in. Probed against CCA ++ // 2026-09-12: `gemini-3.8-flash-high` answered with thoughtsTokenCount=321 and zero ++ // `thought` parts, then 358-652 chars of genuine reasoning once includeThoughts was set. ++ // Scoped to Gemini wire ids — Claude-on-CCA accepts the flag but never returns thought ++ // parts, and gpt-oss rejects it outright (400 INVALID_ARGUMENT, which would break every ++ // gpt-oss turn). Gated on the provider's visible-thinking opt-in so a user who wants ++ // thinking hidden does not pay conversation-history tokens for text nobody renders; ++ // `hideThinkingSummary !== true` is the same per-request gate the response path uses, so ++ // a client that explicitly asked for hidden thinking is not billed for the text either. ++ const includeThoughts = provider.showThinkingSummary === true ++ && parsed.options.hideThinkingSummary !== true ++ && /^gemini-/.test(wireModelId) ++ && !isImageCapableModel(parsed.modelId); + // Effort → thinkingConfig for CCA (CLIProxyAPI proven: request.generationConfig.thinkingConfig). + // Suffix/compat IDs return thinkingLevel=undefined — the suffix IS the effort, no contradiction. +- if (thinkingLevel) { ++ if (thinkingLevel || includeThoughts) { + const gc = (body.generationConfig ?? {}) as Record; +- gc.thinkingConfig = { thinkingLevel }; ++ gc.thinkingConfig = { ++ ...(thinkingLevel ? { thinkingLevel } : {}), ++ ...(includeThoughts ? { includeThoughts: true } : {}), ++ }; + body.generationConfig = gc; + } + // Reasoning continuity: Gemini models re-inject cached thoughtSignatures; Claude-on-Antigravity +diff --git a/src/providers/derive.ts b/src/providers/derive.ts +index 67e6c0522e..7edf28787b 100644 +--- a/src/providers/derive.ts ++++ b/src/providers/derive.ts +@@ -43,6 +43,7 @@ export interface DerivedKeyLoginProvider { + autoToolChoiceOnlyModels?: string[]; + preserveReasoningContentModels?: string[]; + requiresReasoningPlaceholderModels?: string[]; ++ showThinkingSummary?: boolean; + reasoningSplitModels?: string[]; + reasoningDetailsModels?: string[]; + thinkingToggleModels?: string[]; +@@ -267,6 +268,7 @@ export function providerConfigSeed(entry: ProviderRegistryEntry): OcxProviderCon + ...(entry.autoToolChoiceOnlyModels ? { autoToolChoiceOnlyModels: [...entry.autoToolChoiceOnlyModels] } : {}), + ...(entry.preserveReasoningContentModels ? { preserveReasoningContentModels: [...entry.preserveReasoningContentModels] } : {}), + ...(entry.requiresReasoningPlaceholderModels ? { requiresReasoningPlaceholderModels: [...entry.requiresReasoningPlaceholderModels] } : {}), ++ ...(entry.showThinkingSummary !== undefined ? { showThinkingSummary: entry.showThinkingSummary } : {}), + ...(entry.reasoningSplitModels ? { reasoningSplitModels: [...entry.reasoningSplitModels] } : {}), + ...(entry.reasoningDetailsModels ? { reasoningDetailsModels: [...entry.reasoningDetailsModels] } : {}), + ...(entry.thinkingToggleModels ? { thinkingToggleModels: [...entry.thinkingToggleModels] } : {}), +@@ -315,6 +317,7 @@ export function deriveKeyLoginMap(): Record { + ...(entry.autoToolChoiceOnlyModels ? { autoToolChoiceOnlyModels: [...entry.autoToolChoiceOnlyModels] } : {}), + ...(entry.preserveReasoningContentModels ? { preserveReasoningContentModels: [...entry.preserveReasoningContentModels] } : {}), + ...(entry.requiresReasoningPlaceholderModels ? { requiresReasoningPlaceholderModels: [...entry.requiresReasoningPlaceholderModels] } : {}), ++ ...(entry.showThinkingSummary !== undefined ? { showThinkingSummary: entry.showThinkingSummary } : {}), + ...(entry.reasoningSplitModels ? { reasoningSplitModels: [...entry.reasoningSplitModels] } : {}), + ...(entry.reasoningDetailsModels ? { reasoningDetailsModels: [...entry.reasoningDetailsModels] } : {}), + ...(entry.thinkingToggleModels ? { thinkingToggleModels: [...entry.thinkingToggleModels] } : {}), +@@ -567,6 +570,7 @@ export function enrichProviderFromRegistry(name: string, prov: OcxProviderConfig + if (!prov.thinkingToggleModels && seed.thinkingToggleModels) prov.thinkingToggleModels = [...seed.thinkingToggleModels]; + if (!prov.thinkingBudgetModels && seed.thinkingBudgetModels) prov.thinkingBudgetModels = [...seed.thinkingBudgetModels]; + if (prov.escapeBuiltinToolNames === undefined && seed.escapeBuiltinToolNames !== undefined) prov.escapeBuiltinToolNames = seed.escapeBuiltinToolNames; ++ if (prov.showThinkingSummary === undefined && seed.showThinkingSummary !== undefined) prov.showThinkingSummary = seed.showThinkingSummary; + if (prov.keyOptional === undefined && seed.keyOptional !== undefined) prov.keyOptional = seed.keyOptional; + if (prov.freeTier === undefined && seed.freeTier !== undefined) prov.freeTier = seed.freeTier; + if (prov.modelSuffixBracketStrip === undefined && seed.modelSuffixBracketStrip !== undefined) prov.modelSuffixBracketStrip = seed.modelSuffixBracketStrip; +diff --git a/src/providers/registry.ts b/src/providers/registry.ts +index f72bb7650b..e483e9db23 100644 +--- a/src/providers/registry.ts ++++ b/src/providers/registry.ts +@@ -343,6 +343,10 @@ export interface ProviderRegistryEntry { + autoToolChoiceOnlyModels?: string[]; + preserveReasoningContentModels?: string[]; + requiresReasoningPlaceholderModels?: string[]; ++ /** ++ * Opt this provider into visible thinking summaries (see OcxProviderConfig.showThinkingSummary). ++ */ ++ showThinkingSummary?: boolean; + reasoningSplitModels?: string[]; + reasoningDetailsModels?: string[]; + thinkingToggleModels?: string[]; +@@ -367,7 +371,7 @@ export type ProviderConfigSeed = Pick< + | "modelMaxInputTokens" | "defaultMaxOutputTokens" | "modelMaxOutputTokens" + | "reasoningEfforts" | "modelReasoningEfforts" | "modelDefaultReasoningEfforts" | "reasoningEffortMap" | "modelReasoningEffortMap" | "reasoningWireFormat" + | "noVisionModels" | "noReasoningModels" | "noTemperatureModels" | "noTopPModels" | "noPenaltyModels" +- | "autoToolChoiceOnlyModels" | "preserveReasoningContentModels" | "requiresReasoningPlaceholderModels" | "reasoningSplitModels" | "reasoningDetailsModels" | "thinkingToggleModels" | "thinkingBudgetModels" | "escapeBuiltinToolNames" | "openaiChatEofTolerance" ++ | "autoToolChoiceOnlyModels" | "preserveReasoningContentModels" | "requiresReasoningPlaceholderModels" | "reasoningSplitModels" | "reasoningDetailsModels" | "thinkingToggleModels" | "thinkingBudgetModels" | "escapeBuiltinToolNames" | "openaiChatEofTolerance" | "showThinkingSummary" + | "googleMode" | "project" | "location" | "headers" + >; + +@@ -2045,7 +2049,7 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [ + // path must stay RELATIVE: this row sets `allowBaseUrlOverride`, and an absolute `url` would + // retarget a user's custom base back to Google. A leading `./` is required because a bare + // `v1internal:` reads as a URL scheme and `providerModelDiscoverySpecError` rejects it. +- { id: "google-antigravity", alias: "agy", label: "Google Antigravity", adapter: "google", baseUrl: "https://daily-cloudcode-pa.googleapis.com", authKind: "oauth", allowBaseUrlOverride: true, dashboardUrl: "https://antigravity.google", models: ANTIGRAVITY_MODELS, liveModels: true, defaultModel: "gemini-3.8-flash", modelContextWindows: ANTIGRAVITY_MODEL_CONTEXT_WINDOWS, modelInputModalities: ANTIGRAVITY_MODEL_INPUT_MODALITIES, modelReasoningEfforts: ANTIGRAVITY_MODEL_EFFORTS, googleMode: "cloud-code-assist", jawcodeBundle: "google", extraMetadataAliases: ["antigravity", "gemini-antigravity"], modelDiscovery: { path: "./v1internal:fetchAvailableModels" } }, ++ { id: "google-antigravity", alias: "agy", label: "Google Antigravity", adapter: "google", baseUrl: "https://daily-cloudcode-pa.googleapis.com", authKind: "oauth", allowBaseUrlOverride: true, dashboardUrl: "https://antigravity.google", models: ANTIGRAVITY_MODELS, liveModels: true, defaultModel: "gemini-3.8-flash", modelContextWindows: ANTIGRAVITY_MODEL_CONTEXT_WINDOWS, modelInputModalities: ANTIGRAVITY_MODEL_INPUT_MODALITIES, modelReasoningEfforts: ANTIGRAVITY_MODEL_EFFORTS, googleMode: "cloud-code-assist", showThinkingSummary: true, jawcodeBundle: "google", extraMetadataAliases: ["antigravity", "gemini-antigravity"], modelDiscovery: { path: "./v1internal:fetchAvailableModels" } }, + { id: "azure-openai", label: "Azure OpenAI", adapter: "azure-openai", baseUrl: "https://{resource}.openai.azure.com/openai", authKind: "key", featured: true, dashboardUrl: "https://portal.azure.com" }, + { id: "ollama", label: "Ollama (local)", adapter: "openai-chat", baseUrl: "http://localhost:11434/v1", authKind: "local", allowPrivateNetworkByDefault: true, allowBaseUrlOverride: true, featured: true, note: "Local — key usually blank" }, + { id: "vllm", label: "vLLM (local)", adapter: "openai-chat", baseUrl: "http://localhost:8000/v1", authKind: "local", allowPrivateNetworkByDefault: true, allowBaseUrlOverride: true, featured: true, note: "Local — key usually blank" }, +diff --git a/src/router.ts b/src/router.ts +index 55a0326fce..bf2b9b4b98 100644 +--- a/src/router.ts ++++ b/src/router.ts +@@ -410,6 +410,13 @@ export function routedProviderConfig(providerName: string, provider: OcxProvider + ...(provider.preserveResponsesReasoningContent === undefined && registryEntry.preserveResponsesReasoningContent !== undefined + ? { preserveResponsesReasoningContent: registryEntry.preserveResponsesReasoningContent } + : {}), ++ // The request path resolves through routedProviderConfig() and never calls ++ // enrichProviderFromRegistry(), so a saved provider row written before the ++ // registry learned this flag must be backfilled here or route.provider never ++ // carries it and the showThinkingSummary opt-in stays dead. ++ ...(provider.showThinkingSummary === undefined && registryEntry.showThinkingSummary !== undefined ++ ? { showThinkingSummary: registryEntry.showThinkingSummary } ++ : {}), + // Registry-only client-facing repair policy (#938): fill only when the + // saved provider has no explicit policy; clone so runtime never aliases + // the registry constant. +diff --git a/src/server/auth-cors.ts b/src/server/auth-cors.ts +index 3a93246cd0..93377b5573 100644 +--- a/src/server/auth-cors.ts ++++ b/src/server/auth-cors.ts +@@ -885,6 +885,7 @@ const PROVIDER_CONFIG_FIELD_POLICY = { + autoToolChoiceOnlyModels: "editor", + preserveReasoningContentModels: "editor", + requiresReasoningPlaceholderModels: "editor", ++ showThinkingSummary: "editor", + retryOn429: "editor", + transientRetryOn5xx: "editor", + reasoningSplitModels: "editor", +diff --git a/src/server/responses/core.ts b/src/server/responses/core.ts +index cccd942026..852cd9f8b0 100644 +--- a/src/server/responses/core.ts ++++ b/src/server/responses/core.ts +@@ -2467,6 +2467,20 @@ async function resolveSubagentFallbackModelEligibility(args: { + }; + } + ++/** ++ * Whether the client explicitly asked for hidden thinking (`reasoning.summary: "none"`). ++ * ++ * Pinned: parseRequest collapses "omitted" and "none" into one hideThinkingSummary ++ * flag, so the raw request body is the ONLY place that still distinguishes them. ++ * Provider opt-ins like showThinkingSummary must consult this — never the flag ++ * alone — or a future caller that copies only the flag would silently unlock an ++ * explicit opt-out. ++ */ ++function clientExplicitlyHidThinking(parsed: OcxParsedRequest): boolean { ++ const rawReasoning = (parsed._rawBody as { reasoning?: { summary?: unknown } } | undefined)?.reasoning; ++ return typeof rawReasoning === "object" && rawReasoning !== null ++ && (rawReasoning as { summary?: unknown }).summary === "none"; ++} + /** + * Apply every route-dependent request mutation against the final selected route. + * Must run only after subagent fallback has settled the model/provider. +@@ -2508,6 +2522,15 @@ async function applyFinalRouteRequestNormalization(args: { + // this request will actually use (#404). + route.provider = resolveOpenCodeGoTransport(route.provider, getOrAllocateRequestSessionLane(req)); + route.provider = resolveWireProtocolOverride(route.providerName, route.modelId, route.provider, inboundWire); ++ // Provider-opted visible thinking (e.g. google-antigravity): parseRequest hides thinking ++ // whenever the client omits reasoning.summary, which is the Codex default. A provider that ++ // serves genuine user-facing reasoning opts back into the summary channel here, so thought ++ // parts (Gemini thought, content-channel reasoning_text) reach the client instead of only ++ // the hidden replay envelopes. An explicit client reasoning.summary "none" still wins. ++ if (route.provider.showThinkingSummary === true && parsed.options.hideThinkingSummary === true ++ && !clientExplicitlyHidThinking(parsed)) { ++ parsed.options.hideThinkingSummary = false; ++ } + if (preserveAnthropicResponseModel) parsed._responseModelId = responseModelId; + logCtx.model = route.modelId; + logCtx.provider = route.providerName; +diff --git a/src/types/provider.ts b/src/types/provider.ts +index e65130a4fa..b6374a991a 100644 +--- a/src/types/provider.ts ++++ b/src/types/provider.ts +@@ -746,6 +746,15 @@ export interface OcxProviderConfig { + * out explicitly (e.g. MiniMax, where low effort disables thinking). + */ + requiresReasoningPlaceholderModels?: string[]; ++ /** ++ * Opt-in: surface upstream thinking as visible reasoning summaries even when the ++ * client did not send `reasoning.summary`. parseRequest hides thinking by default ++ * (Codex omits the field), which strands genuine reasoning — e.g. Gemini `thought` ++ * parts on the google-antigravity (Cloud Code Assist) wire — in hidden replay ++ * envelopes. An explicit client `reasoning.summary: "none"` still wins. Set `false` ++ * to opt a seeded preset back out. ++ */ ++ showThinkingSummary?: boolean; + /** + * Opt-in same-target 429 retry policy. Codex itself never retries 429 (it retries 5xx only, + * openai/codex#30471), and single-key pools have no failover, so the proxy waits and replays + +``` + +## Reflection corrections accepted + +Explicit wire reasoning.summary:"none" wins. A client that serializes configured none as omission cannot be distinguished from unspecified preference. No client config rewrite or global catalog summary default changes. Summary classification is limited to built CCA Gemini requests; unknown/uninitialized, direct Google/Vertex and CCA Claude/gpt-oss remain raw. Streaming and buffered summary-to-tool continuations assert exact Google signature on correct call; never emit Google signatures as Anthropic thinking_signature. Hidden unsigned summaries may disappear but required tool replay state survives. Exercise final assistant text and terminal order, fallback in both directions, and remove replay-comparison rewrite alongside SSE/JSON rewrite. Desktop appearance remains client-controlled; source patch comments claiming an unconditional placeholder are replaced during adoption. + +## Presentation P revalidation + +Prior D: roadmap locked; next presentation implementation. Both source patches apply to baseline; combined application requires keeping the newer no-rewrite expectation. Shared classifier signatures remain current. Implement CCA-only provider default by recomputing parsed.options.hideThinkingSummary from raw summary each final route for inboundWire responses; other inbound types preserve their existing flag. Missing raw request leaves original hide flag authoritative. CCA Gemini classification records boolean in existing per-request adapter closure on each build; default false. + +C review corrections: fixture summary emission now depends on includeThoughts; provider false + client auto does not request upstream summaries. Signature regression exercises actual SSE and JSON serializers, visible/hidden modes, a tool-ending first turn, exact signature and real matching functionResponse. Final-answer behavior is covered separately by bridge and combo tests. diff --git a/devlog/_plan/260912_thinking_contract/020_transport_hint.md b/devlog/_plan/260912_thinking_contract/020_transport_hint.md new file mode 100644 index 0000000000..abfff918c7 --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/020_transport_hint.md @@ -0,0 +1,323 @@ +# Optional hint suppression + +Class C4 review because client metadata policy changes. Independent of presentation; depends only on roadmap. Adopt #3652 only after independent security/transport review. Public proposal removes exactly two x-codex-safety-buffering headers, metadata.type=safety_buffering events and top-level safety_buffering fields at the client relay boundary. Default false; malformed config must remain off and candidate validation rejects nonbooleans. This suppresses optional transport hints; provider safety decisions/refusals and upstream checks are unchanged. Compact and independent WS/other-provider pathways retain existing policy unless a directly exercised shared boundary already applies. + +MODIFY src/config.ts and src/types/config.ts for validated boolean/default; src/server/relay.ts for allowlisted header removal and SSE terminal-boundary transformation; relay-eager.ts for option forwarding; core.ts to compute option only for canonical OpenAI forward destination and pass it to all relevant headers/client output boundaries; index.ts exports if needed by existing test style. Do not apply to custom gateway/key providers. Preserve errors, response.failed/incomplete and terminal sentinel handling. + +MODIFY tests/responses/passthrough-headers.test.ts, openai-responses-passthrough.test.ts and tests/server/config.test.ts. Scenarios: absent/false/true/malformed config; uppercase headers; unrelated headers; split metadata frames; actual failure carrying hint must still fail; noncanonical provider has identical fields and retains them; eager/non-eager client paths. Add missing canonical route coverage if independent review identifies it. MODIFY English/ja/ko/ru/zh-cn server configuration docs and structure owners. Avoid unsupported claims about models being weaker or provider safety bypass. + +Before/after anchor: createSseTerminalOutputBoundary() -> createSseTerminalOutputBoundary(options?: CodexSafetyBufferingFilterOptions); sanitizePassthroughHeaders(upstream) -> sanitizePassthroughHeaders(upstream, options?); canonical true => filter option, every other provider => undefined. Full public source diff is pinned by #3652 head in 000_plan.md and inspected locally; any needed correction is recorded here before B. + +## Independent design corrections + +H1 accepted: policy rewrite and hint stripping compose. Build policyFailurePayload first, then remove top-level safety_buffering from the effective emitted payload, preserving response.failed/error data and retryable:false. H2 accepted: extend current relaySseWithFailedTail fourth options object with terminalBoundary; never replace upstreamError. Core passes both existing upstreamError and new terminalBoundary; update the existing source-contract assertion to preserve its original guarantee. H3 accepted: native WebSocket codex.response.metadata.headers and /responses/compact are explicitly excluded; their hints remain unfiltered. No new WS metadata filter. Docs must not claim the old WS allowlist excludes these headers. Regression fixtures cover CRLF/split/malformed input, policy error plus hint, EOF upstreamError, canonical true and noncanonical preservation. + +## Pinned source hunks (apply with corrections above) + +```diff +diff --git a/src/config.ts b/src/config.ts +index fdcda9547c..cd0641feb2 100644 +--- a/src/config.ts ++++ b/src/config.ts +@@ -1125,6 +1125,8 @@ const configSchema = z.object({ + configRebaseProvenance: z.unknown().optional(), + // A retry can be billable, so absence and malformed hand edits both stay off. + emptyCompletionRetry: z.boolean().optional().catch(false), ++ // Header suppression changes what Codex sees, so absence and malformed edits stay off. ++ dropCodexSafetyBuffering: z.boolean().optional().catch(false), + // A malformed hand edit must not silently stop opening the browser: fall back + // to undefined, which resolves to the historical auto-open behavior. + oauthOpenBrowser: z.boolean().optional().catch(undefined), +@@ -2613,6 +2615,14 @@ function emptyCompletionRetryError(value: unknown): string | null { + return "schema_invalid: emptyCompletionRetry: must be a boolean or omitted"; + } + ++function dropCodexSafetyBufferingError(value: unknown): string | null { ++ const raw = rawConfigRecord(value); ++ if (!raw || !Object.hasOwn(raw, "dropCodexSafetyBuffering")) return null; ++ const enabled = raw.dropCodexSafetyBuffering; ++ if (enabled === undefined || typeof enabled === "boolean") return null; ++ return "schema_invalid: dropCodexSafetyBuffering: must be a boolean or omitted"; ++} ++ + function oauthOpenBrowserError(value: unknown): string | null { + const raw = rawConfigRecord(value); + if (!raw || !Object.hasOwn(raw, "oauthOpenBrowser")) return null; +@@ -2718,6 +2728,7 @@ export function validateConfigCandidate(value: unknown): { ok: true; config: Ocx + ?? codexQuotaAutoRefreshError(value) + ?? codexAccountPickerEnabledError(value) + ?? emptyCompletionRetryError(value) ++ ?? dropCodexSafetyBufferingError(value) + ?? oauthOpenBrowserError(value) + ?? runtimeRoleError(value) + ?? remoteGuiConfigError(value) +@@ -3684,6 +3695,7 @@ export function getDefaultConfig(): OcxConfig { + return { + port: 10100, + emptyCompletionRetry: false, ++ dropCodexSafetyBuffering: false, + managementUsageMaxReadBytes: 64 * 1024 * 1024, + appOwnedMemoryBudgetMb: DEFAULT_APP_OWNED_MEMORY_BUDGET_BYTES / (1024 * 1024), + // Fresh/re-initialized configs are already written in the current three-tier +diff --git a/src/server/index.ts b/src/server/index.ts +index aedd6bf236..c6ce73b2f1 100644 +--- a/src/server/index.ts ++++ b/src/server/index.ts +@@ -142,6 +142,7 @@ import { + } from "./relay"; + export { + consumeForInspection, ++ codexSafetyBufferingFilterOptions, + relaySseWithFailedTail, + relaySseWithHeartbeat, + relayWithAbort, +diff --git a/src/server/relay-eager.ts b/src/server/relay-eager.ts +index 655997b813..a6e60d3d02 100644 +--- a/src/server/relay-eager.ts ++++ b/src/server/relay-eager.ts +@@ -26,6 +26,7 @@ + + import { + adapterEofIncompleteFrame, ++ type CodexSafetyBufferingFilterOptions, + createSseTerminalOutputBoundary, + doneFrame, + failedTailFrame, +@@ -83,6 +84,8 @@ export type EagerRelayOptions = { + postCancelDrainBytes?: number; + /** Injectable clock for tests. */ + now?: () => number; ++ /** Client output boundary filters (Codex safety-buffering hints). */ ++ terminalBoundary?: CodexSafetyBufferingFilterOptions; + }; + + const DEFAULT_MAX_QUEUE_BYTES = 8 * 1024 * 1024; +@@ -111,7 +114,7 @@ export function relaySseEagerBounded( + const terminalEncoder = new TextEncoder(); + const adapterEofFrame = adapterEofIncompleteFrame(terminalEncoder); + const terminalSentinel = doneFrame(terminalEncoder); +- const terminalBoundary = createSseTerminalOutputBoundary(); ++ const terminalBoundary = createSseTerminalOutputBoundary(opts?.terminalBoundary); + const activeRewrite: SseBlockRewrite | undefined = hooks.rewriteBlocks + ?? (hooks.rewritePayload ? payloadRewriteAsBlockRewrite(hooks.rewritePayload) : undefined); + const encodeFailedTail = (error: unknown): Uint8Array | null => { +diff --git a/src/server/relay.ts b/src/server/relay.ts +index 60b57ea025..d840b2e59c 100644 +--- a/src/server/relay.ts ++++ b/src/server/relay.ts +@@ -162,7 +162,10 @@ export type SseTerminalOutputBoundary = { + * terminal, and drops every later block/byte. A premature [DONE] is held until + * a terminal arrives so clean EOF can synthesize one terminal and one sentinel. + */ +-export function createSseTerminalOutputBoundary(): SseTerminalOutputBoundary { ++export function createSseTerminalOutputBoundary( ++ options?: CodexSafetyBufferingFilterOptions, ++): SseTerminalOutputBoundary { ++ const dropSafetyBuffering = options?.dropCodexSafetyBuffering === true; + const decoder = new TextDecoder(); + const encoder = new TextEncoder(); + const framer = new BoundedSseFrameBuffer(MAX_INSPECTION_SSE_FRAME_BYTES); +@@ -181,6 +184,10 @@ export function createSseTerminalOutputBoundary(): SseTerminalOutputBoundary { + const payload = sseDataPayload(decoder.decode(frame.block)); + const isDone = payload === "[DONE]"; + const parsed = payload === null ? undefined : parseSsePayload(payload); ++ const safetyBuffering = dropSafetyBuffering && parsed !== undefined ++ ? codexSafetyBufferingBlockAction(parsed) ++ : "keep"; ++ if (safetyBuffering === "drop") continue; + const policyError = parsed !== undefined && isPolicyRewriteType(parsed) + ? cyberPolicyTerminalError(parsed) + : undefined; +@@ -189,7 +196,9 @@ export function createSseTerminalOutputBoundary(): SseTerminalOutputBoundary { + decoder.decode(frame.block), + policyFailurePayload(policyError, parsed), + )) +- : frame.block; ++ : safetyBuffering === "strip" ++ ? encoder.encode(stripCodexSafetyBufferingField(decoder.decode(frame.block), parsed)) ++ : frame.block; + if (isDone) { + done = true; + if (responsesTerminal) { +@@ -260,10 +269,11 @@ export function relaySseWithFailedTail( + body: ReadableStream, + upstream: AbortController, + onClientGone?: (reason?: unknown) => void, ++ boundaryOptions?: CodexSafetyBufferingFilterOptions, + ): ReadableStream { + const reader = body.getReader(); + const encoder = new TextEncoder(); +- const terminalBoundary = createSseTerminalOutputBoundary(); ++ const terminalBoundary = createSseTerminalOutputBoundary(boundaryOptions); + let closed = false; + const relayChunk = ( + controller: ReadableStreamDefaultController, +@@ -438,6 +448,29 @@ function isPolicyRewriteType(parsed: unknown): boolean { + return type === "response.failed" || type === "response.incomplete" || type === "error"; + } + ++/** ++ * Codex emits its safety-buffering hint in the SSE body as well as in headers: ++ * a `response.metadata` event whose `metadata.type` is `safety_buffering`, or a ++ * `safety_buffering` field on another event. The metadata event is dropped whole; ++ * the field is stripped so the carrying event is otherwise relayed unchanged. ++ */ ++function codexSafetyBufferingBlockAction(parsed: unknown): "keep" | "drop" | "strip" { ++ const root = asJsonRecord(parsed); ++ if (!root) return "keep"; ++ if (root.type === "response.metadata") { ++ const metadata = asJsonRecord(root.metadata); ++ if (metadata?.type === "safety_buffering") return "drop"; ++ } ++ return Object.hasOwn(root, "safety_buffering") ? "strip" : "keep"; ++} ++ ++function stripCodexSafetyBufferingField(block: string, parsed: unknown): string { ++ const root = asJsonRecord(parsed); ++ if (!root) return block; ++ const { safety_buffering: _safetyBuffering, ...rest } = root; ++ return replaceSseDataPayload(block, JSON.stringify(rest)); ++} ++ + function rewritePolicyTerminalBlock(block: string, payload: string): string { + const newline = block.includes("\r\n") ? "\r\n" : "\n"; + const rewritten = replaceSseDataPayload(block, payload); +@@ -1422,7 +1455,31 @@ export function consumeForResponseLogMetadata( + * body makes the caller (Codex) double-decode / truncate → "stream error" on every gpt passthrough. + * Drop encoding + hop-by-hop headers; relay everything else (content-type, etc.) verbatim. + */ +-export function sanitizePassthroughHeaders(upstream: Headers): Headers { ++export const CODEX_SAFETY_BUFFERING_HEADERS = [ ++ "x-codex-safety-buffering-enabled", ++ "x-codex-safety-buffering-faster-model", ++] as const; ++ ++const CODEX_SAFETY_BUFFERING_HEADER_SET: ReadonlySet = new Set(CODEX_SAFETY_BUFFERING_HEADERS); ++ ++export interface CodexSafetyBufferingFilterOptions { ++ /** ++ * Drop Codex safety-buffering hints: the `x-codex-safety-buffering-*` response ++ * headers and the `safety_buffering` SSE metadata event / field. Absent and ++ * `false` relay everything unchanged. ++ */ ++ dropCodexSafetyBuffering?: boolean; ++} ++ ++/** Resolve the passthrough header policy from the loaded config (absent means "forward everything"). */ ++export function codexSafetyBufferingFilterOptions( ++ config: { dropCodexSafetyBuffering?: boolean }, ++): CodexSafetyBufferingFilterOptions { ++ return { dropCodexSafetyBuffering: config.dropCodexSafetyBuffering === true }; ++} ++ ++export function sanitizePassthroughHeaders(upstream: Headers, options?: CodexSafetyBufferingFilterOptions): Headers { ++ const dropSafetyBuffering = options?.dropCodexSafetyBuffering === true; + const DROP = new Set([ + "content-encoding", + "content-length", +@@ -1439,7 +1496,10 @@ export function sanitizePassthroughHeaders(upstream: Headers): Headers { + ]); + const out = new Headers(); + upstream.forEach((value, key) => { +- if (!DROP.has(key.toLowerCase())) out.set(key, value); ++ const lower = key.toLowerCase(); ++ if (DROP.has(lower)) return; ++ if (dropSafetyBuffering && CODEX_SAFETY_BUFFERING_HEADER_SET.has(lower)) return; ++ out.set(key, value); + }); + return out; + } +diff --git a/src/server/responses/core.ts b/src/server/responses/core.ts +index 9d0eea0d76..e199917968 100644 +--- a/src/server/responses/core.ts ++++ b/src/server/responses/core.ts +@@ -304,6 +304,7 @@ import { + markEagerRelaySseResponse, + markNativePassthroughSseResponse, + relaySseWithFailedTail, ++ codexSafetyBufferingFilterOptions, + relayWithAbort, + sanitizePassthroughHeaders, + } from "../relay"; +@@ -3850,6 +3851,9 @@ async function handleResponsesInner( + let hostAdmissionLease = pendingHostAdmissionLease; + pendingHostAdmissionLease = null; + try { ++ const codexSafetyBufferingOptions = isCanonicalOpenAiForwardProvider(route.provider) ++ ? codexSafetyBufferingFilterOptions(config) ++ : undefined; + const imageGenCallAliases = route.provider.authMode === "forward" + ? new Map() + : imageGenToolCallAliases(toolBridgeMaps.toolNsMap, parsed._rawBody, translatorBudget); +@@ -4732,7 +4736,7 @@ async function handleResponsesInner( + } + break; + } +- const headers = sanitizePassthroughHeaders(upstreamResponse.headers); ++ const headers = sanitizePassthroughHeaders(upstreamResponse.headers, codexSafetyBufferingOptions); + const resolvedModel = headers.get("openai-model")?.trim(); + if (resolvedModel && !logCtx.preserveResolvedModelFromRoute) logCtx.resolvedModel = resolvedModel; + if (isUsageDebugEnabled()) { +@@ -4824,7 +4828,7 @@ async function handleResponsesInner( + return new Response(upstreamResponse.body, { + status: upstreamResponse.status, + statusText: upstreamResponse.statusText, +- headers: sanitizePassthroughHeaders(upstreamResponse.headers), ++ headers: sanitizePassthroughHeaders(upstreamResponse.headers, codexSafetyBufferingOptions), + }); + } + if (!upstreamResponse.ok) { +@@ -5027,6 +5031,7 @@ async function handleResponsesInner( + onDone: () => unregisterTurn(turnAc), + }, { + clientGoneSignal: options.abortSignal, ++ terminalBoundary: codexSafetyBufferingOptions, + ...(inlineEagerRewrite ? { rewriteBudget: translatorBudget } : {}), + }); + // When selected, this relay closes response.completed even if upstream +@@ -5110,7 +5115,8 @@ async function handleResponsesInner( + const rewrittenBody = clientBlockRewrite !== undefined + ? relaySseWithBlockRewrite(nativeBody, clientBlockRewrite, translatorBudget) + : nativeBody; +- const clientBody = relaySseWithFailedTail(rewrittenBody, upstream, reason => clientGone.abort(reason)); ++ const clientBody = relaySseWithFailedTail(rewrittenBody, upstream, reason => clientGone.abort(reason), ++ codexSafetyBufferingOptions); + return markNativePassthroughSseResponse(new Response(clientBody, { + status: upstreamResponse.status, + headers, +@@ -5238,7 +5244,7 @@ async function handleResponsesInner( + } + throw error; + } +- const sseHeaders = sanitizePassthroughHeaders(headers); ++ const sseHeaders = sanitizePassthroughHeaders(headers, codexSafetyBufferingOptions); + sseHeaders.set("content-type", "text/event-stream"); + sseHeaders.set("cache-control", "no-store"); + return new Response(stream, { +diff --git a/src/types/config.ts b/src/types/config.ts +index 8cf1246979..4d2c63fdf1 100644 +--- a/src/types/config.ts ++++ b/src/types/config.ts +@@ -335,6 +335,16 @@ export interface OcxConfig { + client?: OcxClientConnectionConfig; + /** Opt in to one identical-turn retry when a Responses completion has no text or tool call. */ + emptyCompletionRetry?: boolean; ++ /** ++ * Drop the Codex safety-buffering hints from a Codex Responses passthrough: the ++ * `x-codex-safety-buffering-*` response headers, `response.metadata` SSE events of ++ * type `safety_buffering`, and the `safety_buffering` field on other SSE events. ++ * The Codex TUI turns those hints into a "retry with a faster model" prompt whose ++ * default action switches the session to a weaker model, so an unattended session ++ * can lose its model to a stray keystroke. Absent and `false` relay everything ++ * unchanged. ++ */ ++ dropCodexSafetyBuffering?: boolean; + /** + * Whether a login may open a browser on the machine running the proxy. + * + +``` + +## Hint P revalidation + +Previous D: presentation source complete; final hosted CI remains in delivery. This branch starts from the common docs checkpoint bd34120180 and baseline product 69e3dcda. Original #3652 does not apply cleanly because relay upstreamError handling changed. Carry nonconflicting hunks and manually adapt relay/core/config hunks, preserving cancellation and error capture. Independent H1-H3 plan reflection ALIGNED remains applicable. diff --git a/devlog/_plan/260912_thinking_contract/030_spark.md b/devlog/_plan/260912_thinking_contract/030_spark.md new file mode 100644 index 0000000000..3b013f75aa --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/030_spark.md @@ -0,0 +1,74 @@ +# Spark Lite metadata follows body shape + +Class C3 bounded compatibility. Independent of presentation/hint; depends on roadmap. MODIFY src/adapters/openai-responses.ts only inside canonical OpenAI forwarding and final wire model gpt-5.3-codex-spark. Add bodyCarriesLiteToolShape next to existing tool-shape helpers: Array.isArray(body.input) && body.input.some(item => isPlainObject(item) && item.type === "additional_tools" && Array.isArray(item.tools) && item.tools.length > 0). After final Spark body construction, delete all case variants of CODEX_RESPONSES_LITE_HEADER then set it to liteShaped ? "true" : "false". Existing prepareCodexWsRequest projects it onto native frame metadata. + +Before: Spark deletes the header, allowing stale native metadata to survive. After: tool-less/top-level-tool Spark frames advertise false; nonempty Lite catalog frames advertise true despite conflicting inherited header. No retirement, no changes to model availability, no user service changes. + +MODIFY tests/codex-integration/codex-metadata-integrity.test.ts: alias resolved final model, inherited true/false/mixed-case/absent header, Lite tool body true, empty Lite group false, malformed metadata keeps HTTP fallback/body, noncanonical remains unchanged. MODIFY tests/responses/ws-upstream-reuse.test.ts: legacy true socket retires when adapter produces false, replacement same identity reused, raw request immutable. MODIFY all eight existing architecture locale pages and structure/transports/responses.md, referencing body-shape rule from shared area owners. Adopt latest #4130 source diff, preserving author; do not import historical earlier heads. + +Verification: source diff review and final-branch hosted ci.yml lane=all. Tests NOT RUN locally. Success proves framing and connection identity, not a live backend EOF fix or all tool-bearing EOF cases. Remaining acceptance: broader tool-format conversion stays out of scope. + +## Pinned source hunks (apply with corrections above) + +```diff +diff --git a/src/adapters/openai-responses.ts b/src/adapters/openai-responses.ts +index c4aa523ee6..8fbe43816d 100644 +--- a/src/adapters/openai-responses.ts ++++ b/src/adapters/openai-responses.ts +@@ -864,6 +864,21 @@ function promoteClientLoadedTools(body: unknown): unknown { + } + + const MAX_RESPONSES_CALL_ID_LENGTH = 64; ++ ++/** ++ * Whether the outgoing body still delivers tools through the responses-lite shape. ++ * ++ * Lite carries the client catalog as an `additional_tools` input item; the non-Lite wire shape ++ * expects top-level `tools`. Anything that flips the Lite advertisement has to agree with the ++ * shape actually being sent, or the destination silently loses the tool surface. ++ */ ++function bodyCarriesLiteToolShape(body: Record): boolean { ++ if (!Array.isArray(body.input)) return false; ++ return body.input.some(item => ++ isPlainObject(item) && item.type === "additional_tools" ++ && Array.isArray(item.tools) && item.tools.length > 0 ++ ); ++} + const REPAIRED_CALL_ID_PREFIX = "call_ocx_"; + const REPAIRED_CALL_ID_DIGEST_LENGTH = MAX_RESPONSES_CALL_ID_LENGTH - REPAIRED_CALL_ID_PREFIX.length; + +@@ -2515,12 +2530,22 @@ export function createResponsesPassthroughAdapter(provider: OcxProviderConfig): + parsed.modelId, + ); + if (isCanonicalOpenAiForwardProvider(provider)) { +- // Spark closes Responses Lite streams before a terminal completion. Select compatibility +- // from the final wire model so aliases cannot leave the caller or a static header enabled. ++ // Select Spark's Lite compatibility from the final wire model, including aliases, and ++ // let the BODY decide it. The header also overrides native WS metadata downstream, so a ++ // forwarded or statically configured value must never contradict the shape being sent. ++ // ++ // The synchronized catalog keeps `use_responses_lite: true` for Spark precisely because ++ // it selects tool delivery (`input[].additional_tools` instead of top-level `tools`), and ++ // stripSparkCompatibility filters that group in place rather than promoting it. So a ++ // Lite-shaped body is pinned back ON — otherwise an inherited `false` advertises non-Lite ++ // while the tools exist only in the Lite shape, and Spark loses the tool surface. Only a ++ // body with no Lite tool group is downgraded, which is what the stream fix needs. + if (isPlainObject(finalBody) && finalBody.model === "gpt-5.3-codex-spark") { ++ const liteShaped = bodyCarriesLiteToolShape(finalBody); + for (const name of Object.keys(headers)) { + if (name.toLowerCase() === CODEX_RESPONSES_LITE_HEADER) delete headers[name]; + } ++ headers[CODEX_RESPONSES_LITE_HEADER] = liteShaped ? "true" : "false"; + } + const routingHeaders = new Headers(headers); + applyCodexRoutingHint(routingHeaders, finalBody); + +``` + +## Spark P revalidation + +Prior D: hint source/security review PASS, final hosted tests pending. This independent branch starts from bd34120180. Latest #4130 hunks still apply cleanly. CCA summary and hint branches do not modify this adapter. Body-dependent Lite true/false, canonical final wire model and no retirement remain the acceptance contract. + +## Spark design reflection amendments + +S1 accepted: apply the existing modelSuffixBracketStrip normalization to finalBody before deciding Lite, using the same immutable object serialized later. A canonical gpt-5.3-codex-spark[1m] request that strips to Spark gets the policy; a final non-Spark model does not. S2 accepted: all eight architecture paragraphs say nonempty additional_tools tools array, not merely group presence; source PR outstanding documentation finding is addressed. S3 accepted: tests cover catalog filtered empty, surviving functions group, only top-level tools, both alias directions and preserved noncanonical configured Lite. Check both actual serialized body and WS header metadata; shape detection is not tool-support validation. diff --git a/devlog/_plan/260912_thinking_contract/040_delivery.md b/devlog/_plan/260912_thinking_contract/040_delivery.md new file mode 100644 index 0000000000..a70e62bfb6 --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/040_delivery.md @@ -0,0 +1,7 @@ +# Final heads and handoff + +Class C3 delivery evidence. Depends on all dispositions. MODIFY branch-owned numbered completion docs and ignored .tmp/thinking/handoff.md. Read existing .github/PULL_REQUEST_TEMPLATE.md; write every section, credits and precise NOT RUN limitation. Publish only own codex/260912-60plus-thinking* branches with git push --no-verify; PR bases dev for independent units, ordinary parent branch only for actual dependencies. No merge/auto-merge/closures. + +NEW .tmp/thinking/*-ci.json captures gh run view JSON for final SHA plus all jobs. NEW .tmp/thinking/*-review.md captures independent implementation findings with accepted/rebutted disposition. Refresh head/base, native stack membership (unknown if API unsupported), current reviews and CI before handoff. Inspect .github/workflows/ci.yml and dispatch lane=all at each final branch where needed. Existing automatic runs stay untouched. If final-head CI fails, inspect failing logs, repair scoped source or fixtures, commit/push --no-verify and validate new final tip. Do not label skipped/cancelled/old-head runs passing. + +Final handoff fields: own worktree, branch per PR, source PR disposition, exact head, PR URL, dependency order, original author trailers, remaining acceptance, unresolved reviews, CI run id/url/head/result/job conclusions, own cycle records and local tests NOT RUN. Parent performs any subsequent integration. No evidence claims from peer commentary alone. diff --git a/devlog/_plan/260912_thinking_contract/050_refresh.md b/devlog/_plan/260912_thinking_contract/050_refresh.md new file mode 100644 index 0000000000..16d4aa80dd --- /dev/null +++ b/devlog/_plan/260912_thinking_contract/050_refresh.md @@ -0,0 +1,5 @@ +# Integration conflict repair + +Parent explicitly requests own hint branch latest-dev integration with independent resolution-only audit and no-verify push. Latest fetched dev ca5ac39124671ee05349e7873231f672824ea26c also conflicts with presentation; preserve all three independent dev-based PRs. Most collisions are adjacent structure-document additions; core received continuation recovery changes that must survive. No source PR/other worktree modifications or merges into dev. Rebase only owned branches, preserve pre-rebase refs in ignored evidence and compare range-diff; use explicit expected old remote SHA with force-with-lease plus --no-verify. This is branch refresh, not native restacking. Product tests remain NOT RUN. + +Prior D: Spark source audit PASS; hosted tests remain pending. Refresh is a subtask of the already-active delivery cycle; no separate cycle is claimed. MODIFY conflict paths only, retaining source contracts and new dev changes. Verification: git range-diff, git diff --check, docs source validator, independent resolution-only review; hosted tests rerun only on refreshed final heads. diff --git a/devlog/_plan/260912_v2_contracts/000_plan.md b/devlog/_plan/260912_v2_contracts/000_plan.md new file mode 100644 index 0000000000..1646f58064 --- /dev/null +++ b/devlog/_plan/260912_v2_contracts/000_plan.md @@ -0,0 +1,33 @@ +# V2 delegation contracts + +This unit reconciles plaintext prevention (#2495) separately from encrypted task recovery (#3661). Eligible native parents may opt into plaintext V2 calls; recovery continues to use its existing authenticated, bounded path. The replacement candidates #4242/#4243 are compared against the exact issue contract before any adoption. + +Loop: satisfy-spec, triggered by the authorized v2 lane. Goal: scoped carry PRs and final cumulative hosted CI evidence. Non-goals: merges, issue closure, releases, installed service/config changes, native GitHub stacks, local product tests/build/typecheck/install. Local tests are NOT RUN by explicit instruction. Verification: source/diff checks during each cycle; Cross-platform CI on the final published head, with run IDs and conclusions retained. Stop: implementation, audit and CI evidence handed to the integration owner; no claim of integration. Outcomes: DONE with evidence, or an explicit unresolved acceptance/gate. Artifacts: this unit plus ignored `.tmp/v2/` evidence. Escalation: real tool denials and unresolved security/contract blockers are recorded; no new access/settings. Resource bounds: available account/tool permissions, this worktree only, no user token/time/agent-count cap. + +| Cycle | Outcome | Design | +|---|---|---| +| wp0 | Docs-only roadmap locked by independent design reflection and A review | this document | +| wp1 | Exact plaintext request/response contract and regression coverage | [010](010_plaintext.md) | +| wp2 | Bounded encrypted envelope handling and residual disposition | [020](020_recovery.md) | +| wp3 | Restore exact native collaboration dispatch identities | [030](030_native_identity.md) | +| wp4 | Final cumulative hosted verification and durable handoff | [040](040_verification.md) | + +wp1 and wp2 are distinct capabilities; execution order does not itself create a PR dependency. Use independent dev-based PRs if neither consumes the other's changes. A shared final cumulative verification branch may be needed to prove composition; do not silently call intermediate CI final-tip evidence. + +Existing owners: `src/adapters/openai-responses.ts`, `src/server/responses/core.ts`, `src/server/responses/agent-task-recovery.ts`; tests remain under domain directories. Source-of-truth pages are mapped by `structure/INDEX.md`. Reuse these owners, not a second server/recovery subsystem. Do-nothing/config-only alternatives cannot provide the missing wire behavior. + +Generic supported inherited-model subagents provide independent design consultation and separate review. Native architect selection is unavailable and is not claimed. Original contributor attribution follows the adopted source, including Sigurd-git for #2496 and SB Yoon if any #4242 code is carried. Source PRs/issues remain open or closed in their current state until the integration owner decides. + +## Cycle record + +wp0: P entered with own session binding; roadmap in progress. Product validation NOT RUN. + +wp0 A: Gauss GO-WITH-FIXES (blockers=0); WP1-A01 cache ordering and WP2-A01 fragment owner folded into decade docs. Pasteur reflection ALIGNED; generic inherited-model consultation, native architect not selected. + +wp0 check correction: initial D was refused because the roadmap task had not yet been marked done. The subsequent P command re-entered planning; no completed cycle is claimed for that attempt. Re-audit retains the unchanged independent verdict, and a fresh docs-only B/C/D closes the actual cycle after recording its task outcome. + +wp1 D: plaintext implementation published as #4351, static review findings resolved; local tests NOT RUN and hosted proof deferred. +wp2 D: independent multipart implementation published as #4364; static security review PASS. Exact-count, multiplicity, aggregate-byte and mutation regression code added. Token-split reconstruction and live backend fidelity remain issue acceptance, not claimed solved. +wp3 D: source inspection of native Codex at 095da4b7e8b70b01afb5c6131ef926dcb8c0d85d required exact namespace/name restoration. Implementation at 7dc0bf4ea6 received independent static PASS. The earlier helper-only expectations did not establish native dispatch compatibility. Final hosted validation is wp4. + +Disposition: #4242/#4243 were rejected as-is after contract audit; #2496 is the credited adaptation source. #2495 remains open pending integration/retention approval and backend canary judgment. #3661 remains partial. The two carry PRs are independent dev-based siblings; no manual dependency chain or native stack was introduced. Public source/reference facts only are recorded here; detailed security audit material stays in ignored scratch. diff --git a/devlog/_plan/260912_v2_contracts/010_plaintext.md b/devlog/_plan/260912_v2_contracts/010_plaintext.md new file mode 100644 index 0000000000..4551da65cb --- /dev/null +++ b/devlog/_plan/260912_v2_contracts/010_plaintext.md @@ -0,0 +1,38 @@ +# Plaintext V2 prevention + +Class C4 public wire/retention boundary; consumes wp0. Source proposal: #2496 at 1a4cb4aab14200ec2efa71aea00d2a55fc90aca7. Exact public patch is the starting implementation specification, ported to current owners below. #4242 and #4243 are alternatives, not automatically dependencies. + +| Action | Path | Before → after | +|---|---|---| +| NEW | `src/responses/plaintext-v2-agent-messages.ts` | no explicit canonical exception → #2496 request compiler and bounded restoration helper, corrected by D2–D4 below | +| MODIFY | `src/types/config.ts`, `src/config.ts` | absent flag → optional `plaintextV2AgentMessages?: boolean`, unset default; malformed reads drop only field, candidate writes reject | +| MODIFY | `src/types/request.ts` | absent route marker → optional request-local `_plaintextV2AgentMessages` | +| MODIFY | `src/adapters/base.ts`, `src/adapters/openai-responses.ts` | no alias capabilities → adapter-produced request-owned tool-name sets after canonical opt-in rewrite | +| MODIFY | `src/server/responses/core.ts` | direct native passthrough → final-route opt-in preparation, alias metadata refreshed after each build, restoration before client/cache on every JSON/SSE/WS path | +| MODIFY | `src/server/index.ts` | recovery-only warning → separate opt-in plaintext retention warning | +| NEW | `tests/responses/plaintext-v2-agent-messages.test.ts`, `tests/server/plaintext-v2-agent-messages-server.test.ts` | absent → port #2496 tests and add refusal/collision cases | +| MODIFY | `tests/server/config.test.ts`, `tests/server/agent-task-recovery.test.ts`, `tests/responses/ws-upstream.test.ts` | existing adjacent contracts → port applicable #2496 regression deltas | +| MODIFY | `scripts/test-layout/layout.json`, `tests/fixtures/test-layout-expected.json` | no new tests → register both new domain paths | +| MODIFY | English and zh-cn `guides/sub-agent-surface.md`, `reference/configuration/agents.md` under `docs-site/src/content/docs/` | recovery/V1 alternatives → config-only experimental plaintext contract and retention warning | +| MODIFY | applicable `structure/` owners from INDEX | current ownership descriptions → point to canonical plaintext contract without duplicating unrelated subsystem behavior | + +D1: preserve the explicit issue option; no management toggle or routed mirror catalog. +D2: `shouldPreparePlaintextV2AgentMessages`: true only for Responses wire, final canonical ChatGPT forward destination and default top-level collaboration catalog; additional_tools-only catalogs do not activate it. +D3: `preparePlaintextV2AgentMessages`: copy-on-write namespace + three tool aliases. Only `message.encrypted === true` is removed. Scan declaration/reference identity positions, including nested catalogs and qualified alias names, before any rewrite. Refuse all on any collision; foreign namespaces stay untouched. +D4: `restorePlaintextV2AgentMessageCalls*`: restore only request-generated identity capabilities. Preserve marker `encrypted_function_args: []`. Treat malformed JSON, unknown private identities, binding conflict and >10,000 identities as refusal. Bounded JSON returns 502; streams emit response.failed; refusal has no retry and no continuation write. Refresh state per turn/build; no connection/global alias state. + +Field chain: config type → config schema/save → config load/candidate validation → final route marker → adapter body serialization and AdapterRequest metadata → response restoration. Metadata is in-process only, never serialized as response fields or persisted with previous_response_id. Startup consumes config for warning. No new public state enum. + +Activation matrix: disabled/malformed flag, noncanonical/key/Anthropic/routed parent/V1/custom namespace unchanged; true canonical declaration rewritten without input mutation; each collision location leaves whole request unchanged; known aliases restored for JSON/SSE/WS and snapshots; same aliases under foreign namespace unchanged; malformed/overlimit/conflict terminal refused and not cached; next turn disabled and concurrent requests do not inherit prior metadata. Use hosted tests only; no live-account canary is claimed. + +Guard strength: runtime explicit option + compiler/restorer are code-path controls; operator can disable the option, which selects ordinary encrypted behavior. No credential authorization is added. Residual undocumented upstream behavior and plaintext retention are documented, not described as encryption guarantees. + +Port mapping verified against current tree: old `tests/config.test.ts` is now `tests/server/config.test.ts`; old `tests/ws-upstream.test.ts` is now `tests/responses/ws-upstream.test.ts`. A `git apply --check` of supporting #2496 hunks fails at current adapter/config/startup context; manual semantic port is required, not blind cherry-pick. Pure helper and new tests can use their full public source bodies with adjusted imports. Existing test helper `repo-root.ts` supplies repository paths instead of legacy relative directory inference. + +Exact integration replacements: `refreshRoutedNamespaceToolAliases` at core line 4702 becomes `refreshRequestToolAliases`, assigning both alias sets from each AdapterRequest or fresh empty sets. All seven current callers are renamed. At core `rememberPassthroughResponseChecked`, change `const restoredResponse = normalized...` to an intermediate normalized value; run plaintext restoration and return immediately on refusal before `rememberPassthroughResponse`. In blockRewrites, insert plaintext restoration immediately after `createResponsesSnapshotBlockRewrite`, before field backfill and guard. In bounded JSON, apply restoration after `normalizeFunctionCompletionJson` and before model rewrite; a refusal short-circuits before `rememberPassthroughResponseChecked`. At unsupported passthrough fallback, cancel response body and return safe 502 whenever request alias sets are nonempty. + +A synthesis WP1-A01: early raw inspection cannot authorize continuation for plaintext turns. Disable its cache callbacks when aliases are active. Publish only from a post-restoration/post-guard client block observer, and only after shared request-local stream validation accepts the terminal. Bounded JSON caches only its final restored value. Stream malformed/conflicting/overlimit rejection permanently prevents publication. + +WP1 implementation-P revalidation: prior D locked the roadmap. Reuse `createSseInspector` for restored client blocks: append a final block observer after the alias restorer and undeclared-tool guard, feed `${block}\n\n`, and dispose with the composed rewrite. Raw inspector callbacks are suppressed only for active plaintext aliases. The existing collector reconstructs output from accepted events. This gives one validated publication path rather than parallel raw/client cache decisions. Keep the final marker-preserving restorer before this collector. Unknown private identities throw before collection; terminal-only valid complete snapshots are accepted, so there is no invented requirement for prior added frames. + +Exact source-of-truth canonical owner is `structure/subagents.md`; add concise links from the mapped affected owners `runtime.md`, `config.md`, `overview.md`, `catalog.md`, `transports/responses.md`, `transports/streaming-health.md`, `transports/inventory.md`, `data-planes/images.md`, `data-planes/inbound-compat.md`, `providers/openai-tiers.md`, `providers/cursor.md`, `providers/chat-compat.md`, `providers/kiro.md`, `providers/xai-grok.md`, `adapters/registry.md`, `gui-and-management-api.md`, `clients/claude-desktop.md`, `ops/service-and-sidecars.md`, and `ops/docs-and-release.md` where the source-area map requires same-change synchronization. Links distinguish unchanged surfaces from the canonical new contract. diff --git a/devlog/_plan/260912_v2_contracts/020_recovery.md b/devlog/_plan/260912_v2_contracts/020_recovery.md new file mode 100644 index 0000000000..7dbcb075db --- /dev/null +++ b/devlog/_plan/260912_v2_contracts/020_recovery.md @@ -0,0 +1,50 @@ +# Encrypted envelope recovery + +Class C4 authenticated plaintext boundary; consumes roadmap and independent envelope design. #3794 diagnostics and MESSAGE support are already present. This phase preserves them and never adds automatic outage retries. + +| Action | Path | Before → after | +|---|---|---| +| MODIFY | `src/server/responses/agent-task-recovery.ts` | single encryptedIndex/ciphertext → ordered bounded part descriptors and exact envelope snapshot; single backend recovery request; atomic input revalidation before replacement | +| MODIFY | `tests/server/agent-task-recovery.test.ts` | single-part coverage → ordered multipart, invalid/ambiguous fragments, size/count cap, input mutation and cache isolation cases | +| MODIFY | `docs-site/src/content/docs/reference/configuration/agents.md` | narrow recovery description → exact supported multipart shape, no blind retries and residual fragment limitations | +| MODIFY | relevant `structure/` owners | current single-part invariant → canonical bounded envelope contract | + +D5 proposal for design audit: accept a contiguous run of complete structurally valid Fernet strings, at most 32 parts and 2 MiB combined. Keep routing header singular and author/recipient equal to sender/task. Forward original complete token parts in their order to the same fixed backend endpoint once. Partial token strings remain unsupported unless source evidence establishes an unambiguous join contract; do not infer authentication from a plausible Fernet shape. + +`AgentEnvelope` replaces encryptedIndex with an ordered part list. The cache key hashes a length-delimited serialized token array (not ambiguous string concatenation). `recoveryPayload` maps these parts into its one input message. `injectAssignment` reruns envelope parsing and compares the full admitted snapshot (header, identities, positions, all ciphertext parts) before one content splice, then removes agent routing identity fields exactly as today. Existing admission is still before every cache read. Input mutation causes input_changed and cache discard. + +Creation → serialization → consumption: parser builds ordered part descriptors; recoveryPayload emits each validated whole part; cache key binds their order and boundaries; injection validates the original current input and writes one assignment. No new stored config or failure enum is needed; unsupported_envelope remains not attempted and existing typed request failures remain attempted/capacity outcomes. + +Activation matrix: one complete part unchanged; two complete ordered parts reach exactly one mocked backend call and one plaintext replacement; swapped tokens have distinct cache identity; wrong sender/recipient/header rejected without fetch; interleaved plaintext/noncontiguous encrypted parts rejected; empty, malformed, excessive count or total bytes rejected; delayed input mutation refuses assignment; HTTP 5xx yields the existing typed reason after one call; no retry budget increase. Reuse existing helper fixtures; no test execution locally. + +Fragment disposition: this unit does not concatenate split tokens. #3661 contains no fragment association or representation evidence. The runtime's existing plaintext-in-encrypted-slot compatibility must remain. Add end-to-end regression coverage for a consecutive encrypted run whose exact concatenation is structurally one Fernet token: classify that narrowly as unreadable and refuse without recovery, while ordinary plaintext slots still normalize. If no sound discriminator is found, retain the issue residual explicitly; never claim full #3661 closure from whole-token support. + + +Concrete replacement contract: + +```ts +// AgentEnvelope +// - encryptedIndex: number; ciphertext: string; +// + encryptedStartIndex: number; ciphertexts: readonly string[]; +// + inputSnapshot: string; +// Parser: collect {index, token} only when token list has exactly one member +// and that member === raw encrypted_content. Reject missing header, +// >32 entries, >2 MiB aggregate, and nonconsecutive indexes. Capture +// JSON.stringify(item) at admission after all identity checks. +// Cache replaces .update(envelope.ciphertext) with +.update(JSON.stringify(envelope.ciphertexts)) +// Fixed recovery endpoint content replaces its single encrypted part with +...envelope.ciphertexts.map(encrypted_content => ({ + type: "encrypted_content", encrypted_content, +})) +// Injection verifies original item bytes before touching content: +if (JSON.stringify(item) !== envelope.inputSnapshot) return false; +content.splice(envelope.encryptedStartIndex, envelope.ciphertexts.length, + { type: "input_text", text: assignment }); +``` + +The snapshot is request-local and not logged/persisted. JSON request parsing is the input boundary, so getters/cycles are not supported client states. Tests use the existing Request/recovery public entrypoints, not exported parser internals. + +Reflection amendment: also MODIFY `src/server/responses/encrypted-payload.ts` only for the narrow multi-slot discriminator and MODIFY `tests/server/agent-task-recovery.test.ts` with `post()` integration assertions that recovery is not attempted and routed fetch is absent. Whole-token recovery tests remain at the recovery API. The discriminator runs before sanitization; for otherwise unreadable envelopes, matched fragments do not reach the routed provider. General malformed payload detection remains outside this claim. + +wp2 reflection synthesis: preserve only identified fragment objects during sanitization, not an entire content array. Independent plaintext slots still normalize. Fragment refusal applies only when no independent readable task text remains, retaining current mixed-content policy; mixed input is explicitly outside the refusal claim. All encrypted slots in a recovery envelope must be valid consecutive whole tokens, including malformed non-string slots (which refuse). MODIFY `tests/server/agent-task-recovery-security.test.ts`: replace formerly unsupported duplicate-whole-token fixture with a genuinely noncontiguous encrypted run; keep fragment and admission-negative coverage, add positive multipart regression separately. diff --git a/devlog/_plan/260912_v2_contracts/030_native_identity.md b/devlog/_plan/260912_v2_contracts/030_native_identity.md new file mode 100644 index 0000000000..14fa64a80c --- /dev/null +++ b/devlog/_plan/260912_v2_contracts/030_native_identity.md @@ -0,0 +1,13 @@ +# Native plaintext tool identity correction + +Prior D: wp2 source and static audit complete, hosted acceptance pending. Final consumer tracing found a missing identity component; split correction from final hosted verification rather than accepting helper-only mock expectations. + +MODIFY `src/responses/plaintext-v2-agent-messages.ts`: private bare and qualified aliases in calls/selectors must restore both `namespace: "collaboration"` and the unqualified declared child name. Namespace-member declarations restore only their child name, without injecting a redundant namespace field. Foreign namespaces remain untouched. Add an explicit namespace-member traversal context so declarations and selectors are not conflated. + +Before: an unqualified `start_delegated_task` becomes bare `spawn_agent`, or a qualified private name becomes `collaboration__spawn_agent`. After: a call becomes `{namespace:"collaboration",name:"spawn_agent"}`, preserving encrypted_function_args. A declaration inside restored namespace has `{type:"function",name:"spawn_agent"}`. + +MODIFY `tests/responses/plaintext-v2-agent-messages.test.ts`, `tests/server/plaintext-v2-agent-messages-server.test.ts`, `tests/responses/ws-upstream.test.ts`: pin exact namespace+child identity for bare, dotted, double-underscore, JSON, SSE and WS restoration; assert a compatible namespace/name plus empty marker selects the documented native plaintext path. Keep foreign and opaque data negatives. + +MODIFY `structure/subagents.md`: canonical dispatch identity is namespace plus unqualified child name. Source authority: locally inspected upstream `protocol/src/tool_name.rs` constructor preserves name literally; with_default_namespace assigns functions to absent namespace. `core/src/tools/router.rs` direct_source requires collaboration plus exact spawn_agent/send_message/followup_task and empty marker. This is source evidence, not a live backend canary. + +No new settings or APIs; same request alias metadata and collision gates. Product checks remain hosted-only; local tests/build/typecheck/install NOT RUN. Independent design and A review precede code; final evidence remains wp4. This amendment adds work and does not remove any original acceptance requirement. diff --git a/devlog/_plan/260912_v2_contracts/040_verification.md b/devlog/_plan/260912_v2_contracts/040_verification.md new file mode 100644 index 0000000000..c93afe24a0 --- /dev/null +++ b/devlog/_plan/260912_v2_contracts/040_verification.md @@ -0,0 +1,11 @@ +# Final hosted verification and handoff + +Consumes published implementation heads. No product change is planned unless exact hosted failure or independent audit identifies a defect; then amend this design with the concrete source delta before repair. + +MODIFY this unit's cycle records with actual outcomes. MODIFY ignored `.tmp/v2/handoff.md` and NEW ignored `.tmp/v2/final-ci.json` with own branch/worktree/session, original dispositions, credit, carry PR URLs, exact heads, chain order if any, remaining issue acceptance, unresolved review/security judgments and local NOT RUN. + +Commands: `git diff --check` observes whitespace only. `gh pr view` observes live head/base/reviews. `gh run list --commit ` finds hosted runs; `gh run view --json headSha,status,conclusion,jobs,url` provides final evidence. Inspect `.github/workflows/ci.yml` or actual workflow source for full lane dispatch. Do not claim skipped/cancelled jobs passed. CI failure repairs are additional PABCD cycles when they form a separate work-phase. + +Before publish, inspect exact diff and original contributor commits; push only owned branches using `git push --no-verify`. Populate Summary/Verification/Checklist template honestly with NOT RUN local tests. No closure or merge. Capture remote PR head equality with local final SHA and final Cross-platform CI result. A source scan or receipt wrapper is not product test evidence. Independent review has a source SHA and limitations. Any live upstream canary absent remains explicit. + +Integration refresh: fetched dev81f0c78d7a after parent integrated other lanes. Read-only merge-tree previews identified only shared structure-document EOF additions as conflicts; source hunks combined without conflict. Move this lane's added paragraphs to separate existing section boundaries while preserving their bytes, then verify clean merge-tree previews for both independent PRs. No branch merge, rebase or force push is needed. Joint runtime composition remains the integration owner's verification duty. diff --git a/devlog/_plan/260912_v2_contracts/050_resume_evidence.md b/devlog/_plan/260912_v2_contracts/050_resume_evidence.md new file mode 100644 index 0000000000..3034547702 --- /dev/null +++ b/devlog/_plan/260912_v2_contracts/050_resume_evidence.md @@ -0,0 +1,11 @@ +# Resumed V2 verification + +The existing task resumed without reset or a replacement worktree. Its persisted phase remains C; the host goal remains blocked and has not been modified. Local suites, focused checks, builds, typechecks and installs remain NOT RUN. + +Remote correction `bda46fe8f0056fa8d1eaeceb04e04f54e204359b` already repaired nullable namespace handling and added pure/JSON/SSE/WS regressions. The clean task branch fast-forwarded to that commit; the correction was not recreated. Independent source/security re-review is required for this delta. + +Historical full CI runs 34675376969 (plaintext head 4319c1eae0) and 34675235190 (recovery head c9544c1445) concluded failure. The observed Windows failures concern Devin discovery/credential fixtures, pnpm generated shims, and, on the plaintext run, the live autostart-owner lock fixture. They are not declared flakes or a green baseline. Other lanes retain those fixes; this lane does not alter their files. Passing Linux/macOS jobs do not make either full run pass. + +Read-only merge-tree inspection against dev392e182a004d61b38c7cf652642e63b9a11d9a65 found a plaintext conflict only in `structure/transports/responses.md`, where the newly added nullable-namespace paragraph shared an append point. Move that exact paragraph beside this lane's existing plaintext contract. No product code, branch rebase, integration merge or foreign paragraph changes are needed for this collision. + +Both existing PRs remain open. Final hosted verification must use the new plaintext head after this documentation checkpoint; recovery's unchanged failed head is not blindly rerun. Detailed logs, reviewed hashes, CI run identities, and current blockers remain in the task-local ignored handoff. diff --git a/docs-site/src/content/docs/fr/guides/combos.md b/docs-site/src/content/docs/fr/guides/combos.md index 9073372327..dc18f68971 100644 --- a/docs-site/src/content/docs/fr/guides/combos.md +++ b/docs-site/src/content/docs/fr/guides/combos.md @@ -218,20 +218,16 @@ Le basculement est intentionnellement limité. Il facilite la disponibilité, l' ## Effort de raisonnement par défaut -`defaultEffort` fournit `reasoning.effort` uniquement lorsque toutes ces conditions sont vraies : +`defaultEffort` complète un `reasoning.effort` absent si le combo possède une valeur par défaut non nulle et si la liste des niveaux acceptés par la cible est connue et non vide. La valeur configurée est conservée si elle est acceptée ; sinon, le niveau accepté le plus élevé ne la dépassant pas est choisi, ou le niveau le plus bas si aucun n’est inférieur. Une liste inconnue ou vide n’ajoute aucune valeur par défaut. -1. le combo a un défaut non nul ; -2. l'appelant n'a pas fait d'effort ; et -3. le catalogue de la cible sélectionnée annonce cet effort précis. +Cette étape conserve un effort existant et les autres champs reasoning. La normalisation des capacités ci-dessous peut supprimer séparément les paramètres effort/thinking non acceptés. Valeurs possibles : `low`, `medium`, `high`, `xhigh`, `max`, `ultra` ; l’absence du champ ou `null` désactive l’ajout. -Si la requête n'a pas d'objet `reasoning`, opencodex en crée un. Si `reasoning` existe sans -`effort`, il préserve les autres champs et ajoute la valeur par défaut. Un effort fourni par l’appelant n’est -jamais écrasé. -Lorsque la capacité cible est inconnue ou n'inclut pas l'effort configuré, opencodex omet le -par défaut et laisse le comportement de la cible inchangé. Les valeurs prises en charge sont `low`, `medium`, -`high`, `xhigh`, `max` et `ultra` ; omettez le champ ou réglez-le sur `null` pour laisser l'effort entièrement à -l'appelant et la cible. +## Capacités reasoning mixtes + +`reasoningEffortMode` vaut `"strict"` par défaut : le catalogue publie l’intersection des listes effort de toutes les cibles, y compris les listes explicitement vides. `"adaptive"` exclut ces listes vides pour conserver le sélecteur dans un combo mixte. Une liste inconnue ne limite l’intersection dans aucun des deux modes. + +À l’envoi, une liste explicitement vide supprime les paramètres effort et thinking dans les deux modes ; une liste inconnue les supprime uniquement en adaptive. `reasoning.summary` et les autres champs hors effort sont conservés. La résolution des cibles connues non vides reste inchangée. Les cibles inconnues en strict et les déclarations inconnues du native Chat ordinaire conservent les paramètres de l’appelant. L’ajout d’une valeur par défaut ne remplace pas un effort existant, mais cette normalisation peut supprimer les paramètres non pris en charge. ## Capacité d’entrée d’images / multimodale @@ -337,6 +333,7 @@ Les combos sont stockés dans l'objet `combos` de niveau supérieur, saisi par l | `strategy` | Non | `"failover"` | Valeurs autorisées : `"failover"`, `"round-robin"`, `"random"`, `"least-used"` et `"reset-window"`. | | `stickyLimit` | Non | `1` | Nombre entier de 1 à 100 requêtes réussies par sélection à tour de rôle. S’applique uniquement à `round-robin`. | | `defaultEffort` | Non | `null` | `low`, `medium`, `high`, `xhigh`, `max` ou `ultra` ; appliqué uniquement lorsque l'appelant omet ses efforts et que la cible annonce son soutien. | +| `reasoningEffortMode` | Non | `"strict"` | `strict` ou `adaptive` ; choisit l’intersection des capacités et la normalisation par cible. | | `imageInput` | Non | `"auto"` | `"auto"` ou `"disabled"`. `"auto"` publie les images uniquement si toutes les cibles les prennent en charge ; `"disabled"` impose le texte seul, retire les images des modalités publiées et rejette les requêtes qui en contiennent avant leur distribution. | | `alias` | Non | aucun | Identifiant de modèle public tronqué facultatif ; utilisez les règles d'alias ci-dessus. Une valeur vide est stockée sans alias. | | `nativeAlias` | Non | `false` | Autoriser explicitement un `alias` natif nu actuellement pris en charge à avoir la priorité sur le routage et le catalogue. Jamais déduit de l'alias. | diff --git a/docs-site/src/content/docs/fr/reference/architecture.md b/docs-site/src/content/docs/fr/reference/architecture.md index f197d85c11..746ea7f726 100644 --- a/docs-site/src/content/docs/fr/reference/architecture.md +++ b/docs-site/src/content/docs/fr/reference/architecture.md @@ -89,6 +89,15 @@ Par défaut, `server/index.ts` sert HTTP/SSE sur `/v1/responses`. Si Codex tente Indépendamment de ce réglage côté client, les requêtes canoniques transmises à ChatGPT avec `stream: true` à la racine peuvent utiliser le transport WebSocket en amont de Codex avec une version stable de Bun 1.4.0 ou ultérieure. La version intégrée Bun 1.3.14, les préversions et les identités de runtime impossibles à vérifier utilisent HTTP/SSE. Les réponses WS en amont qui réussissent conservent le contrat SSE en aval et contournent `tee()` au moyen d’un relais borné à lecteur unique et avide (4 MiB par trame brute/enveloppée et une file de production de 8 MiB). Le dépassement de la file ferme la connexion en amont et émet en aval un événement terminal `response.failed`, suivi de `[DONE]`. +Pour le modèle sortant final `gpt-5.3-codex-spark`, la transmission canonique à ChatGPT +désactive explicitement Responses Lite dans l’en-tête HTTP et les métadonnées natives des +trames WS, même lorsqu’un alias sélectionne Spark — uniquement si le corps sortant ne porte pas +de groupe `additional_tools` contenant un tableau `tools` non vide. Ce groupe EST la forme Lite de livraison des outils : un corps Spark +qui l’utilise conserve Lite ACTIF même si un en-tête appelant ou configuré disait l’inverse. Un changement d’identité Lite retire +l’ancien socket ; les requêtes admissibles suivantes ayant la même identité peuvent réutiliser +le nouveau socket. Les autres modèles et passerelles conservent leur politique Lite. +Des métadonnées natives mal formées entraînent toujours un repli HTTP, sans modifier le corps. + Le compactage du contexte Codex fonctionne avec les modèles routés. `server/responses/compact.ts` traite `POST /v1/responses/compact` en exécutant un tour interne de synthèse routé et en renvoyant un historique compacté, tandis que `responses/parser.ts` et `bridge.ts` traitent les tours de compactage distant v2 `compaction_trigger` en émettant exactement un élément de sortie synthétique `compaction`. ## Mise en cache et catalogue diff --git a/docs-site/src/content/docs/fr/reference/configuration/routing.md b/docs-site/src/content/docs/fr/reference/configuration/routing.md index 110d2fcd66..64713d645c 100644 --- a/docs-site/src/content/docs/fr/reference/configuration/routing.md +++ b/docs-site/src/content/docs/fr/reference/configuration/routing.md @@ -58,7 +58,8 @@ Chaque clé de combinaison est un identifiant conforme à `[A-Za-z0-9][A-Za-z0-9 | `targets` | `{ provider: string; model: string; weight?: number }[]` | requis | Routes concrètes ordonnées. `weight` est compris entre 1 et 10000 et vaut `1` par défaut. | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | Stratégie de sélection. L’ordre des cibles définit la priorité de `failover` ; les poids déterminent les sélections de `round-robin` et de `random` ; `least-used` suit les réussites enregistrées ; `reset-window` suit la réinitialisation de quota la plus proche. | | `stickyLimit?` | `number` | `1` | Nombre de requêtes réussies conservées dans un même lot de rotation. Plage de 1 à 100. | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | non défini | Appliqué uniquement lorsque l’appelant ne précise aucun effort et que la cible sélectionnée annonce le niveau demandé. | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | non défini | `defaultEffort` complète un `reasoning.effort` absent si le combo possède une valeur par défaut non nulle et si la liste des niveaux acceptés par la cible est connue et non vide. La valeur configurée est conservée si elle est acceptée ; sinon, le niveau accepté le plus élevé ne la dépassant pas est choisi, ou le niveau le plus bas si aucun n’est inférieur. Une liste inconnue ou vide n’ajoute aucune valeur par défaut. | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` calcule l’intersection des listes connues, y compris les listes vides ; `"adaptive"` exclut les listes vides. Les listes inconnues ne limitent l’intersection dans aucun des deux modes. À l’envoi, les listes explicitement vides suppriment les paramètres effort/thinking dans les deux modes ; les listes inconnues les suppriment seulement en adaptive. `reasoning.summary` est conservé. La résolution des listes connues non vides ainsi que le choix et l’ordre des cibles restent inchangés. | | `imageInput?` | `"auto" \| "disabled"` | `"auto"` | `"auto"` publie les images uniquement lorsque toutes les cibles les prennent en charge ; `"disabled"` impose le texte seul, retire les images des modalités publiées et rejette les requêtes qui en contiennent avant leur distribution. | | `alias?` | `string` | — | Identifiant public facultatif du modèle, à la place du slug canonique du sélecteur. | | `nativeAlias?` | `boolean` | `false` | Permet à un identifiant natif non qualifié actuellement pris en charge de prendre la priorité uniquement pour cet identifiant. Les identifiants non qualifiés `gpt-5.6-*` utilisent les identifiants Codex Pool/Direct. Les routes qualifiées par un compte restent distinctes. Les routes qualifiées par un fournisseur, telles que `openai-apikey/gpt-5.6-*`, utilisent la route configurée avec sa clé d’API et ne passent jamais par l’alias natif. | diff --git a/docs-site/src/content/docs/guides/combos.md b/docs-site/src/content/docs/guides/combos.md index ef5ccde16c..979a84ffcd 100644 --- a/docs-site/src/content/docs/guides/combos.md +++ b/docs-site/src/content/docs/guides/combos.md @@ -256,20 +256,10 @@ instead of growing memory without a bound. ## Default reasoning effort -`defaultEffort` supplies `reasoning.effort` only when all of these are true: +`defaultEffort` fills an absent `reasoning.effort` when the combo has a non-null default and the selected target has a known, nonempty supported ladder. If the target supports the configured value, it is retained; otherwise the highest supported rung at or below it is used, or the lowest supported rung when none is lower. Unknown or empty ladders omit the default. -1. the combo has a non-null default; -2. the caller did not set an effort; and -3. the selected target's catalog advertises that exact effort. +The default-injection step preserves existing effort and other reasoning fields. Capability normalization can separately remove unsupported effort/thinking controls as described below. Supported defaults are `low`, `medium`, `high`, `xhigh`, `max`, and `ultra`; omit the field or use `null` to disable default injection. -If the request has no `reasoning` object, opencodex creates one. If `reasoning` exists without an -`effort` property, it preserves the other fields and adds the default. A caller-provided effort is -never overwritten. - -When target capability is unknown or does not include the configured effort, opencodex omits the -default and leaves the target's own behavior unchanged. Supported values are `low`, `medium`, -`high`, `xhigh`, `max`, and `ultra`; omit the field or set it to `null` to leave effort entirely to -the caller and target. ### Mixed-capability groups (`reasoningEffortMode`) @@ -297,9 +287,12 @@ as wildcards in both modes. } ``` -The default is `"strict"`, which keeps the original behavior. This setting changes published -catalog metadata only — it does not change target order, failover policy, or which effort a given -target receives at dispatch. In the dashboard it is the **Adaptive reasoning ladder** switch in a +The default is `"strict"`, which keeps the original picker behavior. This setting does not change +target order or failover policy. At dispatch, an explicitly empty target ladder has its unsupported +effort/thinking controls removed in either mode while preserving supported non-effort reasoning fields +such as `reasoning.summary`; `"adaptive"` applies the same normalization to an unknown target +capability, while known non-empty targets keep their existing per-target effort resolution. +In the dashboard it is the **Adaptive reasoning ladder** switch in a combo's Capabilities section. ## Image / multimodal capability @@ -414,7 +407,7 @@ Combos are stored in the top-level `combos` object, keyed by combo id: | `cooldownMs` | No | unset → upstream fallback (5 s for request-rate 429 codes `1302`/`1305`, otherwise 60 s) | Integer from 1 to 600000. When set, applies as the per-target cooldown whenever no usable upstream `Retry-After` or Codex reset signal exists, including request-rate 429s; when unset, uses the upstream fallback. | | `waitForCooldownMs` | No | `0` | Integer from 0 to 600000. Maximum time to wait for the earliest eligible cooling target before returning `combo_unavailable`; abort cancels the wait. | | `defaultEffort` | No | `null` | `low`, `medium`, `high`, `xhigh`, `max`, or `ultra`; applied only when the caller omits effort and the target advertises support. | -| `reasoningEffortMode` | No | `"strict"` | `"strict"` intersects every known target ladder, so one target advertising no effort control empties the combo's picker. `"adaptive"` excludes those empty ladders from the published intersection. Metadata only; dispatch is unchanged. | +| `reasoningEffortMode` | No | `"strict"` | `"strict"` intersects every known target ladder, so one target advertising no effort control empties the combo's picker. `"adaptive"` excludes those empty ladders from the published intersection. At dispatch, explicit empty or adaptive unknown ladders remove unsupported effort/thinking controls while preserving supported non-effort reasoning fields such as `reasoning.summary`; known non-empty targets keep existing effort resolution. | | `imageInput` | No | `"auto"` | `"auto"` or `"disabled"`. `"auto"` publishes image support only when every target supports images; `"disabled"` forces text-only (drops image from published modalities and rejects image-bearing requests before dispatch). | | `alias` | No | none | Optional trimmed public model id; use the alias rules above. An empty value is stored as no alias. | | `nativeAlias` | No | `false` | Explicitly permit a currently supported bare native `alias` to take routing and catalog precedence. Never inferred from the alias. | diff --git a/docs-site/src/content/docs/guides/sub-agent-surface.md b/docs-site/src/content/docs/guides/sub-agent-surface.md index 49faeb9370..912268325f 100644 --- a/docs-site/src/content/docs/guides/sub-agent-surface.md +++ b/docs-site/src/content/docs/guides/sub-agent-surface.md @@ -328,3 +328,17 @@ tier that Codex converts to `max`; opencodex then maps or clamps the value for t The model context cap is independent of sub-agent mode. Configure it on the Models page; native OpenAI models retain their real context windows. + +The experimental `plaintextV2AgentMessages` field is unset in a fresh config and runs only when set +to `true`. The caller must use the Responses wire, and the final destination must use +`adapter: "openai-responses"`, `authMode: "forward"`, and the exact base URL +`https://chatgpt.com/backend-api/codex`. OpenAI API-key providers, custom compatible gateways, +routes to other providers, and non-Responses callers are excluded. For an eligible new native +ChatGPT v2 tool call, the option assigns request-scoped aliases to the namespace and three reserved +message-tool names, removes the message marker, and restores the original identities in the +response. It handles +`spawn_agent`, `send_message`, and `followup_task` and adds no recovery request. HTTPS remains +encrypted, but task text can be retained in Codex history, routed-provider requests, and local +response/debug state. Existing ciphertext is unchanged, and the option depends on undocumented +ChatGPT and Codex behavior. See +[Agent configuration: Plaintext v2 agent messages](/reference/configuration/agents/#plaintext-v2-agent-messages). diff --git a/docs-site/src/content/docs/ja/guides/combos.md b/docs-site/src/content/docs/ja/guides/combos.md index 70c902ec6e..7b2c62a9ec 100644 --- a/docs-site/src/content/docs/ja/guides/combos.md +++ b/docs-site/src/content/docs/ja/guides/combos.md @@ -139,15 +139,16 @@ ocx combo set balanced \ ## デフォルトの推論負荷 -`defaultEffort` は、次のすべてが当てはまる場合にのみ `reasoning.effort` を提供します。 +`defaultEffort` は、コンボの既定値が null でなく、対象の対応リストが既知で空でない場合に、省略された `reasoning.effort` を補います。設定値に対応していればその値を使い、そうでなければ設定値以下で最も高い段階を選びます。それもなければ最も低い対応段階を使います。不明または空のリストでは既定値を省略します。 -1. コンボには null 以外のデフォルトがあります。 -2. 呼び出し側は努力を設定しませんでした。そして -3. 選択したターゲットのカタログは、その正確な取り組みを宣伝します。 +既定値の補完は既存の effort と他の reasoning フィールドを保持します。以下の capability 正規化は、別途、非対応の effort/thinking 制御を削除できます。設定可能な既定値は `low`、`medium`、`high`、`xhigh`、`max`、`ultra` です。省略または `null` で補完を無効にします。 -リクエストに `reasoning` オブジェクトがない場合、opencodex はオブジェクトを作成します。 `reasoning` が `effort` プロパティなしで存在する場合、他のフィールドは保持され、デフォルトが追加されます。呼び出し元が提供した努力は決し​​て上書きされません。 -ターゲットの機能が不明な場合、または設定されたエフォートが含まれていない場合、opencodex はデフォルトを省略し、ターゲット自体の動作を変更しないままにします。サポートされている値は、`low`、`medium`、`high`、`xhigh`、`max`、および `ultra` です。このフィールドを省略するか、`null` に設定して、呼び出し元とターゲットに作業を完全に任せます。 +## 異なる reasoning capability の組み合わせ + +`reasoningEffortMode` の既定値は `"strict"` です。明示的な空リストを含む全対象の effort リストの共通部分を公開します。`"adaptive"` は空リストを共通部分の計算から除外し、混在するコンボでも選択肢を維持します。不明なリストは、どちらのモードでもカタログの共通部分を制限しません。 + +送信時には、明示的な空リストの対象で effort と thinking の制御を両モードとも削除します。不明な対象で削除するのは adaptive のみです。`reasoning.summary` と effort 以外のフィールドは保持し、既知の空でない対象は従来どおり effort を解決します。strict の不明な対象と通常の native Chat の不明な宣言は保持します。既定値の補完は既存の effort を上書きしませんが、この正規化は非対応の制御を削除できます。 ## 暗号化された v2 サブエージェント タスク @@ -237,6 +238,7 @@ ocx combo remove --yes | `cooldownMs` |いいえ | 未設定 → アップストリーム フォールバック(リクエストレート 429 コード `1302`/`1305` では 5 秒、それ以外では 60 秒) | 1 ~ 600000 の整数。設定時は、使用可能なアップストリーム `Retry-After` または Codex リセットシグナルがない場合に、リクエストレート 429 を含むターゲットごとのクールダウンとして適用されます。未設定時はアップストリーム フォールバックを使用します。 | | `waitForCooldownMs` |いいえ | `0` | 0 ~ 600000 の整数。最も早く利用可能になる冷却中のターゲットを待ってから `combo_unavailable` を返すまでの最大待機時間。中止すると待機はキャンセルされます。 | | `defaultEffort` |いいえ | `null` | `low`、`medium`、`high`、`xhigh`、`max`、または `ultra`;呼び出し元が努力を省略し、ターゲットがサポートをアドバタイズした場合にのみ適用されます。 | +| `reasoningEffortMode` | いいえ | `"strict"` | `strict` または `adaptive`。混在する capability の共通部分と対象別の制御正規化を選択します。 | | `alias` |いいえ |なし |オプションのトリミングされたパブリック モデル ID。上記のエイリアス ルールを使用します。空の値はエイリアスなしで保存されます。 | | `nativeAlias` |いいえ | `false` | 現在サポートされている bare native alias に routing/catalog の優先権を明示的に与えます。 | | `displayName` |いいえ |なし | catalog 表示専用ラベル。`nativeAlias` が true の場合は必須です。 | diff --git a/docs-site/src/content/docs/ja/reference/architecture.md b/docs-site/src/content/docs/ja/reference/architecture.md index cd2ce3db56..3500cbaf86 100644 --- a/docs-site/src/content/docs/ja/reference/architecture.md +++ b/docs-site/src/content/docs/ja/reference/architecture.md @@ -97,6 +97,15 @@ HTTP の境界は `server/index.ts` が担い、Responses データプレーン `server/index.ts` はデフォルトで `/v1/responses` を HTTP/SSE で提供します。`websockets` が `false` の状態で Codex が Responses WebSocket アップグレードを試みると、opencodex は `426 upgrade_required` を返し、Codex はそのセッションで HTTP にフォールバックします。`"websockets": true` を設定すると同じエンドポイントがアップグレードを受け入れ WebSocket ブリッジを使います。 +最終送信モデルが `gpt-5.3-codex-spark` の場合、canonical ChatGPT 転送は HTTP ヘッダーと +ネイティブ WS フレームのメタデータの両方で Responses Lite を明示的に無効にします。 +エイリアスで Spark を選択した場合も同様です。ただし無効化は、送信本文が空でない `tools` 配列を持つ `additional_tools` +項目を持たない場合に限ります。このグループ自体が Lite のツール受け渡し形式なので、それを +使う Spark 本文は呼び出し元や設定のヘッダーに関わらず Lite を有効のまま保ちます。Lite の識別値が変わると古いソケットは退役し、 +以後の条件を満たす同じ識別値のリクエストは新しいソケットを再利用できます。他のモデルと +ゲートウェイの Lite ポリシーは維持されます。不正なネイティブメタデータは引き続き、 +本文を変更せずに HTTP にフォールバックします。 + Codex コンテキスト compaction はルーティングされたモデルでも動作します。`server/responses/compact.ts` は `POST /v1/responses/compact` を内部ルーティング要約ターンとして扱い、圧縮されたヒストリーを返します。 `responses/parser.ts` と `bridge.ts` は remote compaction v2 の `compaction_trigger` ターンを扱い、合成 `compaction` 出力項目を正確に 1 つ送ります。 diff --git a/docs-site/src/content/docs/ja/reference/configuration/routing.md b/docs-site/src/content/docs/ja/reference/configuration/routing.md index 18b238c8bf..c6a6d80772 100644 --- a/docs-site/src/content/docs/ja/reference/configuration/routing.md +++ b/docs-site/src/content/docs/ja/reference/configuration/routing.md @@ -69,7 +69,8 @@ picker catalog の convergence だけが保留中で routing change は失われ | `targets` | `{ provider: string; model: string; weight?: number }[]` |必須 |具体的なルートを指示しました。 `weight` は 1 ~ 10000 で、デフォルトは `1` です。 | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` |選択戦略。ターゲットの順序は `failover` の優先順位となり、`weight` は `round-robin` と `random` の抽選に影響し、`least-used` は記録された成功数に従い、`reset-window` は最も早いクォータリセットに従います。 | | `stickyLimit?` | `number` | `1` |成功したリクエストは 1 つのラウンドロビン バッチに保持されます。範囲は 1 ~ 100。 | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` |設定を解除する |呼び出し元が努力を省略し、選択されたターゲットが要求されたラングをアドバタイズする場合にのみ適用されます。 | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` |設定を解除する | `defaultEffort` は、コンボの既定値が null でなく、対象の対応リストが既知で空でない場合に、省略された `reasoning.effort` を補います。設定値に対応していればその値を使い、そうでなければ設定値以下で最も高い段階を選びます。それもなければ最も低い対応段階を使います。不明または空のリストでは既定値を省略します。 | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` は空リストを含む既知の対応リストの共通部分を公開し、`"adaptive"` は空リストを除外します。不明なリストは両モードで共通部分を制限しません。送信時、明示的な空リストは両モードで effort/thinking 制御を削除し、不明なリストでは adaptive のみ削除します。`reasoning.summary` は保持されます。既知の空でない対象の effort 解決と対象の選択・順序は変わりません。 | | `alias?` | `string` | — |正規のピッカー スラグの代わりのオプションのパブリック モデル ID。 | | `nativeAlias?` | `boolean` | `false` | 現在サポートされている bare native id に限り、その未修飾 id で優先します。アカウント修飾およびプロバイダー修飾の OpenAI ルートは別のままです。 | | `displayName?` | `string` | — | catalog 表示専用ラベル。native alias では空でない値が必須です。 | diff --git a/docs-site/src/content/docs/ja/reference/configuration/server.md b/docs-site/src/content/docs/ja/reference/configuration/server.md index 0dd9cf59e4..e86387f80e 100644 --- a/docs-site/src/content/docs/ja/reference/configuration/server.md +++ b/docs-site/src/content/docs/ja/reference/configuration/server.md @@ -13,6 +13,7 @@ description: リスナー、リモート アクセス、アドミッション | `hostname?` | `string` | `"127.0.0.1"` |バインドアドレス。非ループバック バインドには `OPENCODEX_API_AUTH_TOKEN` が必要です。 | | `proxy?` | `string` | — |送信 HTTP(S) プロキシ URL または `${ENV_VAR}`。これらの変数が設定されていない場合にのみ、`HTTP_PROXY` / `HTTPS_PROXY` に適用されます。ループバックは `NO_PROXY` に残ります。 | | `emptyCompletionRetry?` | `boolean` | `false` | テキストもツール呼び出しもない Responses ターンを、ターミナルイベント前にストリームが終了した場合も含め、同一リクエストで 1 回再試行するよう明示的に有効化します。再試行は課金対象になる場合があります。`OCX_EMPTY_COMPLETION_RETRY=0` で設定を変更せず無効化できます。combo と routed-compaction turn は対象外です。 | +| `dropCodexSafetyBuffering?` | `boolean` | `false` | Codex Responses パススルーから Codex の safety-buffering ヒントを除去します。対象は `x-codex-safety-buffering-enabled` / `x-codex-safety-buffering-faster-model` 応答ヘッダー、`safety_buffering` 型の `response.metadata` SSE イベント、およびその他の SSE イベントにある `safety_buffering` フィールドです。Codex TUI はこれらを、既定の操作でセッションをより弱いモデルに切り替える「より高速なモデルで再試行」プロンプトとして表示します。その他の `x-codex-*` ヘッダーと SSE イベントの内容は、そのフィールドの除去を除いて変更せずに転送されます。既定ではオフです。 | | `stallTimeoutSec?` | `number` | `300` | `response.incomplete` より前にアップストリーム データがない秒数。最小 1。 | `connectTimeoutMs?` | `number` | `200000` |試行ごとの DNS/TCP/TLS/最終ヘッダーの期限。本体が生成される前に終了します。 | | `shutdownTimeoutMs?` | `number` | `5000` |アクティブなターンが中止される前の正常な排出期限。 | @@ -171,3 +172,5 @@ Anthropic OAuth サイドカーは、opencodex の既存のクロード コー ## Codex クォータのネットワーク診断 メイン Codex アカウント行の `quotaRefresh` はクォータ取得の診断情報であり、残量やモデルへのアクセス権を示すものではありません。キャッシュ利用時や取得を行わない場合は省略されることがあります。取得には操作中のシェルではなく、実行中のプロキシサービスの環境が使われます。`proxy` 未設定では既存の環境を維持し、`"auto"` は起動時に Windows の静的プロキシ設定だけを読みます。PAC/WPAD、SOCKS のみの設定、実行中の変更は自動反映されません。TUN での成功だけでは HTTP プロキシ経路の正常性は確認できません。[コマンドと状態の説明(英語)](/reference/configuration/server/#codex-quota-network-diagnostics)を参照してください。 + +`dropCodexSafetyBuffering`: プロバイダーの安全性の適用と拒否応答は変更しません。native `codex.response.metadata.headers` WebSocket メタデータと `/responses/compact` は対象外です。 diff --git a/docs-site/src/content/docs/ko/guides/combos.md b/docs-site/src/content/docs/ko/guides/combos.md index 80eac32c2d..e507889c6f 100644 --- a/docs-site/src/content/docs/ko/guides/combos.md +++ b/docs-site/src/content/docs/ko/guides/combos.md @@ -145,15 +145,16 @@ ocx combo set balanced \ ## 기본 reasoning effort -`defaultEffort`는 다음 조건이 모두 참일 때만 `reasoning.effort`를 채웁니다. +`defaultEffort`는 콤보 기본값이 null이 아니고, 선택한 대상의 지원 목록이 알려져 있으며 비어 있지 않을 때 생략된 `reasoning.effort`를 채웁니다. 설정값을 지원하면 그대로 사용합니다. 그렇지 않으면 설정값 이하의 가장 높은 지원 단계를 사용하고, 그런 단계가 없으면 가장 낮은 지원 단계를 사용합니다. 지원 목록이 없거나 비어 있으면 기본값을 생략합니다. -1. 콤보에 null이 아닌 기본값이 있습니다. -2. 호출자가 effort를 설정하지 않았습니다. -3. 선택된 대상의 카탈로그가 그 정확한 effort를 광고합니다. +기본값 주입은 기존 effort와 다른 reasoning 필드를 보존합니다. 아래의 capability 정규화는 별도로 지원되지 않는 effort·thinking 제어를 제거할 수 있습니다. 기본값은 `low`, `medium`, `high`, `xhigh`, `max`, `ultra`이며, 필드를 생략하거나 `null`로 설정하면 주입하지 않습니다. -요청에 `reasoning` 객체가 없으면 opencodex가 새로 만듭니다. `reasoning`은 있지만 `effort` 속성이 없으면 다른 필드는 그대로 두고 기본값만 추가합니다. 호출자가 준 effort는 절대 덮어쓰지 않습니다. -대상 기능을 알 수 없거나 설정한 effort를 포함하지 않으면 opencodex는 기본값을 생략하고 대상의 동작은 그대로 둡니다. 지원 값은 `low`, `medium`, `high`, `xhigh`, `max`, `ultra`입니다. effort를 호출자와 대상에 완전히 맡기려면 이 필드를 생략하거나 `null`로 설정하십시오. +## 서로 다른 reasoning capability + +`reasoningEffortMode`의 기본값은 `"strict"`입니다. 모든 대상의 effort 목록을 교집합으로 계산하므로 명시적 빈 목록도 반영합니다. `"adaptive"`는 빈 목록을 교집합에서 제외해 혼합 콤보에서도 선택기를 유지합니다. 알 수 없는 목록은 두 모드 모두 카탈로그 교집합을 제한하지 않습니다. + +전송 시 명시적 빈 목록은 두 모드 모두에서 effort·thinking 제어를 제거하고, 알 수 없는 목록은 adaptive에서만 제거합니다. `reasoning.summary`와 다른 비-effort 필드는 보존하며, 알려진 비어 있지 않은 대상은 기존 방식으로 effort를 결정합니다. strict의 unknown 대상과 일반 native Chat의 unknown 선언은 그대로 유지됩니다. 기본값 주입은 기존 effort를 덮어쓰지 않지만, 이 capability 정규화는 지원되지 않는 제어를 제거할 수 있습니다. ## 암호화된 v2 서브에이전트 작업 @@ -241,6 +242,7 @@ ocx combo remove --yes | `cooldownMs` | 아니요 | 미설정 → 업스트림 폴백(요청 속도 제한 429 코드 `1302`/`1305`는 5초, 그 외는 60초) | 1에서 600000 사이의 정수입니다. 설정하면 사용 가능한 업스트림 `Retry-After` 또는 Codex 재설정 신호가 없을 때 요청 속도 제한 429를 포함한 대상별 쿨다운으로 적용됩니다. 설정하지 않으면 업스트림 폴백을 사용합니다. | | `waitForCooldownMs` | 아니요 | `0` | 0에서 600000 사이의 정수입니다. `combo_unavailable`을 반환하기 전에 가장 먼저 적합해지는 쿨다운 중인 대상을 기다리는 최대 시간입니다. 중단하면 대기가 취소됩니다. | | `defaultEffort` | 아니요 | `null` | `low`, `medium`, `high`, `xhigh`, `max`, 또는 `ultra`입니다. 호출자가 effort를 생략하고 대상이 지원을 광고할 때만 적용됩니다. | +| `reasoningEffortMode` | 아니요 | `"strict"` | `strict` 또는 `adaptive`; 혼합 capability의 교집합과 대상별 제어 정규화를 선택합니다. | | `alias` | 아니요 | 없음 | 선택적으로 앞뒤 공백을 제거한 공개 모델 ID입니다. 위의 alias 규칙을 따릅니다. 빈 값은 alias 없음으로 저장됩니다. | | `nativeAlias` | 아니요 | `false` | 현재 지원되는 bare native alias가 routing/catalog 우선권을 갖도록 명시적으로 허용합니다. | | `displayName` | 아니요 | 없음 | catalog 표시 전용 label입니다. `nativeAlias`가 true이면 필수입니다. | diff --git a/docs-site/src/content/docs/ko/reference/architecture.md b/docs-site/src/content/docs/ko/reference/architecture.md index fa3056fb49..d95bf391ad 100644 --- a/docs-site/src/content/docs/ko/reference/architecture.md +++ b/docs-site/src/content/docs/ko/reference/architecture.md @@ -127,6 +127,15 @@ envelope를 각각 4 MiB로 제한하고 8 MiB producer queue 상한이 있는 b relay를 거칩니다. queue overflow 시 업스트림을 닫고 downstream에는 terminal `response.failed` 이벤트와 `[DONE]`을 내보냅니다. +최종 전송 모델이 `gpt-5.3-codex-spark`이면 canonical ChatGPT forward 경로는 HTTP 헤더와 +네이티브 WS 프레임 메타데이터 모두에서 Responses Lite를 명시적으로 끕니다. 별칭으로 Spark를 +선택해도 동일합니다. 다만 이 비활성화는 전송 본문에 비어 있지 않은 `tools` 배열을 가진 `additional_tools` 항목이 없을 때만 +적용됩니다. 이 그룹 자체가 Lite의 도구 전달 형식이므로, 그것을 사용하는 Spark 본문은 호출자나 +설정 헤더가 무엇이든 Lite를 켠 상태로 유지합니다. Lite 식별값이 바뀌면 기존 소켓은 사용을 종료하며, 이후 같은 식별값으로 +재사용 조건을 충족하는 요청은 새 소켓을 재사용할 수 있습니다. 다른 모델과 게이트웨이는 기존 +Lite 정책을 유지합니다. 네이티브 메타데이터 형식이 잘못된 경우에는 본문을 바꾸지 않고 +기존처럼 HTTP로 폴백합니다. + Codex 컨텍스트 compaction은 라우팅된 모델에서도 동작합니다. `server/responses/compact.ts`는 `POST /v1/responses/compact`를 내부 라우팅 요약 턴으로 처리해 압축된 히스토리를 반환합니다. `responses/parser.ts`와 `bridge.ts`는 remote compaction v2의 `compaction_trigger` 턴을 처리해 diff --git a/docs-site/src/content/docs/ko/reference/configuration/routing.md b/docs-site/src/content/docs/ko/reference/configuration/routing.md index cfc20d07fc..2f617695ce 100644 --- a/docs-site/src/content/docs/ko/reference/configuration/routing.md +++ b/docs-site/src/content/docs/ko/reference/configuration/routing.md @@ -68,7 +68,8 @@ Codex Auth 페이지에서 이 picker 동작을 opt-in할 수 있습니다. 비 | `targets` | `{ provider: string; model: string; weight?: number }[]` | required | 순서가 있는 concrete route입니다. `weight`는 1–10000이며 기본값은 `1`입니다. | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | 선택 전략입니다. 대상 순서는 `failover` 우선순위이고, 가중치는 `round-robin`과 `random` 추첨 비율을 결정하며, `least-used`는 기록된 성공 횟수를 따르고, `reset-window`는 가장 가까운 할당량 재설정을 따릅니다. | | `stickyLimit?` | `number` | `1` | 한 round-robin 배치에서 유지되는 성공 요청 수입니다. 범위는 1–100입니다. | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | 호출자가 effort를 생략했고 선택된 대상이 요청한 rung를 광고할 때만 적용됩니다. | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | `defaultEffort`는 콤보 기본값이 null이 아니고, 선택한 대상의 지원 목록이 알려져 있으며 비어 있지 않을 때 생략된 `reasoning.effort`를 채웁니다. 설정값을 지원하면 그대로 사용합니다. 그렇지 않으면 설정값 이하의 가장 높은 지원 단계를 사용하고, 그런 단계가 없으면 가장 낮은 지원 단계를 사용합니다. 지원 목록이 없거나 비어 있으면 기본값을 생략합니다. | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"`는 빈 목록을 포함한 알려진 대상 지원 목록의 교집합을 사용하고, `"adaptive"`는 빈 목록을 제외합니다. 알 수 없는 목록은 두 모드 모두 교집합을 제한하지 않습니다. 전송 시 명시적 빈 목록은 두 모드에서 effort·thinking 제어를 제거하고, 알 수 없는 목록은 adaptive에서만 제거합니다. `reasoning.summary`는 보존됩니다. 알려진 비어 있지 않은 대상의 effort 결정과 대상 선택·순서는 그대로입니다. | | `alias?` | `string` | — | 정규화된 picker slug 대신 쓰는 선택적 공개 model id입니다. | | `nativeAlias?` | `boolean` | `false` | 현재 지원되는 bare native id가 해당 비수식 id에만 우선하도록 합니다. 계정 또는 프로바이더로 수식된 OpenAI route는 별도로 유지됩니다. | | `displayName?` | `string` | — | catalog 표시 전용 label이며 native alias에서는 비어 있지 않아야 합니다. | diff --git a/docs-site/src/content/docs/ko/reference/configuration/server.md b/docs-site/src/content/docs/ko/reference/configuration/server.md index 577379fed9..5f60fc7c6c 100644 --- a/docs-site/src/content/docs/ko/reference/configuration/server.md +++ b/docs-site/src/content/docs/ko/reference/configuration/server.md @@ -13,6 +13,7 @@ description: 리스너, 원격 접근, admission 키, 타임아웃, 저장소, | `hostname?` | `string` | `"127.0.0.1"` | 바인드 주소입니다. 루프백이 아닌 바인드에는 데이터 admission 토큰이 필요하며, `OPENCODEX_API_AUTH_TOKEN` → `OCX_API_TOKEN_FILE` → 설치된 owner-only `service-api-token` 순서로 결정됩니다. 손으로 내보낼 값은 없습니다. [Remote access](#remote-access)를 보세요. | | `proxy?` | `string` | — | 송신용 HTTP(S) 프록시 URL 또는 `${ENV_VAR}`입니다. 해당 변수가 비어 있을 때만 `HTTP_PROXY` / `HTTPS_PROXY`에 적용되며, 루프백은 `NO_PROXY`에 그대로 남습니다. | | `emptyCompletionRetry?` | `boolean` | `false` | 텍스트나 도구 호출이 없는 Responses 턴을, 터미널 이벤트 전에 스트림이 종료된 경우를 포함해 동일한 요청으로 한 번 재시도하도록 선택합니다. 재시도에는 비용이 발생할 수 있습니다. `OCX_EMPTY_COMPLETION_RETRY=0`은 설정을 바꾸지 않고 비활성화하며, combo 및 routed-compaction turn은 제외됩니다. | +| `dropCodexSafetyBuffering?` | `boolean` | `false` | Canonical Codex Responses 응답의 선택적 safety-buffering 헤더 두 개와 SSE 힌트를 제거합니다. 공급자의 안전 정책이나 거절 응답은 바뀌지 않습니다. Native WS 메타데이터와 compact는 제외됩니다. | | `stallTimeoutSec?` | `number` | `300` | 업스트림 데이터가 없을 때 `response.incomplete`가 되기까지의 초 수입니다. 최소 1입니다. | | `connectTimeoutMs?` | `number` | `200000` | 시도별 DNS/TCP/TLS/최종 헤더 기한입니다. 본문 생성 전에 끝납니다. | | `shutdownTimeoutMs?` | `number` | `5000` | 진행 중인 turn을 중단하기 전에 허용하는 정상 종료 드레인 기한입니다. | diff --git a/docs-site/src/content/docs/reference/architecture.md b/docs-site/src/content/docs/reference/architecture.md index e0fbc8bcba..a9bdf48978 100644 --- a/docs-site/src/content/docs/reference/architecture.md +++ b/docs-site/src/content/docs/reference/architecture.md @@ -157,6 +157,15 @@ upstream WS responses keep the downstream SSE contract and bypass `tee()` throug single-reader relay (4 MiB per raw/enveloped frame and an 8 MiB producer queue). Queue overflow closes the upstream and emits a terminal downstream `response.failed` event followed by `[DONE]`. +For the final outgoing model `gpt-5.3-codex-spark`, canonical ChatGPT forwarding explicitly +disables Responses Lite in both the HTTP header and native WS frame metadata, including when +an alias selects Spark — but only when the outgoing body carries no `additional_tools` item with a nonempty `tools` array. +That group IS the Lite tool-delivery shape, so a Spark body that still uses it keeps Lite ON even +if a caller or configured header said otherwise; otherwise the frame would advertise non-Lite +while the tools exist only in the Lite shape. A changed Lite identity retires the old socket; subsequent eligible +requests with the same identity can reuse the new socket. Other models and gateways keep +their existing Lite policy. Malformed native metadata still falls back to HTTP with its body unchanged. + When a provider rejects a streaming request with HTTP 413 before SSE begins, OpenCodex emits one terminal `response.failed` event with `context_length_exceeded` instead of relaying the retryable unknown status. This lets Codex stop its reconnect loop and apply its own context-compaction policy diff --git a/docs-site/src/content/docs/reference/configuration/agents.md b/docs-site/src/content/docs/reference/configuration/agents.md index a113813b40..819aa3f360 100644 --- a/docs-site/src/content/docs/reference/configuration/agents.md +++ b/docs-site/src/content/docs/reference/configuration/agents.md @@ -37,6 +37,7 @@ still depends on upstream support for your account. | `subagentModelFallbackPollMs?` | `number` | `60000` | Availability-probe cache interval. Values below 1000 ms fall back to the default. | | `effortCap?` | `string` | — | Hard ceiling for qualifying v2 main turns and marked spawned-child turns. Accepts `low` through `ultra`. | | `subagentEffortCap?` | `string` | — | Additional ceiling for spawned-child turns only. When both caps apply, the lower wins. | +| `plaintextV2AgentMessages?` | `boolean` | — (unset) | Experimental opt-in. It runs only when explicitly set to `true` and asks eligible native ChatGPT v2 parents to emit `spawn_agent`, `send_message`, and `followup_task` message arguments as plaintext. See [Plaintext v2 agent messages](#plaintext-v2-agent-messages). | | `agentTaskRecovery?` | `object` | — | Experimental opt-in recovery for backend-encrypted v2 tasks sent to routed providers. Disabled unless `enabled: true`; see [Encrypted v2 task recovery](#encrypted-v2-task-recovery). | Manage the surface with the dashboard or @@ -162,6 +163,57 @@ on a mid-thread model switch. } ``` +## Plaintext v2 agent messages + +`plaintextV2AgentMessages` is unset in a fresh config and runs only when explicitly set to `true`. +The caller must use the Responses wire, and the final destination must use the canonical ChatGPT +Codex forward transport: `adapter: "openai-responses"`, `authMode: "forward"`, and the exact base URL +`https://chatgpt.com/backend-api/codex`. OpenAI API-key providers, custom OpenAI-compatible +gateways, routes whose final destination is another provider, and non-Responses callers are never +rewritten. + +For an eligible v2 request, opencodex recognizes the catalog by a top-level `collaboration` +namespace with a direct `spawn_agent` child. It removes +`parameters.properties.message.encrypted: true`, when present, only from `spawn_agent`, +`send_message`, and `followup_task`. ChatGPT reserves both the `collaboration` namespace and those +three tool names, so the request uses fixed private aliases for all four identities. Before making +that change, opencodex checks top-level and `additional_tools` catalogs, nested namespaces, +`tool_search_output` declarations, `tool_choice`, and prior call items for the private namespace and +fixed aliases. Any conflict leaves the entire request unchanged. OpenCodex restores only the +request-scoped aliases in JSON, SSE, and WebSocket responses before Codex receives the tool call. +The `encrypted_function_args: []` field is preserved so compatible Codex clients recognize the +message as plaintext. + +This path adds no recovery request and therefore does not spend the extra ChatGPT quota used by a +cache miss in `agentTaskRecovery`. It cannot change tasks that are already encrypted. If the request +already declares the private alias or a conflicting reference, opencodex leaves that request +unchanged; separately enabled recovery can still handle a routed task that is later encrypted. If +ChatGPT rejects or ignores the modified schema, or the Codex client does not recognize the plaintext +response fields, the call can fail. OpenCodex does not retry the parent request with the original +schema because doing so could duplicate quota use or tool calls. + +Restoration uses a 10,000-identity traversal budget for each response payload. If a payload exhausts +that budget, bounded JSON returns HTTP 502 and a stream returns `response.failed`; neither path +sends the private aliases to Codex or saves the refused response for `previous_response_id` +continuation. + +For successfully rewritten calls, the option removes application-layer encryption from agent +message arguments. HTTPS still encrypts network transport, but message text can appear in Codex +task history, routed-provider requests, `responses-state.json` or its spill files, and +`usage-debug.jsonl` when debug capture is enabled. The behavior depends on undocumented ChatGPT +schema and response fields and may stop working after a backend or client update. Startup prints a +warning while it is enabled. + +```json +{ + "plaintextV2AgentMessages": true +} +``` + +The equivalent CLI command is `ocx config set plaintextV2AgentMessages true`. Restart the proxy +after changing the setting. + + ## Encrypted v2 task recovery `agentTaskRecovery` is an experimental compatibility path for backend-encrypted v2 tasks that reach diff --git a/docs-site/src/content/docs/reference/configuration/providers.md b/docs-site/src/content/docs/reference/configuration/providers.md index 67e072177c..eb61194c94 100644 --- a/docs-site/src/content/docs/reference/configuration/providers.md +++ b/docs-site/src/content/docs/reference/configuration/providers.md @@ -208,6 +208,7 @@ Providers can expose a built-in shorthand, such as `agy` for `google-antigravity | `preserveReasoningContentModels?` | `string[]` | Models requiring prior assistant `reasoning_content` in chat history. | | `reasoningDetailsModels?` | `string[]` | Models whose endpoint returns thinking as a structured `reasoning_details` array (MiniMax M-series with `reasoning_split`); stream deltas are cumulative snapshots that are prefix-diffed, and preserved reasoning replays as a `reasoning_details` array instead of a `reasoning_content` string. | | `requiresReasoningPlaceholderModels?` | `string[]` | Models whose upstream rejects a tool_call continuation missing `reasoning_content` (DeepSeek thinking mode); a minimal placeholder is injected when the replay cache misses. Defaults to `preserveReasoningContentModels`; set `[]` to opt out. | +| `showThinkingSummary?` | `boolean` | Display provider-authored summaries when a Responses client omits `reasoning.summary`. Explicit wire `"none"` wins; a client that serializes its preference as omission cannot be distinguished. Raw reasoning remains content and is never relabeled as a summary. The `google-antigravity` preset defaults to `true`; explicit `false` disables that default. CCA Gemini requests also opt into `generationConfig.thinkingConfig.includeThoughts` when display is enabled; image, Claude and gpt-oss requests do not. This does not change client configuration or global catalog summary defaults. | | `thinkingToggleModels?` | `string[]` | Chat models using `thinking.enabled` rather than an effort ladder. | | `thinkingBudgetModels?` | `string[]` | Chat models using integer `thinking_budget`; effort maps to a budget fraction. | | `noVisionModels?` | `string[]` | Text-only models sent through the vision sidecar; matching tolerates an Ollama `:size` tag. | diff --git a/docs-site/src/content/docs/reference/configuration/routing.md b/docs-site/src/content/docs/reference/configuration/routing.md index 25bd32f64e..2934c337a6 100644 --- a/docs-site/src/content/docs/reference/configuration/routing.md +++ b/docs-site/src/content/docs/reference/configuration/routing.md @@ -90,8 +90,8 @@ namespace, and cannot use reserved bare native families such as `gpt-*`, `o1-*`, | `stickyLimit?` | `number` | `1` | Successful requests retained in one round-robin batch. Range 1–100. Applies only to round-robin. | | `cooldownMs?` | `number` | unset → upstream fallback (5 s for request-rate 429 codes `1302`/`1305`, otherwise 60 s) | Range 1–600000. When set, applies whenever no usable upstream `Retry-After` or Codex reset signal exists, including request-rate 429s; when unset, uses the upstream fallback. Upstream signals take precedence and all cooldowns are capped at 10 minutes. | | `waitForCooldownMs?` | `number` | `0` | Maximum wait for the earliest eligible cooling target on each selection attempt before returning `combo_unavailable`. Range 0–600000; an abort cancels the wait. | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | Applied only when the caller omits effort and the selected target advertises the requested rung. | -| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` intersects every known target effort ladder, so a target advertising no effort control empties the combo's picker. `"adaptive"` excludes those empty ladders from the published intersection. Picker metadata only; target selection and dispatch are unchanged. | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | `defaultEffort` fills an absent `reasoning.effort` when the combo has a non-null default and the selected target has a known, nonempty supported ladder. If the target supports the configured value, it is retained; otherwise the highest supported rung at or below it is used, or the lowest supported rung when none is lower. Unknown or empty ladders omit the default. | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` intersects all known target ladders, including empty ones; `"adaptive"` excludes empty ladders. Unknown ladders are catalog wildcards in both modes. At dispatch, explicit empty ladders remove effort/thinking controls in both modes; unknown ladders do so only in adaptive. `reasoning.summary` is preserved. Known nonempty targets retain their effort resolution, and target selection/order is unchanged. | | `imageInput?` | `"auto" \| "disabled"` | `"auto"` | `"auto"` publishes image only when every target supports images; `"disabled"` forces text-only (drops image from published modalities and rejects image-bearing requests before dispatch). | | `alias?` | `string` | — | Optional public model id in place of the canonical picker slug. | | `nativeAlias?` | `boolean` | `false` | Let a currently supported bare native id take precedence only for that unqualified id. Bare `gpt-5.6-*` ids use Codex Pool/Direct credentials. Account-qualified routes remain distinct. Provider-qualified routes such as `openai-apikey/gpt-5.6-*` use their configured API-key route and never fall through to the native alias. | diff --git a/docs-site/src/content/docs/reference/configuration/server.md b/docs-site/src/content/docs/reference/configuration/server.md index 108d2e0925..3781d5846e 100644 --- a/docs-site/src/content/docs/reference/configuration/server.md +++ b/docs-site/src/content/docs/reference/configuration/server.md @@ -15,6 +15,7 @@ runs helper features around provider requests. | `proxy?` | `string` | — | Outbound HTTP(S) proxy URL, `${ENV_VAR}`, or `"auto"`. Applied to `HTTP_PROXY` / `HTTPS_PROXY` only when those variables are unset; loopback remains in `NO_PROXY`. `"auto"` reads the Windows system proxy (WinINET `ProxyEnable`/`ProxyServer`, `https=` then `http=` entry) once at process start and logs the host it chose. On other platforms, or when the system proxy is off, SOCKS-only, or unreadable, it uses direct egress and says so. PAC/WPAD and live proxy changes are not followed; restart the service after changing the system proxy. | | `noProxy?` | `string \| string[]` | — | Hosts that bypass `proxy`, merged with inherited `NO_PROXY` and loopback entries. A string may use comma-separated `NO_PROXY` syntax or `${ENV_VAR}`. | | `emptyCompletionRetry?` | `boolean` | `false` | Opt in to one identical Responses retry when a turn has no text or tool call, including a stream that ends before a terminal event. The retry may be billable. `OCX_EMPTY_COMPLETION_RETRY=0` disables it without changing config; combo and routed-compaction turns remain excluded. | +| `dropCodexSafetyBuffering?` | `boolean` | `false` | Remove optional client-facing hints from canonical Codex Responses passthrough: the two `x-codex-safety-buffering-enabled` / `x-codex-safety-buffering-faster-model` response headers, `response.metadata` events whose metadata type is `safety_buffering`, and top-level `safety_buffering` fields. Other headers, response data, policy refusals and failures are preserved. This does not disable provider safety enforcement or upstream buffering. Native `codex.response.metadata.headers` WebSocket metadata and `/responses/compact` are outside this filter. | | `stallTimeoutSec?` | `number` | `300` | Seconds without upstream data before `response.incomplete`. Minimum 1. | | `oauthOpenBrowser?` | `boolean` | `true` | Whether a login may open a browser on the machine running the proxy. Absent and `true` both open, so an existing install is unchanged; only an explicit `false` declines. Decline when you need the authorization link in a different browser profile, or when the dashboard is not on the proxy's machine — the login still starts and the URL is still returned and displayed. `POST /api/oauth/login` and `POST /api/codex-auth/login` accept a per-request `openBrowser` boolean that overrides this, and the dashboard exposes the same choice beside the login button. Device-code flows never open a browser either way. | | `connectTimeoutMs?` | `number` | `200000` | Per-attempt DNS/TCP/TLS/final-header deadline; it ends before body generation. | diff --git a/docs-site/src/content/docs/reference/proxy-formats.md b/docs-site/src/content/docs/reference/proxy-formats.md index c5fa271bf1..2273962329 100644 --- a/docs-site/src/content/docs/reference/proxy-formats.md +++ b/docs-site/src/content/docs/reference/proxy-formats.md @@ -24,6 +24,14 @@ should select among several targets. Credential-bearing model, image, video, and search requests do not automatically follow HTTP redirects, including same-origin redirects. Configure the final upstream API URL instead of a redirecting alias. A redirect does not cause the server to resend credentials or the request body to its destination. The response owner retains its existing error or relay behavior; native Responses and compact routes can return the original 3xx and `Location` to the client. Client redirect behavior is separate from this server transport policy. +## Console upload rejections + +An exact Console or Console Go `Invalid upload request.` HTTP 400 from a canonical +OpenCode Zen/Go generation endpoint receives one retry after 800 ms. The proxy reuses +the same serialized request and records the recovery in Logs. Other 400 errors, +custom destinations, cancellations and repeated upload rejections remain failures. +This does not retry filtered model responses or interrupted streams. + ## Endpoint overview | Client surface | Endpoint | Successful non-stream result | Successful stream or socket result | diff --git a/docs-site/src/content/docs/ru/guides/combos.md b/docs-site/src/content/docs/ru/guides/combos.md index b873f4ed03..ddd16f6175 100644 --- a/docs-site/src/content/docs/ru/guides/combos.md +++ b/docs-site/src/content/docs/ru/guides/combos.md @@ -177,20 +177,16 @@ Failover намеренно ограничен. Он помогает при п ## Effort по умолчанию -`defaultEffort` подставляет `reasoning.effort` только если одновременно выполняются все условия: +`defaultEffort` заполняет отсутствующий `reasoning.effort`, если задан `defaultEffort`, отличный от `null`, и список поддерживаемых уровней цели известен и непуст. Поддерживаемое настроенное значение сохраняется; иначе выбирается максимальный поддерживаемый уровень не выше него, а если такого нет — минимальный поддерживаемый уровень. При неизвестном или пустом списке default не добавляется. -1. у combo задано ненулевое значение по умолчанию; -2. вызывающая сторона сама не указала effort; и -3. каталог выбранной цели объявляет поддержку именно этого effort. +Подстановка сохраняет существующий effort и остальные поля reasoning. Нормализация возможностей ниже может отдельно удалить неподдерживаемые параметры effort/thinking. Значения default: `low`, `medium`, `high`, `xhigh`, `max`, `ultra`; отсутствие поля или `null` отключает подстановку. -Если в запросе нет объекта `reasoning`, opencodex создаёт его. Если `reasoning` есть, но в нём нет -свойства `effort`, остальные поля сохраняются, а значение по умолчанию добавляется. Effort, -заданный вызывающей стороной, никогда не перезаписывается. -Если возможности цели неизвестны или не включают настроенный effort, opencodex опускает значение -по умолчанию и оставляет нативное поведение цели без изменений. Поддерживаются `low`, `medium`, -`high`, `xhigh`, `max` и `ultra`; опустите поле или задайте `null`, чтобы полностью оставить выбор -effort вызывающей стороне и цели. +## Разные возможности reasoning в одном combo + +`reasoningEffortMode` по умолчанию равен `"strict"`: публикуется пересечение списков effort всех целей, включая явно пустые списки. `"adaptive"` исключает пустые списки из пересечения, сохраняя выбор effort для смешанного combo. Неизвестные списки не ограничивают пересечение каталога в обоих режимах. + +При отправке явно пустой список удаляет параметры effort и thinking в обоих режимах; неизвестный список — только в adaptive. `reasoning.summary` и остальные поля, не задающие effort, сохраняются. Для известных непустых списков разрешение effort не меняется. Неизвестные цели в strict и неизвестные объявления обычного native Chat сохраняют параметры вызывающей стороны. Подстановка значения по умолчанию не заменяет существующий effort, но нормализация возможностей может удалить неподдерживаемые параметры. ## Шифрованные задачи подагентов v2 @@ -292,6 +288,7 @@ Combo хранятся в объекте верхнего уровня `combos`, | `cooldownMs` | No | не задано → fallback upstream (5 с для rate-limit 429 с кодами `1302`/`1305`, иначе 60 с) | Целое число от 1 до 600000. Если задано, применяется как cooldown каждой цели, когда нет пригодного upstream `Retry-After` или сигнала сброса Codex, включая rate-limit 429; если не задано, используется fallback upstream. | | `waitForCooldownMs` | No | `0` | Целое число от 0 до 600000. Максимальное время ожидания самой ранней подходящей цели в cooldown перед возвратом `combo_unavailable`; отмена запроса отменяет ожидание. | | `defaultEffort` | No | `null` | `low`, `medium`, `high`, `xhigh`, `max` или `ultra`; применяется только когда вызывающая сторона не указала effort, а цель объявляет поддержку. | +| `reasoningEffortMode` | Нет | `"strict"` | `strict` или `adaptive`; задаёт пересечение возможностей и нормализацию параметров конкретной цели. | | `alias` | No | none | Необязательный обрезанный публичный id модели; используйте правила alias выше. Пустое значение хранится как отсутствие alias. | | `nativeAlias` | No | `false` | Явно разрешает поддерживаемому сейчас bare native alias перехватить приоритет routing/catalog только для неквалифицированного id. Bare `gpt-5.6-*` использует учётные данные Codex Pool/Direct; маршруты с квалификатором аккаунта сохраняют свою идентичность, а provider-qualified `openai-apikey/gpt-5.6-*` использует API-ключ и никогда не переходит на native alias. | | `displayName` | No | none | Метка только для отображения в catalog; обязательна при `nativeAlias: true`. | diff --git a/docs-site/src/content/docs/ru/reference/architecture.md b/docs-site/src/content/docs/ru/reference/architecture.md index c46da773c0..cd3f776efb 100644 --- a/docs-site/src/content/docs/ru/reference/architecture.md +++ b/docs-site/src/content/docs/ru/reference/architecture.md @@ -151,6 +151,15 @@ loopback; настроенные записи `corsAllowOrigins` расширя `426 upgrade_required`; Codex тогда откатывается на HTTP для этой сессии. Когда установлено `"websockets": true`, та же конечная точка принимает апгрейд и использует WebSocket-мост. +Для итоговой исходящей модели `gpt-5.3-codex-spark` каноническая пересылка в ChatGPT явно +отключает Responses Lite в HTTP-заголовке и нативных метаданных WS-кадра, в том числе при +выборе Spark через псевдоним — но только если в исходящем теле нет элемента `additional_tools` с непустым массивом `tools`. +Эта группа и ЕСТЬ Lite-форма доставки инструментов, поэтому тело Spark, которое её использует, +сохраняет Lite ВКЛЮЧЁННЫМ независимо от заголовка вызывающего клиента или конфигурации. Изменение идентичности Lite выводит старый сокет из использования; +последующие подходящие запросы с той же идентичностью могут повторно использовать новый сокет. +Другие модели и шлюзы сохраняют прежнюю политику Lite. Некорректные нативные метаданные +по-прежнему приводят к откату на HTTP без изменения тела запроса. + Compaction контекста Codex работает для маршрутизируемых моделей. `server/responses/compact.ts` обрабатывает `POST /v1/responses/compact`, выполняя внутренний маршрутизируемый ход суммаризации и возвращая сжатую историю, а `responses/parser.ts` и `bridge.ts` обрабатывают ходы diff --git a/docs-site/src/content/docs/ru/reference/configuration/routing.md b/docs-site/src/content/docs/ru/reference/configuration/routing.md index dbc7181044..595916aebd 100644 --- a/docs-site/src/content/docs/ru/reference/configuration/routing.md +++ b/docs-site/src/content/docs/ru/reference/configuration/routing.md @@ -87,7 +87,8 @@ selector-qualified строки и возвращает обычные GPT-ст | `targets` | `{ provider: string; model: string; weight?: number }[]` | required | Упорядоченные конкретные маршруты. `weight` находится в диапазоне 1–10000 и по умолчанию равен `1`. | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | Стратегия выбора. Порядок целей задаёт приоритет `failover`; значения `weight` определяют взвешивание выборов `round-robin` и `random`; `least-used` следует числу зарегистрированных успешных запросов; `reset-window` следует ближайшему сбросу квоты. | | `stickyLimit?` | `number` | `1` | Число успешных запросов, удерживаемых в одной партии round-robin. Диапазон 1–100. | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | Применяется, только если вызывающая сторона не задала effort, а выбранная цель объявляет эту ступень. | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | `defaultEffort` заполняет отсутствующий `reasoning.effort`, если задан `defaultEffort`, отличный от `null`, и список поддерживаемых уровней цели известен и непуст. Поддерживаемое настроенное значение сохраняется; иначе выбирается максимальный поддерживаемый уровень не выше него, а если такого нет — минимальный поддерживаемый уровень. При неизвестном или пустом списке default не добавляется. | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` вычисляет пересечение известных списков уровней, включая пустые; `"adaptive"` исключает пустые списки. Неизвестные списки не ограничивают пересечение в обоих режимах. При отправке явно пустой список удаляет параметры effort/thinking в обоих режимах, а неизвестный — только в adaptive. `reasoning.summary` сохраняется. Разрешение effort для известных непустых списков, выбор и порядок целей не меняются. | | `alias?` | `string` | — | Необязательный публичный id модели вместо канонического slug в селекторе. | | `nativeAlias?` | `boolean` | `false` | Даёт поддерживаемому bare native id приоритет только для этого неквалифицированного id. Bare `gpt-5.6-*` использует учётные данные Codex Pool/Direct. Маршруты с квалификатором аккаунта остаются отдельными. Провайдер-квалифицированные маршруты, например `openai-apikey/gpt-5.6-*`, используют настроенный API-ключ и никогда не переходят на native alias. | | `displayName?` | `string` | — | Метка только для catalog; для native alias обязательна и не может быть пустой. | diff --git a/docs-site/src/content/docs/ru/reference/configuration/server.md b/docs-site/src/content/docs/ru/reference/configuration/server.md index 84b3caf183..000a194b5c 100644 --- a/docs-site/src/content/docs/ru/reference/configuration/server.md +++ b/docs-site/src/content/docs/ru/reference/configuration/server.md @@ -14,6 +14,7 @@ description: Listener, удалённый доступ, admission key, тайм | `hostname?` | `string` | `"127.0.0.1"` | Адрес bind'а. Не-loopback bind требует `OPENCODEX_API_AUTH_TOKEN`. | | `proxy?` | `string` | — | URL исходящего HTTP(S)-прокси или `${ENV_VAR}`. Применяется к `HTTP_PROXY` / `HTTPS_PROXY` только когда эти переменные не заданы; loopback всегда остаётся в `NO_PROXY`. | | `emptyCompletionRetry?` | `boolean` | `false` | Явно включает один идентичный повтор Responses, если в turn нет ни текста, ни tool call, включая случай, когда stream завершается до terminal event. Повтор может тарифицироваться. `OCX_EMPTY_COMPLETION_RETRY=0` отключает его без изменения config; combo и routed-compaction turn исключены. | +| `dropCodexSafetyBuffering?` | `boolean` | `false` | Удаляет подсказки Codex safety-buffering из passthrough-ответов Codex Responses: заголовки `x-codex-safety-buffering-enabled` / `x-codex-safety-buffering-faster-model`, SSE-события `response.metadata` типа `safety_buffering` и поле `safety_buffering` в других SSE-событиях. Codex TUI отображает их как предложение повторить запрос с более быстрой моделью, действие по умолчанию в котором переключает сессию на более слабую модель. Остальные заголовки `x-codex-*` и содержимое других SSE-событий передаются без изменений, кроме удаления этого поля. По умолчанию выключено. | | `stallTimeoutSec?` | `number` | `300` | Секунды без upstream-данных до `response.incomplete`. Минимум 1. | | `connectTimeoutMs?` | `number` | `200000` | Дедлайн одной попытки DNS/TCP/TLS/final-header; он завершается до генерации тела ответа. | | `shutdownTimeoutMs?` | `number` | `5000` | Дедлайн graceful-drain до принудительного прерывания активных turn'ов. | @@ -219,3 +220,5 @@ opencodex. Перед использованием прогоните soak-test ## Сетевая диагностика квоты Codex Поле `quotaRefresh` в строке основного аккаунта Codex описывает получение квоты, а не её остаток или право доступа к модели. Оно может отсутствовать при чтении кэша или если запрос не выполнялся. Используется окружение работающего прокси-сервиса, а не текущего терминала. Если `proxy` не задан, существующее окружение сохраняется; `"auto"` читает только статические настройки прокси Windows при запуске. PAC/WPAD, настройки только SOCKS и изменения во время работы автоматически не учитываются. Успех через TUN сам по себе не подтверждает исправность пути HTTP-прокси. См. [команды и состояния на английском](/reference/configuration/server/#codex-quota-network-diagnostics). + +`dropCodexSafetyBuffering`: не меняет проверки безопасности провайдера или отказы. Native WebSocket `codex.response.metadata.headers` и `/responses/compact` не входят в область фильтра. diff --git a/docs-site/src/content/docs/tr/guides/combos.md b/docs-site/src/content/docs/tr/guides/combos.md index b8cd5bad0d..4ac9e0febf 100644 --- a/docs-site/src/content/docs/tr/guides/combos.md +++ b/docs-site/src/content/docs/tr/guides/combos.md @@ -249,22 +249,16 @@ veya politika retlerini gizlemez. ## Varsayılan akıl yürütme çabası -`defaultEffort`, yalnızca bunların tümü doğru olduğunda `reasoning.effort` -sağlar: +`defaultEffort`, combo varsayılanı null değilse ve hedefin desteklenen seviye listesi bilinen ve boş olmayan bir listeyse eksik `reasoning.effort` değerini doldurur. Yapılandırılmış değer destekleniyorsa korunur; değilse bu değeri aşmayan en yüksek desteklenen seviye, böyle bir seviye yoksa en düşük desteklenen seviye kullanılır. Liste bilinmiyor veya boşsa varsayılan eklenmez. -1. kombonun boş olmayan (non-null) bir varsayılanı vardır; -2. arayan bir çaba ayarlamamıştır; ve -3. seçilen hedefin kataloğu tam olarak bu çabayı bildirmektedir. +Varsayılan ekleme mevcut effort ve diğer reasoning alanlarını korur. Aşağıdaki yetenek normalizasyonu desteklenmeyen effort/thinking denetimlerini ayrıca kaldırabilir. Desteklenen varsayılanlar: `low`, `medium`, `high`, `xhigh`, `max`, `ultra`; alanı atlamak veya `null` kullanmak eklemeyi kapatır. -İstekte bir `reasoning` nesnesi yoksa opencodex bir tane oluşturur. Bir `effort` -özelliği olmadan `reasoning` varsa diğer alanları korur ve varsayılanı ekler. -Arayan tarafından sağlanan bir çabanın üzerine asla yazılmaz. -Hedef yeteneği bilinmediğinde veya yapılandırılan çabayı içermediğinde opencodex -varsayılanı atlar ve hedefin kendi davranışını değiştirmeden bırakır. -Desteklenen değerler `low`, `medium`, `high`, `xhigh`, `max` ve `ultra`'dır; -çabayı tamamen arayana ve hedefe bırakmak için alanı atlayın veya `null` olarak -ayarlayın. +## Farklı reasoning yetenekleri + +`reasoningEffortMode` varsayılan olarak `"strict"` kullanır: açıkça boş listeler dahil tüm hedeflerin effort listelerinin kesişimi yayımlanır. `"adaptive"`, karma kombolarda seçiciyi korumak için boş listeleri kesişimden çıkarır. Bilinmeyen listeler her iki modda da katalog kesişimini sınırlamaz. + +Gönderim sırasında açıkça boş liste her iki modda effort ve thinking denetimlerini kaldırır; bilinmeyen liste bunları yalnızca adaptive modunda kaldırır. `reasoning.summary` ve effort dışındaki alanlar korunur. Bilinen, boş olmayan hedeflerin effort çözümü değişmez. strict modundaki bilinmeyen hedefler ve normal native Chat bilinmeyen bildirimleri çağıranın denetimlerini korur. Varsayılan değer ekleme mevcut effort değerini değiştirmez; yetenek normalizasyonu desteklenmeyen denetimleri kaldırabilir. ## Şifrelenmiş v2 alt ajan görevleri @@ -373,6 +367,7 @@ saklanır: | `strategy` | Hayır | `"failover"` | İzin verilen değerler: `"failover"`, `"round-robin"`, `"random"`, `"least-used"`, `"reset-window"`. | | `stickyLimit` | Hayır | `1` | Yalnızca `round-robin` için geçerlidir; seçim başına 1 ile 100 arasında başarılı istek tam sayısı. | | `defaultEffort` | Hayır | `null` | `low`, `medium`, `high`, `xhigh`, `max` veya `ultra`; yalnızca arayan çabayı atladığında ve hedef desteği bildirdiğinde uygulanır. | +| `reasoningEffortMode` | Hayır | `"strict"` | `strict` veya `adaptive`; karma yetenek kesişimini ve hedefe özel normalizasyonu seçer. | | `alias` | Hayır | yok | İsteğe bağlı kırpılmış genel model kimliği; yukarıdaki takma ad kurallarını kullanın. Boş bir değer takma ad yok olarak saklanır. | | `nativeAlias` | Hayır | `false` | Şu anda desteklenen yalın bir yerel `alias`'ın yönlendirme ve katalog önceliği almasına açıkça izin verin. Asla takma addan çıkarılmaz. | | `displayName` | Hayır | yok | Sınırlı salt görüntüleme katalog etiketi. `nativeAlias` true olduğunda gerekli ve boş değildir. | diff --git a/docs-site/src/content/docs/tr/reference/architecture.md b/docs-site/src/content/docs/tr/reference/architecture.md index 24d2bbb7aa..5dd9b99957 100644 --- a/docs-site/src/content/docs/tr/reference/architecture.md +++ b/docs-site/src/content/docs/tr/reference/architecture.md @@ -170,6 +170,15 @@ opencodex `426 upgrade_required` döndürür; Codex daha sonra bu oturum için HTTP'ye geri döner. `"websockets": true` ayarlandığında aynı uç nokta yükseltmeyi kabul eder ve WebSocket köprüsünü kullanır. +Son gönderilen model `gpt-5.3-codex-spark` olduğunda, kanonik ChatGPT iletimi HTTP başlığında +ve yerel WS çerçevesi meta verilerinde Responses Lite'ı açıkça kapatır; Spark bir takma adla +seçildiğinde de bu geçerlidir — ancak yalnızca giden gövde boş olmayan `tools` dizisine sahip bir `additional_tools` grubu +taşımıyorsa. Bu grup Lite'ın araç teslim biçiminin kendisidir; onu kullanan bir Spark gövdesi, +çağıran veya yapılandırılmış başlık ne derse desin Lite'ı AÇIK tutar. Lite kimliği değişince eski soket kullanım dışı bırakılır; +aynı kimliğe sahip sonraki uygun istekler yeni soketi yeniden kullanabilir. Diğer modeller ve +ağ geçitleri mevcut Lite politikalarını korur. Bozuk yerel meta verilerde, istek gövdesi +değiştirilmeden HTTP'ye geri dönülmeye devam edilir. + Codex bağlam sıkıştırması yönlendirilen modeller için çalışır. `server/responses/compact.ts`, dahili bir yönlendirilen özetleme turu çalıştırarak ve sıkıştırılmış geçmişi döndürerek `POST /v1/responses/compact`'ı @@ -220,4 +229,3 @@ Dahili model `types.ts` içinde yer alır: `OcxParsedRequest`, `OcxContext`, `OcxProviderConfig`). İki yardımcı yaygın olarak kullanılır: `namespacedToolName()` ve `modelInList()` (`noVisionModels` / `noReasoningModels` için toleranslı `:size` etiketi eşleştirmesi). - diff --git a/docs-site/src/content/docs/tr/reference/configuration/routing.md b/docs-site/src/content/docs/tr/reference/configuration/routing.md index 74d239ee14..b1cbe4b484 100644 --- a/docs-site/src/content/docs/tr/reference/configuration/routing.md +++ b/docs-site/src/content/docs/tr/reference/configuration/routing.md @@ -110,7 +110,8 @@ aileleri kullanamaz. | `targets` | `{ provider: string; model: string; weight?: number }[]` | gerekli | Sıralı somut rotalar. `weight` 1–10000 arasındadır ve varsayılan olarak `1`'dir. | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | Seçim stratejisi. Hedef sırası `failover` önceliğini belirler; `weight` değerleri `round-robin` ve `random` seçimlerini biçimlendirir; `least-used` kaydedilen başarılı istekleri izler; `reset-window` en yakın kota sıfırlamasını izler. | | `stickyLimit?` | `number` | `1` | Tek bir round-robin grubunda tutulan başarılı istekler. Aralık 1–100. | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | ayarlanmamış | Yalnızca arayan çabayı atladığında ve seçilen hedef istenen basamağı bildirdiğinde uygulanır. | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | ayarlanmamış | `defaultEffort`, combo varsayılanı null değilse ve hedefin desteklenen seviye listesi bilinen ve boş olmayan bir listeyse eksik `reasoning.effort` değerini doldurur. Yapılandırılmış değer destekleniyorsa korunur; değilse bu değeri aşmayan en yüksek desteklenen seviye, böyle bir seviye yoksa en düşük desteklenen seviye kullanılır. Liste bilinmiyor veya boşsa varsayılan eklenmez. | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"`, boş listeler dahil bilinen hedef seviye listelerinin kesişimini alır; `"adaptive"` boş listeleri çıkarır. Bilinmeyen listeler iki modda da kesişimi sınırlamaz. Gönderimde açıkça boş listeler iki modda effort/thinking denetimlerini kaldırır; bilinmeyen listeler bunu yalnızca adaptive modunda yapar. `reasoning.summary` korunur. Bilinen boş olmayan hedeflerin effort çözümü ve hedef seçimi/sırası değişmez. | | `alias?` | `string` | — | Kurallı seçici slug'ı yerine isteğe bağlı genel model kimliği. | | `nativeAlias?` | `boolean` | `false` | Şu anda desteklenen bir yalın yerel kimliğin yalnızca o niteliksiz kimlik için öncelikli olmasına izin verin. Yalın `gpt-5.6-*` kimlikleri Codex Havuz/Direct kimlik bilgilerini kullanır. Hesap nitelikli rotalar ayrı kalır. `openai-apikey/gpt-5.6-*` gibi sağlayıcı nitelikli rotalar yapılandırılmış API anahtarı rotalarını kullanır ve asla yerel takma ada düşmez. | | `displayName?` | `string` | — | Yalnızca görüntüleme amaçlı katalog etiketi, yerel bir takma ad için gerekli ve boş olmamalıdır. | diff --git a/docs-site/src/content/docs/zh-cn/guides/combos.md b/docs-site/src/content/docs/zh-cn/guides/combos.md index abe32ae786..7c8c8efd63 100644 --- a/docs-site/src/content/docs/zh-cn/guides/combos.md +++ b/docs-site/src/content/docs/zh-cn/guides/combos.md @@ -165,15 +165,16 @@ combo 失败分为 **跳转** 失败和 **终止** 失败。 ## 默认推理力度 -只有在以下所有条件都满足时,`defaultEffort` 才会提供 `reasoning.effort`: +当 combo 配置了非 null 默认值且目标支持列表已知且非空时,`defaultEffort` 会填充省略的 `reasoning.effort`。目标支持配置值时保留该值,否则选择不高于配置值的最高支持档位;若不存在更低档位,则使用最低支持档位。未知或空列表不会注入默认值。 -1. combo 有一个非空默认值; -2. 调用方没有设置 effort;并且 -3. 选中的目标目录明确声明了该精确的 effort。 +默认值注入保留已有 effort 和其他 reasoning 字段。下述能力归一化可单独移除不支持的 effort/thinking 控制。默认值支持 `low`、`medium`、`high`、`xhigh`、`max`、`ultra`;省略字段或设为 `null` 可关闭注入。 -如果请求没有 `reasoning` 对象,opencodex 会创建一个。如果 `reasoning` 存在但没有 `effort` 属性,它会保留其他字段并添加默认值。调用方提供的 effort 永远不会被覆盖。 -当目标能力未知,或者不包含配置的 effort 时,opencodex 会省略默认值,并保持目标自身行为不变。支持的值是 `low`、`medium`、`high`、`xhigh`、`max` 和 `ultra`;省略该字段或将其设为 `null`,就会把 effort 完全交给调用方和目标。 +## 混合 reasoning 能力 + +`reasoningEffortMode` 默认为 `"strict"`,发布所有目标 effort 列表的交集,包括显式空列表。`"adaptive"` 在计算交集时排除空列表,让混合 combo 保留选择器。未知列表在两种模式下都不限制目录交集。 + +发送时,显式空列表在两种模式下都会移除 effort 和 thinking 控制;未知列表仅在 adaptive 下移除这些控制。`reasoning.summary` 和其他非 effort 字段保持不变,已知非空目标继续按现有规则解析 effort。strict 的未知目标及普通 native Chat 的未知声明保留调用方控制。默认值填充不会覆盖现有 effort,但能力归一化可移除不支持的控制。 ## 图片 / 多模态能力 @@ -266,6 +267,7 @@ combo 会存储在顶层的 `combos` 对象中,并以 combo id 作为键: | `cooldownMs` | 否 | 未设置 → 上游回退值(请求速率限制代码为 `1302`/`1305` 的 429 为 5 秒,否则为 60 秒) | 1 到 600000 的整数。设置后,只要没有可用的上游 `Retry-After` 或 Codex 重置信号,就会作为每个目标的冷却时间应用,包括请求速率限制 429;未设置时使用上游回退值。 | | `waitForCooldownMs` | 否 | `0` | 0 到 600000 的整数。在返回 `combo_unavailable` 前等待最早恢复资格的冷却中目标的最长时间;请求中止会取消等待。 | | `defaultEffort` | 否 | `null` | `low`、`medium`、`high`、`xhigh`、`max` 或 `ultra`;仅当调用方省略 effort 且目标声明支持时才会应用。 | +| `reasoningEffortMode` | 否 | `"strict"` | `strict` 或 `adaptive`;选择混合能力交集和目标级控制归一化。 | | `imageInput` | 否 | `"auto"` | `"auto"` 或 `"disabled"`。`"auto"` 仅在每个目标都支持图片时发布图片能力;`"disabled"` 强制仅文本(从对外能力中去掉图片,并在分发前拒绝带图请求)。 | | `alias` | 否 | 无 | 可选的、已修剪的公开模型 id;使用上面的别名规则。空值会以“无别名”形式存储。 | | `nativeAlias` | 否 | `false` | 显式允许当前受支持的裸原生 alias 接管路由和 catalog 优先级;绝不会根据 alias 自动推断。 | diff --git a/docs-site/src/content/docs/zh-cn/guides/sub-agent-surface.md b/docs-site/src/content/docs/zh-cn/guides/sub-agent-surface.md index b0d1dc6536..390ad72d4a 100644 --- a/docs-site/src/content/docs/zh-cn/guides/sub-agent-surface.md +++ b/docs-site/src/content/docs/zh-cn/guides/sub-agent-surface.md @@ -172,3 +172,14 @@ curl -X PUT http://localhost:10100/api/injection-model \ ### 上下文上限 模型上下文上限与子代理模式无关。请在 Models 页面配置它;原生 OpenAI 模型会保留其真实的上下文窗口。 + +新配置不会写入实验性的 `plaintextV2AgentMessages` 字段,只有显式设置为 `true` 才会启用。 +调用方必须使用 Responses 格式,最终目标必须采用 `adapter: "openai-responses"`、 +`authMode: "forward"` 和准确的基础地址 `https://chatgpt.com/backend-api/codex`。OpenAI API key +provider、自定义兼容网关、最终发往其他 provider 的请求,以及非 Responses 调用都不会被改写。 +对于符合条件的新原生 ChatGPT v2 工具调用,该选项会临时改写 namespace 和三个保留工具名, +并删除消息字段的加密标记;响应返回 Codex 前会恢复原 namespace 与工具名。它处理 +`spawn_agent`、`send_message` 和 `followup_task`,不会增加恢复请求。 +HTTPS 仍会加密网络传输,但任务文字可能保存在 Codex 历史、外部模型请求和本地响应或调试文件中。 +已有密文不会改变。该选项依赖 ChatGPT 和 Codex 未公开的行为。详见 +[明文 v2 代理消息](/zh-cn/reference/configuration/agents/#明文-v2-代理消息)。 diff --git a/docs-site/src/content/docs/zh-cn/reference/architecture.md b/docs-site/src/content/docs/zh-cn/reference/architecture.md index 94c5eb8ea0..b1edbd1e8b 100644 --- a/docs-site/src/content/docs/zh-cn/reference/architecture.md +++ b/docs-site/src/content/docs/zh-cn/reference/architecture.md @@ -134,6 +134,13 @@ thread affinity 位于 `codex/` 下,不会出现在管理 API 响应中。请 session 中回退到 HTTP。设置 `"websockets": true` 后,同一 endpoint 会接受 upgrade 并使用 WebSocket bridge。 +当最终发送的模型为 `gpt-5.3-codex-spark` 时,canonical ChatGPT 转发会在 HTTP 请求头和 +原生 WS 帧元数据中明确关闭 Responses Lite,通过别名选择 Spark 时也一样;但这仅适用于发送正文 +不含带有非空 `tools` 数组的 `additional_tools` 分组的情况。该分组本身就是 Lite 的工具投递形态,因此仍使用它的 Spark +正文会保持 Lite 开启,无论调用方或配置的请求头如何。Lite 标识变化时, +旧 socket 会退出使用;后续标识相同且满足复用条件的请求可以复用新 socket。其他模型和网关 +保留原有 Lite 策略。原生元数据格式不合法时,仍会回退到 HTTP,并保持请求正文不变。 + Codex context compaction 同样适用于路由模型。`server/responses/compact.ts` 处理 `POST /v1/responses/compact`,运行一次内部路由 summarization turn 并返回压缩后的历史; `responses/parser.ts` 与 `bridge.ts` 则处理 remote compaction v2 的 `compaction_trigger` turn, diff --git a/docs-site/src/content/docs/zh-cn/reference/configuration/agents.md b/docs-site/src/content/docs/zh-cn/reference/configuration/agents.md index fcc84ab87f..dacc763c03 100644 --- a/docs-site/src/content/docs/zh-cn/reference/configuration/agents.md +++ b/docs-site/src/content/docs/zh-cn/reference/configuration/agents.md @@ -21,6 +21,7 @@ description: 多代理界面、委派引导、首选模型、回退链、原生 | `subagentModelFallbackPollMs?` | `number` | `60000` | 可用性探测缓存间隔。低于 1000 ms 的值会回退到默认值。 | | `effortCap?` | `string` | — | 对符合条件的 v2 主轮次和标记的派生子轮次设置硬上限。接受 `low` 到 `ultra`。 | | `subagentEffortCap?` | `string` | — | 仅针对派生子轮次的额外上限。两个上限同时适用时,较低者生效。 | +| `plaintextV2AgentMessages?` | `boolean` | —(未设置) | 实验性选项。只有显式设置为 `true` 才会启用。符合条件的新 `spawn_agent`、`send_message` 和 `followup_task` 调用会使用明文消息参数。详见[明文 v2 代理消息](#明文-v2-代理消息)。 | 通过仪表板或 `ocx v2 status|on|off|mode |threads ` 管理该界面。模式变更会应用于新会话。`maxConcurrentThreadsPerSession` 是 `PUT /api/v2` 字段,不是 `config.json` 键;`ocx v2 threads ` 会在启用 v2 后,将 `max_concurrent_threads_per_session` 写入 Codex 的 `$CODEX_HOME/config.toml` 中的 `[features.multi_agent_v2]` 下。 @@ -80,6 +81,42 @@ opencodex 会跳过已禁用、不可路由、不健康、处于冷却中,或 } ``` +## 明文 v2 代理消息 + +新配置不会写入 `plaintextV2AgentMessages`。只有显式设置为 `true` 才会启用。调用方必须使用 Responses +格式,最终目标必须采用规范的 ChatGPT Codex 转发配置,即 `adapter: "openai-responses"`、 +`authMode: "forward"` 和准确的基础地址 `https://chatgpt.com/backend-api/codex`。OpenAI API key +provider、自定义 OpenAI 兼容网关、最终发往其他 provider 的请求,以及非 Responses 调用都不会被改写。 + +对于符合条件的 v2 请求,opencodex 只识别顶层 `collaboration` namespace,而且它必须直接包含 +`spawn_agent`。原生 ChatGPT 收到请求前,opencodex 会删除 `spawn_agent`、`send_message` 和 +`followup_task` 中已有的 `parameters.properties.message.encrypted: true`。ChatGPT 会按保留的 +`collaboration` namespace 和三个工具名处理消息,因此请求会给这四个名称使用固定的临时别名。 +修改前,opencodex 会检查顶层和 `additional_tools` 工具目录、嵌套 namespace、 +`tool_search_output` 声明、`tool_choice` 和历史调用项。只要发现私有 namespace 或固定别名冲突, +整个请求就保持原样。opencodex 只会在 JSON、SSE 和 WebSocket 响应中恢复本次请求生成的别名, +并保留 `encrypted_function_args: []`,让兼容的 Codex 客户端把参数识别为明文。 + +这个选项不会增加恢复请求,也不会使用 `agentTaskRecovery` 在缓存未命中时产生的额外 ChatGPT +配额。它只能影响新工具调用,不能修改已有密文。请求已占用私有名称或有冲突引用时,opencodex 会保持该请求不变;若它后来生成加密的路由子任务,单独启用的 `agentTaskRecovery` 仍可处理。ChatGPT 拒绝或忽略修改后的 schema,或 Codex 客户端不识别明文响应字段时,调用可能失败。opencodex 不会用原 schema 自动重发父请求,因为重发可能重复消耗配额或重复执行工具。 + +每个响应的恢复检查最多处理 10,000 个身份位置。达到限制时,JSON 响应返回 HTTP 502,流式响应返回 +`response.failed`。这两种情况都不会把私有别名发给 Codex,也不会保存该响应供后续 +`previous_response_id` 继续使用。 + +成功改写后,这个选项会取消代理消息参数的应用层加密。HTTPS 仍会加密网络传输,但消息文字可能出现在 Codex +任务历史、外部模型请求、`responses-state.json` 及其 spill 文件,以及启用调试记录时的 +`usage-debug.jsonl`。该行为依赖 ChatGPT 未公开的 schema 和响应字段,后端或客户端更新后可能失效。 +服务启动时会打印警告。 + +```json +{ + "plaintextV2AgentMessages": true +} +``` + +等价命令是 `ocx config set plaintextV2AgentMessages true`。修改后重启代理。 + ## Effort 上限 上限只适用于 v2 协作功能:当主轮次的工具暴露 v2 时,它就符合条件;当子轮次在 `x-codex-turn-metadata` 中带有 codex-rs 的精确 `x-openai-subagent: collab_spawn` 或 `"subagent_kind": "thread_spawn"` 标记时,它也符合条件,即使叶子工具已经不再暴露协作。V1 主轮次、`multiAgentMode: "v1"`、压缩、审查以及记忆整合轮次都会绕过上限。 diff --git a/docs-site/src/content/docs/zh-cn/reference/configuration/routing.md b/docs-site/src/content/docs/zh-cn/reference/configuration/routing.md index 609c9ab48f..fa7d04ffc5 100644 --- a/docs-site/src/content/docs/zh-cn/reference/configuration/routing.md +++ b/docs-site/src/content/docs/zh-cn/reference/configuration/routing.md @@ -73,7 +73,8 @@ Codex Auth 页面将此 picker 行为作为选择加入项。关闭它会隐藏 | `targets` | `{ provider: string; model: string; weight?: number }[]` | required | 有序的具体路由。`weight` 范围为 1–10000,默认值为 `1`。 | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | 选择策略。目标顺序表示 `failover` 优先级;`weight` 决定 `round-robin` 和 `random` 的抽取权重;`least-used` 根据记录的成功次数选择;`reset-window` 跟随最近的额度重置。 | | `stickyLimit?` | `number` | `1` | 在单个轮询批次中保留的成功请求数。范围 1–100。 | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | 仅在调用方省略 effort 且所选目标声明了请求的档位时应用。 | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | unset | 当 combo 配置了非 null 默认值且目标支持列表已知且非空时,`defaultEffort` 会填充省略的 `reasoning.effort`。目标支持配置值时保留该值,否则选择不高于配置值的最高支持档位;若不存在更低档位,则使用最低支持档位。未知或空列表不会注入默认值。 | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` 对所有已知目标档位列表取交集,包括空列表;`"adaptive"` 排除空列表。未知列表在两种模式下都不限制目录交集。发送时,显式空列表在两种模式下都会移除 effort/thinking 控制;未知列表仅在 adaptive 下移除。`reasoning.summary` 保持不变。已知非空目标的 effort 解析、目标选择和顺序不变。 | | `imageInput?` | `"auto" \| "disabled"` | `"auto"` | `"auto"` 仅在每个目标都支持图片时发布图片能力;`"disabled"` 强制仅文本(从对外能力中去掉图片,并在分发前拒绝带图请求)。 | | `alias?` | `string` | — | 可选的公开 model id,用于替代规范化的选择器 slug。 | | `nativeAlias?` | `boolean` | `false` | 仅让当前受支持的裸原生 id 对该不带限定前缀的 id 优先;带账号或提供方限定的 OpenAI 路由仍是独立路由。 | diff --git a/docs-site/src/content/docs/zh-cn/reference/configuration/server.md b/docs-site/src/content/docs/zh-cn/reference/configuration/server.md index 4031f9ff19..8e8050de64 100644 --- a/docs-site/src/content/docs/zh-cn/reference/configuration/server.md +++ b/docs-site/src/content/docs/zh-cn/reference/configuration/server.md @@ -14,6 +14,7 @@ description: 监听、远程访问、准入密钥、超时、存储、侧车、 | `hostname?` | `string` | `"127.0.0.1"` | 绑定地址。非回环绑定需要 `OPENCODEX_API_AUTH_TOKEN`。 | | `proxy?` | `string` | — | 出站 HTTP(S) 代理 URL,或 `${ENV_VAR}`。仅当 `HTTP_PROXY` / `HTTPS_PROXY` 未设置时才会应用;回环地址始终保留在 `NO_PROXY` 中。 | | `emptyCompletionRetry?` | `boolean` | `false` | 显式启用:当 Responses turn 既无文本也无工具调用时,使用相同请求重试一次,包括流在终止事件之前结束的情况。重试可能产生费用。`OCX_EMPTY_COMPLETION_RETRY=0` 可在不修改配置的情况下禁用;combo 与 routed-compaction turn 不参与。 | +| `dropCodexSafetyBuffering?` | `boolean` | `false` | 从 Codex Responses 透传响应中移除 Codex safety-buffering 提示:`x-codex-safety-buffering-enabled` / `x-codex-safety-buffering-faster-model` 响应头、类型为 `safety_buffering` 的 `response.metadata` SSE 事件,以及其他 SSE 事件中的 `safety_buffering` 字段。Codex TUI 会将这些提示显示为“使用更快模型重试”的提示框,其默认操作会把会话切换到较弱的模型。其他 `x-codex-*` 响应头和其他所有 SSE 事件内容均保持不变,但会移除该字段。默认关闭。 | | `stallTimeoutSec?` | `number` | `300` | 在上游没有数据之前可等待的秒数,超过后返回 `response.incomplete`。最小值为 1。 | | `connectTimeoutMs?` | `number` | `200000` | 每次尝试的 DNS/TCP/TLS/最终响应头截止时间;它在正文生成之前结束。 | | `shutdownTimeoutMs?` | `number` | `5000` | 优雅停机截止时间,超过后会中止仍在进行中的请求。 | @@ -185,3 +186,5 @@ Anthropic OAuth 侧车会复用 opencodex 现有的 Claude Code OAuth 指纹。 ## Codex 额度网络诊断 主 Codex 账户行中的 `quotaRefresh` 描述额度查询结果,并不代表剩余额度或模型访问权限。读取缓存或未执行查询时,该字段可能省略。查询使用正在运行的代理服务的环境,而不是当前终端的环境。未设置 `proxy` 时保留现有环境;`"auto"` 只在启动时读取 Windows 静态代理设置,不自动处理 PAC/WPAD、仅 SOCKS 的设置或运行中的更改。TUN 测试成功并不能单独证明 HTTP 代理路径正常。命令和状态说明见[英文网络诊断章节](/reference/configuration/server/#codex-quota-network-diagnostics)。 + +`dropCodexSafetyBuffering`: 不会改变供应商安全策略或拒绝响应。原生 WebSocket `codex.response.metadata.headers` 和 `/responses/compact` 不在过滤范围内。 diff --git a/docs-site/src/content/docs/zh-tw/guides/combos.md b/docs-site/src/content/docs/zh-tw/guides/combos.md index ce3ad70a94..be566c77cd 100644 --- a/docs-site/src/content/docs/zh-tw/guides/combos.md +++ b/docs-site/src/content/docs/zh-tw/guides/combos.md @@ -177,15 +177,16 @@ Failover 是刻意受限的。它有助於目標特定的可用性、認證、 ## 預設推理 effort -`defaultEffort` 僅在以下全為真時提供 `reasoning.effort`: +當 combo 設定非 null 預設值且目標支援清單已知且非空時,`defaultEffort` 會補入省略的 `reasoning.effort`。目標支援設定值時保留該值,否則選擇不高於設定值的最高支援層級;若沒有更低層級,則使用最低支援層級。未知或空清單不會注入預設值。 -1. combo 有非 null 預設值; -2. 呼叫者未設定 effort;且 -3. 所選目標的目錄宣告該精確 effort。 +預設值補入會保留既有 effort 與其他 reasoning 欄位。下述能力正規化可另外移除不支援的 effort/thinking 控制。預設值支援 `low`、`medium`、`high`、`xhigh`、`max`、`ultra`;省略欄位或設為 `null` 可關閉注入。 -若請求沒有 `reasoning` 物件,opencodex 建立一個。若 `reasoning` 存在但無 `effort` 屬性,它保留其他欄位並加入預設值。呼叫者提供的 effort 永不被覆寫。 -當目標能力未知或不包含設定的 effort 時,opencodex 省略預設值並保持目標自身行為不變。支援的值為 `low`、`medium`、`high`、`xhigh`、`max` 與 `ultra`;省略欄位或設為 `null` 可將 effort 完全交給呼叫者與目標。 +## 混合 reasoning 能力 + +`reasoningEffortMode` 預設為 `"strict"`,發布所有目標 effort 清單的交集,包括明確空清單。`"adaptive"` 計算交集時排除空清單,讓混合 combo 保留選擇器。未知清單在兩種模式下都不限制目錄交集。 + +傳送時,明確空清單在兩種模式下都會移除 effort 與 thinking 控制;未知清單只在 adaptive 移除這些控制。`reasoning.summary` 與其他非 effort 欄位保持不變,已知非空目標仍按現有規則解析 effort。strict 的未知目標及一般 native Chat 的未知宣告保留呼叫者控制。預設值補入不會覆寫現有 effort,但能力正規化可移除不支援的控制。 ## 加密的 v2 子代理任務 @@ -269,6 +270,7 @@ Combo 儲存於頂層 `combos` 物件中,以 combo id 為 key: | `strategy` | 否 | `"failover"` | 可用值為 `"failover"`、`"round-robin"`、`"random"`、`"least-used"`、`"reset-window"`。 | | `stickyLimit` | 否 | `1` | 僅適用於 `round-robin`:每次選擇的成功請求數,1 到 100 的整數。 | | `defaultEffort` | 否 | `null` | `low`、`medium`、`high`、`xhigh`、`max` 或 `ultra`;僅在呼叫者省略 effort 且目標宣告支援時套用。 | +| `reasoningEffortMode` | 否 | `"strict"` | `strict` 或 `adaptive`;選擇混合能力交集及目標層級控制正規化。 | | `alias` | 否 | 無 | 可選的修剪後公開模型 id;使用上述別名規則。空值儲存為無別名。 | ## 疑難排解 diff --git a/docs-site/src/content/docs/zh-tw/reference/architecture.md b/docs-site/src/content/docs/zh-tw/reference/architecture.md index daf9bc9080..be0b2c5003 100644 --- a/docs-site/src/content/docs/zh-tw/reference/architecture.md +++ b/docs-site/src/content/docs/zh-tw/reference/architecture.md @@ -134,6 +134,14 @@ thread affinity 位於 `codex/` 下,不會出現在管理 API 回應中。請 session 中回退到 HTTP。設定 `"websockets": true` 後,同一 endpoint 會接受 upgrade 並使用 WebSocket bridge。 +當最終傳送的模型為 `gpt-5.3-codex-spark` 時,canonical ChatGPT 轉送會在 HTTP 請求標頭與 +原生 WS 訊框中繼資料中明確關閉 Responses Lite,透過別名選擇 Spark 時也一樣;但僅限於傳送本文 +不含帶有非空 `tools` 陣列的 `additional_tools` 群組的情況。該群組本身就是 Lite 的工具傳遞形態,因此仍使用它的 Spark +本文會保持 Lite 開啟,無論呼叫端或設定的標頭為何。Lite 識別值 +改變時,舊 socket 會停止使用;後續識別值相同且符合重用條件的請求可以重用新 socket。 +其他模型與閘道保留既有 Lite 政策。原生中繼資料格式不合法時,仍會退回 HTTP,並保持 +請求本文不變。 + Codex context compaction 同樣適用於路由模型。`server/responses/compact.ts` 處理 `POST /v1/responses/compact`,執行一次內部路由 summarization turn 並回傳壓縮後的歷史; `responses/parser.ts` 與 `bridge.ts` 則處理 remote compaction v2 的 `compaction_trigger` turn, diff --git a/docs-site/src/content/docs/zh-tw/reference/configuration/routing.md b/docs-site/src/content/docs/zh-tw/reference/configuration/routing.md index 13a4e2622e..2001cf14a9 100644 --- a/docs-site/src/content/docs/zh-tw/reference/configuration/routing.md +++ b/docs-site/src/content/docs/zh-tw/reference/configuration/routing.md @@ -56,7 +56,8 @@ Codex Auth 頁面將此 picker 行為作為選擇加入功能暴露。停用它 | `targets` | `{ provider: string; model: string; weight?: number }[]` | 必填 | 有序的具體路由。`weight` 為 1–10000,預設 `1`。 | | `strategy?` | `"failover" \| "round-robin" \| "random" \| "least-used" \| "reset-window"` | `"failover"` | 選擇策略。目標順序為 `failover` 優先序;`weight` 塑造 `round-robin` 與 `random` 抽選;`least-used` 依循已記錄的成功次數;`reset-window` 依循最早的配額重設。 | | `stickyLimit?` | `number` | `1` | 在一個 round-robin 批次中保留的成功請求數。範圍 1–100。 | -| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | 未設定 | 僅在呼叫者省略 effort 且所選目標廣告請求的階層時套用。 | +| `defaultEffort?` | `"low" \| "medium" \| "high" \| "xhigh" \| "max" \| "ultra" \| null` | 未設定 | 當 combo 設定非 null 預設值且目標支援清單已知且非空時,`defaultEffort` 會補入省略的 `reasoning.effort`。目標支援設定值時保留該值,否則選擇不高於設定值的最高支援層級;若沒有更低層級,則使用最低支援層級。未知或空清單不會注入預設值。 | +| `reasoningEffortMode?` | `"strict" \| "adaptive"` | `"strict"` | `"strict"` 對所有已知目標層級清單取交集,包括空清單;`"adaptive"` 排除空清單。未知清單在兩種模式下都不限制目錄交集。傳送時,明確空清單在兩種模式下都會移除 effort/thinking 控制;未知清單只在 adaptive 移除。`reasoning.summary` 保持不變。已知非空目標的 effort 解析、目標選擇及順序不變。 | | `alias?` | `string` | — | 可選的公開模型 id,取代標準 picker slug。 | | `nativeAlias?` | `boolean` | `false` | 讓目前支援的裸原生 id 僅對該未限定 id 取得優先。裸 `gpt-5.6-*` id 使用 Codex 池/Direct 憑證。帳號限定路由保持獨立。供應商限定路由(如 `openai-apikey/gpt-5.6-*`)使用其設定的 API-key 路由,且永不會落到原生別名。 | | `displayName?` | `string` | — | 僅顯示的目錄標籤,對原生別名為必填且非空。 | diff --git a/gui/src/i18n/de.ts b/gui/src/i18n/de.ts index db6b8c889f..07f9c5ad05 100644 --- a/gui/src/i18n/de.ts +++ b/gui/src/i18n/de.ts @@ -796,6 +796,7 @@ export const de: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth ratenbegrenzt (429)", "logs.detail.attempt.recovery.image413": "Bildnutzlast zu groß (413)", "logs.detail.attempt.recovery.emptyCompletion": "Wiederholung nach leerer Antwort", + "logs.detail.attempt.recovery.consoleGoUpload": "Console-Upload erneut versucht", "logs.detail.attempt.recovery.unknown": "Unbekannter Wiederherstellungsgrund", "logs.detail.reason.usage_missing": "Nutzung wurde nicht gemeldet.", "logs.detail.reason.usage_unsupported": "Dieser Anbieter meldet keine Nutzung.", diff --git a/gui/src/i18n/en.ts b/gui/src/i18n/en.ts index 74e4ded4d1..d19275be73 100644 --- a/gui/src/i18n/en.ts +++ b/gui/src/i18n/en.ts @@ -845,6 +845,7 @@ export const en = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth rate-limited (429)", "logs.detail.attempt.recovery.image413": "Image payload too large (413)", "logs.detail.attempt.recovery.emptyCompletion": "Empty completion retry", + "logs.detail.attempt.recovery.consoleGoUpload": "Console upload retry", "logs.detail.attempt.recovery.unknown": "Unknown recovery reason", "logs.detail.reason.usage_missing": "Usage was not reported.", "logs.detail.reason.usage_unsupported": "This provider does not report usage.", diff --git a/gui/src/i18n/fr.ts b/gui/src/i18n/fr.ts index 2bc3d5ab51..7b245b524f 100644 --- a/gui/src/i18n/fr.ts +++ b/gui/src/i18n/fr.ts @@ -821,6 +821,7 @@ export const fr: Record = { "logs.detail.attempt.recovery.transient5xx": "Erreur 5xx temporaire", "logs.detail.attempt.recovery.connectionReset": "Réinitialisation de la connexion", "logs.detail.attempt.recovery.emptyCompletion": "Nouvelle tentative après une réponse vide", + "logs.detail.attempt.recovery.consoleGoUpload": "Nouvelle tentative d’envoi Console", "logs.detail.attempt.recovery.oauth401": "Réauthentification OAuth", "logs.detail.attempt.recovery.key429": "Clé soumise à une limitation de débit (429)", "logs.detail.attempt.recovery.rateLimit429": "Limitation de débit (429)", diff --git a/gui/src/i18n/ja.ts b/gui/src/i18n/ja.ts index a787ab735b..551fbe3e69 100644 --- a/gui/src/i18n/ja.ts +++ b/gui/src/i18n/ja.ts @@ -758,6 +758,7 @@ export const ja: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth レート制限 (429)", "logs.detail.attempt.recovery.image413": "画像ペイロードが大きすぎます (413)", "logs.detail.attempt.recovery.emptyCompletion": "空の完了を再試行", + "logs.detail.attempt.recovery.consoleGoUpload": "Console アップロード再試行", "logs.detail.attempt.recovery.unknown": "不明なリカバリ理由", "logs.detail.reason.usage_missing": "使用量が報告されませんでした。", "logs.detail.reason.usage_unsupported": "このプロバイダーは使用量を報告しません。", diff --git a/gui/src/i18n/ko.ts b/gui/src/i18n/ko.ts index ccd2b06735..883ab327a1 100644 --- a/gui/src/i18n/ko.ts +++ b/gui/src/i18n/ko.ts @@ -827,6 +827,7 @@ export const ko: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth 요청 한도 초과 (429)", "logs.detail.attempt.recovery.image413": "이미지 페이로드가 너무 큼 (413)", "logs.detail.attempt.recovery.emptyCompletion": "빈 응답 재시도", + "logs.detail.attempt.recovery.consoleGoUpload": "Console 업로드 재시도", "logs.detail.attempt.recovery.unknown": "알 수 없는 복구 사유", "logs.detail.reason.usage_missing": "usage가 보고되지 않았습니다.", "logs.detail.reason.usage_unsupported": "이 프로바이더는 usage 보고를 지원하지 않습니다.", diff --git a/gui/src/i18n/ru.ts b/gui/src/i18n/ru.ts index ceb0c18a92..5d0eb4f527 100644 --- a/gui/src/i18n/ru.ts +++ b/gui/src/i18n/ru.ts @@ -813,6 +813,7 @@ export const ru: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth ограничен (429)", "logs.detail.attempt.recovery.image413": "Слишком большой размер изображения (413)", "logs.detail.attempt.recovery.emptyCompletion": "Повтор пустого завершения", + "logs.detail.attempt.recovery.consoleGoUpload": "Повтор загрузки Console", "logs.detail.attempt.recovery.unknown": "Неизвестная причина восстановления", "logs.detail.reason.usage_missing": "Данные об использовании не были сообщены.", "logs.detail.reason.usage_unsupported": "Этот провайдер не сообщает данные об использовании.", diff --git a/gui/src/i18n/tr.ts b/gui/src/i18n/tr.ts index 26a8f93d09..1d404f7ce2 100644 --- a/gui/src/i18n/tr.ts +++ b/gui/src/i18n/tr.ts @@ -832,6 +832,7 @@ export const tr: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth kısıtlandı (429)", "logs.detail.attempt.recovery.image413": "Görsel boyutu çok büyük (413)", "logs.detail.attempt.recovery.emptyCompletion": "Boş tamamlama yeniden denemesi", + "logs.detail.attempt.recovery.consoleGoUpload": "Console yüklemesi yeniden denendi", "logs.detail.attempt.recovery.unknown": "Bilinmeyen kurtarma nedeni", "logs.detail.reason.usage_missing": "Kullanım bildirilmedi.", "logs.detail.reason.usage_unsupported": "Bu sağlayıcı kullanım bildirmeyebilir.", diff --git a/gui/src/i18n/zh-TW.ts b/gui/src/i18n/zh-TW.ts index 1fb8387f4d..943a7224a2 100644 --- a/gui/src/i18n/zh-TW.ts +++ b/gui/src/i18n/zh-TW.ts @@ -2210,6 +2210,7 @@ export const zhTW: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth 被限流 (429)", "logs.detail.attempt.recovery.image413": "圖片承載過大 (413)", "logs.detail.attempt.recovery.emptyCompletion": "空白完成重試", + "logs.detail.attempt.recovery.consoleGoUpload": "Console 上傳重試", "logs.detail.attempt.recovery.unknown": "未知的復原原因", "logs.detail.estimate.provider_cost_overlay": "已使用供應商設定的價格覆蓋。", "logs.detail.estimate.priority_lower_bound": "無法取得已確認的 Priority 價格;目前顯示的估算是已知下限。", diff --git a/gui/src/i18n/zh.ts b/gui/src/i18n/zh.ts index 4e2cd831cc..fdd0b574ba 100644 --- a/gui/src/i18n/zh.ts +++ b/gui/src/i18n/zh.ts @@ -808,6 +808,7 @@ export const zh: Record = { "logs.detail.attempt.recovery.anthropicOauth429": "Anthropic OAuth 被限流 (429)", "logs.detail.attempt.recovery.image413": "图片载荷过大 (413)", "logs.detail.attempt.recovery.emptyCompletion": "空完成重试", + "logs.detail.attempt.recovery.consoleGoUpload": "Console 上传重试", "logs.detail.attempt.recovery.unknown": "未知的恢复原因", "logs.detail.reason.usage_missing": "未上报 usage。", "logs.detail.reason.usage_unsupported": "该提供方不支持上报 usage。", diff --git a/gui/src/pages/Logs.tsx b/gui/src/pages/Logs.tsx index 774efc455a..d7fc3ab5c4 100644 --- a/gui/src/pages/Logs.tsx +++ b/gui/src/pages/Logs.tsx @@ -113,7 +113,8 @@ type AttemptRecoveryKind = | "rate-limit-429" | "anthropic-oauth-429" | "image-413" - | "empty-completion"; + | "empty-completion" + | "console-go-upload-retry"; interface LogAttempt { ordinal: number; @@ -307,6 +308,7 @@ const RECOVERY_KIND_KEYS = { "anthropic-oauth-429": "logs.detail.attempt.recovery.anthropicOauth429", "image-413": "logs.detail.attempt.recovery.image413", "empty-completion": "logs.detail.attempt.recovery.emptyCompletion", + "console-go-upload-retry": "logs.detail.attempt.recovery.consoleGoUpload", } as const satisfies Record; /** Map a metric-unavailable reason to its i18n key. */ diff --git a/scripts/test-layout/layout.json b/scripts/test-layout/layout.json index c4363a7305..00558b39ec 100644 --- a/scripts/test-layout/layout.json +++ b/scripts/test-layout/layout.json @@ -994,6 +994,8 @@ "pi-path-contract.test.ts": "clients", "pinned-http.test.ts": "lib", "pinned-https-get.test.ts": "images", + "plaintext-v2-agent-messages-server.test.ts": "server", + "plaintext-v2-agent-messages.test.ts": "responses", "plan-video.test.ts": "videos", "plan.test.ts": "images", "policy-execution.test.ts": "routing", @@ -1093,6 +1095,7 @@ "responses-account-label.test.ts": "responses", "responses-compaction-routing.test.ts": "responses", "responses-compaction.test.ts": "responses", + "responses-console-go-upload-retry.test.ts": "responses", "responses-context-overflow.test.ts": "responses", "responses-custom-tool-guidance.test.ts": "responses", "responses-custom-tool-repair.test.ts": "responses", @@ -1116,10 +1119,10 @@ "responses-pool-401-refresh.test.ts": "responses", "responses-pool-refresh-attribution.test.ts": "responses", "responses-reasoning-summary-passthrough.test.ts": "responses", - "responses-reasoning-summary-rewrite.test.ts": "responses", "responses-routed-web-search-fields.test.ts": "responses", "responses-self-named-namespace-scrub.test.ts": "responses", "responses-shadow-intercept.test.ts": "responses", + "responses-show-thinking-summary.test.ts": "responses", "responses-snapshot-repair-server.test.ts": "responses", "responses-snapshot-repair.test.ts": "responses", "responses-state-write-amplification.test.ts": "responses", diff --git a/src/adapters/base.ts b/src/adapters/base.ts index 0affb33b55..7c00d6a3b1 100644 --- a/src/adapters/base.ts +++ b/src/adapters/base.ts @@ -91,6 +91,10 @@ export interface AdapterRequest { convertedRoutedToolSearchNames?: ReadonlySet; /** Upstream-only aliases for namespace tools flattened in this request. */ convertedRoutedNamespaceToolAliases?: ReadonlyMap; + /** Request-declared collaboration child names eligible for plaintext-v2 alias restoration. */ + plaintextV2AgentMessageToolNames?: ReadonlySet; + /** Collaboration message-tool names actually rewritten to fixed aliases in this request. */ + plaintextV2AgentMessageAliasedToolNames?: ReadonlySet; /** Upstream-only <=64-char aliases for Meta Muse tool names rewritten in this request. */ convertedMuseToolNameAliases?: ReadonlyMap; /** Releases observation of a serialized request body after its final fetch attempt settles. */ diff --git a/src/adapters/google-wire-compiler.ts b/src/adapters/google-wire-compiler.ts index 88c482ba7d..aa835e50b4 100644 --- a/src/adapters/google-wire-compiler.ts +++ b/src/adapters/google-wire-compiler.ts @@ -130,12 +130,20 @@ function compileGenerationConfig(value: unknown): JsonObject | undefined { ))].slice(0, 5); if (stopSequences.length > 0) out.stopSequences = stopSequences; } - if (isObject(value.thinkingConfig) && typeof value.thinkingConfig.thinkingLevel === "string") { - const raw = value.thinkingConfig.thinkingLevel.toLowerCase(); - const thinkingLevel = GOOGLE_THINKING_LEVELS.has(raw) - ? raw - : (["xhigh", "max", "ultra"].includes(raw) ? "high" : undefined); - if (thinkingLevel) out.thinkingConfig = { thinkingLevel }; + if (isObject(value.thinkingConfig)) { + const thinking: JsonObject = {}; + if (typeof value.thinkingConfig.thinkingLevel === "string") { + const raw = value.thinkingConfig.thinkingLevel.toLowerCase(); + const thinkingLevel = GOOGLE_THINKING_LEVELS.has(raw) + ? raw + : (["xhigh", "max", "ultra"].includes(raw) ? "high" : undefined); + if (thinkingLevel) thinking.thinkingLevel = thinkingLevel; + } + // The one key that makes Google return `thought: true` text. Cloud Code Assist serves + // thinking either way (thoughtsTokenCount stays non-zero) but withholds the text unless the + // request opts in, so dropping it here silently reinstates the missing-thinking behavior. + if (value.thinkingConfig.includeThoughts === true) thinking.includeThoughts = true; + if (Object.keys(thinking).length > 0) out.thinkingConfig = thinking; } if (Array.isArray(value.responseModalities)) { const valid = value.responseModalities.filter((m): m is string => typeof m === "string" && ["TEXT", "IMAGE", "AUDIO"].includes(m)); diff --git a/src/adapters/google.ts b/src/adapters/google.ts index 7fcc88ba59..518aca3903 100644 --- a/src/adapters/google.ts +++ b/src/adapters/google.ts @@ -568,13 +568,15 @@ function googleToolCallMetadataFromPart( * Keep that provider visibility bit authoritative here so the streaming and buffered parsers * cannot accidentally expose the same hidden reasoning through different event types. */ -function googlePartTextEvent(part: GoogleResponsePart): AdapterEvent | undefined { +function googlePartTextEvent(part: GoogleResponsePart, thoughtSummary = false): AdapterEvent | undefined { // A malformed scalar/object is not text and must not cross the AdapterEvent boundary. Dropping // only this optional field preserves the rest of the part without inventing assistant output by // coercion; an empty string keeps its existing no-event behavior. if (typeof part.text !== "string" || part.text.length === 0) return undefined; return part.thought === true - ? { type: "reasoning_raw_delta", text: part.text } + ? thoughtSummary + ? { type: "thinking_delta", thinking: part.text } + : { type: "reasoning_raw_delta", text: part.text } : { type: "text_delta", text: part.text }; } @@ -721,6 +723,7 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte // Per-request closure: resolveAdapter builds a fresh adapter per request (server.ts), so buildRequest // can stash the CCA model/session for parseStream's reasoning-replay observation. let antigravityModel: string | undefined; + let returnsThoughtSummaries = false; let antigravitySession: string | undefined; // Vertex returns the same opaque Gemini thought signatures as CCA, but its replay namespace // must stay transport-scoped: a signature minted by one Google backend must never be sent to @@ -795,6 +798,8 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte : provider.googleMode === "vertex" ? parsed.modelId : resolveDirectGeminiWireModelId(parsed.modelId, provider.directGeminiWireRenames !== false); + returnsThoughtSummaries = provider.googleMode === "cloud-code-assist" + && /^gemini-/.test(routedModelId) && !isImageCapableModel(parsed.modelId); // AI Studio's `-tiered` spelling is wire-only; CCA aliases may migrate to another generation. const identityModelId = provider.googleMode === "cloud-code-assist" ? routedModelId : parsed.modelId; const stripRejectedClaudeSdkParagraph = provider.googleMode === "cloud-code-assist" @@ -866,11 +871,20 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte ); antigravityModel = wireModelId; antigravitySession = sessionId; + // CCA Gemini exposes provider-authored thought summaries with includeThoughts. + // Other CCA model families do not share this request contract. + const includeThoughts = provider.showThinkingSummary === true + && parsed.options.hideThinkingSummary !== true + && /^gemini-/.test(wireModelId) + && !isImageCapableModel(parsed.modelId); // Effort → thinkingConfig for CCA (CLIProxyAPI proven: request.generationConfig.thinkingConfig). // Suffix/compat IDs return thinkingLevel=undefined — the suffix IS the effort, no contradiction. - if (thinkingLevel) { + if (thinkingLevel || includeThoughts) { const gc = (body.generationConfig ?? {}) as Record; - gc.thinkingConfig = { thinkingLevel }; + gc.thinkingConfig = { + ...(thinkingLevel ? { thinkingLevel } : {}), + ...(includeThoughts ? { includeThoughts: true } : {}), + }; body.generationConfig = gc; } // Reasoning continuity: Gemini models re-inject cached thoughtSignatures; Claude-on-Antigravity @@ -1139,7 +1153,7 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte if (part.thought === true && sig && isLikelyRealThoughtSignature(sig)) { pendingStreamThoughtSig = sig; } - const textEvent = googlePartTextEvent(part); + const textEvent = googlePartTextEvent(part, returnsThoughtSummaries); if (textEvent) { emittedContentEvent = true; yield textEvent; @@ -1415,7 +1429,7 @@ export function createGoogleAdapter(provider: OcxProviderConfig): ProviderAdapte if (part.thought === true && sig && isLikelyRealThoughtSignature(sig)) { pendingThoughtSig = sig; } - const textEvent = googlePartTextEvent(part); + const textEvent = googlePartTextEvent(part, returnsThoughtSummaries); if (textEvent) events.push(textEvent); const inline = (part as { inlineData?: { mimeType?: string; data?: string } }).inlineData; if (inline && typeof inline.data === "string") { diff --git a/src/adapters/openai-chat.ts b/src/adapters/openai-chat.ts index e338d845ad..e4136ecb3e 100644 --- a/src/adapters/openai-chat.ts +++ b/src/adapters/openai-chat.ts @@ -130,6 +130,10 @@ export function buildOpenAIChatPassthroughRequest( for (const field of CHAT_PASSTHROUGH_FIELDS) { if (rawBody[field] !== undefined) body[field] = rawBody[field]; } + const rawEfforts = modelRecordValue(provider.modelReasoningEfforts, modelId) ?? provider.reasoningEfforts; + if (modelInList(provider.noReasoningModels, modelId) || rawEfforts?.length === 0) { + delete body.reasoning_effort; + } const openRouterRouting = resolveOpenRouterRouting(provider, modelId); if (openRouterRouting) body.provider = openRouterProviderPayload(openRouterRouting); diff --git a/src/adapters/openai-responses.ts b/src/adapters/openai-responses.ts index 92de305f6d..5a5d370faf 100644 --- a/src/adapters/openai-responses.ts +++ b/src/adapters/openai-responses.ts @@ -24,6 +24,7 @@ import type { TranslatorBudget } from "../lib/translator-budget"; import { rewriteRoutedCustomToolsForUpstream } from "../responses/custom-tool-compat"; import { rewriteRoutedToolSearchForUpstream } from "../responses/tool-search-compat"; import { rewriteRoutedNamespaceToolsForUpstream } from "../responses/namespace-tool-compat"; +import { preparePlaintextV2AgentMessages } from "../responses/plaintext-v2-agent-messages"; import { isMetaAiResponsesDestination, rewriteMuseToolNamesForUpstream } from "../responses/muse-tool-name-alias"; import { openaiResponsesUrl } from "./openai-responses-url"; import { normalizeResponsesCodeMode } from "./responses-code-mode"; @@ -866,6 +867,21 @@ function promoteClientLoadedTools(body: unknown): unknown { } const MAX_RESPONSES_CALL_ID_LENGTH = 64; + +/** + * Whether the outgoing body still delivers tools through the responses-lite shape. + * + * Lite carries the client catalog as an `additional_tools` input item; the non-Lite wire shape + * expects top-level `tools`. Anything that flips the Lite advertisement has to agree with the + * shape actually being sent, or the destination silently loses the tool surface. + */ +function bodyCarriesLiteToolShape(body: Record): boolean { + if (!Array.isArray(body.input)) return false; + return body.input.some(item => + isPlainObject(item) && item.type === "additional_tools" + && Array.isArray(item.tools) && item.tools.length > 0 + ); +} const REPAIRED_CALL_ID_PREFIX = "call_ocx_"; const REPAIRED_CALL_ID_DIGEST_LENGTH = MAX_RESPONSES_CALL_ID_LENGTH - REPAIRED_CALL_ID_PREFIX.length; @@ -2371,6 +2387,8 @@ export function createResponsesPassthroughAdapter(provider: OcxProviderConfig): let routedCustomToolRepairNames: Set | undefined; let convertedRoutedToolSearchNames: Set | undefined; let convertedRoutedNamespaceToolAliases: Map | undefined; + let plaintextV2AgentMessageToolNames: ReadonlySet | undefined; + let plaintextV2AgentMessageAliasedToolNames: ReadonlySet | undefined; let convertedMuseToolNameAliases: Map | undefined; const unexpandedMiss = !!parsed.previousResponseId && parsed._previousResponseInputExpanded !== true; let outBody = stripPreviousResponseId( @@ -2489,6 +2507,14 @@ export function createResponsesPassthroughAdapter(provider: OcxProviderConfig): // Run after routed compaction so nested input_image parts are replaced before a malformed // tool output is flattened to text and can no longer be inspected structurally. outBody = repairUnidentifiedToolOutputItems(outBody); + if (parsed._plaintextV2AgentMessages === true && isCanonicalOpenAiForwardProvider(provider)) { + const prepared = preparePlaintextV2AgentMessages(outBody); + outBody = prepared.body; + if (prepared.namespaceAliased) { + plaintextV2AgentMessageToolNames = prepared.toolNames; + plaintextV2AgentMessageAliasedToolNames = prepared.aliasedAgentMessageToolNames; + } + } const threadServingIdentityChanged = parsed._stripReasoningEncryptedContent === true; const sanitizedBody = normalizeToolSchemas( stripSparkCompatibility( @@ -2516,7 +2542,7 @@ export function createResponsesPassthroughAdapter(provider: OcxProviderConfig): ), isXaiSchemaTarget(provider), ); - const finalBody = stripDisabledVerbosity( + const unnormalizedBody = stripDisabledVerbosity( stripDisabledReasoningSummaries( normalizeConfiguredReasoningSummaryDelivery(sanitizedBody, provider, parsed.modelId), provider, @@ -2525,13 +2551,32 @@ export function createResponsesPassthroughAdapter(provider: OcxProviderConfig): provider, parsed.modelId, ); + // Normalize the wire model before deriving model-dependent transport metadata. + const finalBody = + provider.modelSuffixBracketStrip + && unnormalizedBody !== null + && typeof unnormalizedBody === "object" + && !Array.isArray(unnormalizedBody) + && typeof (unnormalizedBody as { model?: unknown }).model === "string" + ? { ...(unnormalizedBody as Record), model: stripBracketedModelSuffix((unnormalizedBody as { model: string }).model) } + : unnormalizedBody; if (isCanonicalOpenAiForwardProvider(provider)) { - // Spark closes Responses Lite streams before a terminal completion. Select compatibility - // from the final wire model so aliases cannot leave the caller or a static header enabled. + // Select Spark's Lite compatibility from the final wire model, including aliases, and + // let the BODY decide it. The header also overrides native WS metadata downstream, so a + // forwarded or statically configured value must never contradict the shape being sent. + // + // The synchronized catalog keeps `use_responses_lite: true` for Spark precisely because + // it selects tool delivery (`input[].additional_tools` instead of top-level `tools`), and + // stripSparkCompatibility filters that group in place rather than promoting it. So a + // Lite-shaped body is pinned back ON — otherwise an inherited `false` advertises non-Lite + // while the tools exist only in the Lite shape, and Spark loses the tool surface. Only a + // body with no Lite tool group is downgraded, which is what the stream fix needs. if (isPlainObject(finalBody) && finalBody.model === "gpt-5.3-codex-spark") { + const liteShaped = bodyCarriesLiteToolShape(finalBody); for (const name of Object.keys(headers)) { if (name.toLowerCase() === CODEX_RESPONSES_LITE_HEADER) delete headers[name]; } + headers[CODEX_RESPONSES_LITE_HEADER] = liteShaped ? "true" : "false"; } const routingHeaders = new Headers(headers); applyCodexRoutingHint(routingHeaders, finalBody); @@ -2558,15 +2603,7 @@ export function createResponsesPassthroughAdapter(provider: OcxProviderConfig): // here, on the serialized body, not on the parsed selector. One place covers both the // HTTP and the WebSocket outbound, because the WS path transports this same request // instead of rebuilding it. - const body = JSON.stringify( - provider.modelSuffixBracketStrip - && finalBody !== null - && typeof finalBody === "object" - && !Array.isArray(finalBody) - && typeof (finalBody as { model?: unknown }).model === "string" - ? { ...(finalBody as Record), model: stripBracketedModelSuffix((finalBody as { model: string }).model) } - : finalBody, - ); + const body = JSON.stringify(finalBody); const releaseBodyObservation = translatorBudget.observeExternallyCapped( "passthrough_serialization", new TextEncoder().encode(body).byteLength, @@ -2581,6 +2618,8 @@ export function createResponsesPassthroughAdapter(provider: OcxProviderConfig): ...(routedCustomToolRepairNames ? { routedCustomToolRepairNames } : {}), ...(convertedRoutedToolSearchNames ? { convertedRoutedToolSearchNames } : {}), ...(convertedRoutedNamespaceToolAliases ? { convertedRoutedNamespaceToolAliases } : {}), + ...(plaintextV2AgentMessageToolNames ? { plaintextV2AgentMessageToolNames } : {}), + ...(plaintextV2AgentMessageAliasedToolNames ? { plaintextV2AgentMessageAliasedToolNames } : {}), ...(convertedMuseToolNameAliases ? { convertedMuseToolNameAliases } : {}), ...(tierLog ? { tierLog } : {}), }; diff --git a/src/bridge.ts b/src/bridge.ts index 3f5f529d3f..612b60ba36 100644 --- a/src/bridge.ts +++ b/src/bridge.ts @@ -663,16 +663,13 @@ export function bridgeToResponsesSSE( const closeCurrentRawReasoning = () => { if (!currentRawReasoning) return; rawReasoningForNextToolCall = currentRawReasoning.text; - emit("response.reasoning_summary_text.done", { - item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, summary_index: 0, text: currentRawReasoning.text, - }); - emit("response.reasoning_summary_part.done", { - item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, summary_index: 0, - part: { type: "summary_text", text: currentRawReasoning.text }, + emit("response.reasoning_text.done", { + item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, content_index: 0, text: currentRawReasoning.text, }); const item = { type: "reasoning", id: currentRawReasoning.itemId, - summary: [{ type: "summary_text", text: currentRawReasoning.text }], + summary: [] as never[], + content: [{ type: "reasoning_text", text: currentRawReasoning.text }], }; emit("response.output_item.done", { output_index: currentRawReasoning.outputIndex, item }); retainFinishedItem(item as OutputItem, currentRawReasoning.textBytes, "reasoning"); @@ -1111,10 +1108,6 @@ export function bridgeToResponsesSSE( const itemId = `rs_${uuid()}`; const item = { type: "reasoning", id: itemId, summary: [] as { type: string; text: string }[] }; emit("response.output_item.added", { output_index: outputIndex, item }); - emit("response.reasoning_summary_part.added", { - item_id: itemId, output_index: outputIndex, summary_index: 0, - part: { type: "summary_text", text: "" }, - }); currentRawReasoning = { itemId, outputIndex, text: "", textBytes: 0 }; } ({ value: currentRawReasoning.text, bytes: currentRawReasoning.textBytes } = appendString( @@ -1123,9 +1116,12 @@ export function bridgeToResponsesSSE( event.text, "reasoning", )); - emit("response.reasoning_summary_text.delta", { + // Raw reasoning (openai-chat reasoning_content, kiro tags) rides the CONTENT + // channel. Clients control raw-reasoning display; this text is not a + // provider-authored summary. + emit("response.reasoning_text.delta", { item_id: currentRawReasoning.itemId, output_index: currentRawReasoning.outputIndex, - summary_index: 0, delta: event.text, + content_index: 0, delta: event.text, }); break; } @@ -1783,7 +1779,8 @@ function buildResponseJSONWithBudget( } pushOutput({ type: "reasoning", id: `rs_${uuid()}`, - summary: [{ type: "summary_text", text: currentRawReasoning }], + summary: [], + content: [{ type: "reasoning_text", text: currentRawReasoning }], }, currentRawReasoningBytes, "reasoning"); currentRawReasoning = ""; currentRawReasoningBytes = 0; diff --git a/src/combos/request.ts b/src/combos/request.ts index abafccc525..63c5ba7fca 100644 --- a/src/combos/request.ts +++ b/src/combos/request.ts @@ -1,4 +1,4 @@ -import type { OcxComboDefaultEffort, OcxComboTarget, OcxConfig } from "../types"; +import type { OcxComboDefaultEffort, OcxComboReasoningEffortMode, OcxComboTarget, OcxConfig } from "../types"; import { resolveEffortAtOrBelow } from "../reasoning-effort"; import { resolveComboId } from "./types"; @@ -59,9 +59,14 @@ export function concreteComboRequestBody( target: Pick, defaultEffort: OcxComboDefaultEffort | null, targetReasoningEfforts: readonly string[] | undefined, + reasoningEffortMode: OcxComboReasoningEffortMode = "strict", ): Record { const clone = structuredClone(body) as Record; clone.model = `${target.provider}/${target.model}`; + if (targetReasoningEfforts?.length === 0 + || (reasoningEffortMode === "adaptive" && targetReasoningEfforts === undefined)) { + stripUnsupportedReasoningControls(clone); + } if (!defaultEffort) return clone; const reasoning = clone.reasoning; const needsDefault = reasoning === undefined || ( @@ -104,3 +109,16 @@ export function concreteComboRequestBody( } return clone; } + +function stripUnsupportedReasoningControls(body: Record): void { + const reasoning = body.reasoning; + if (reasoning && typeof reasoning === "object" && !Array.isArray(reasoning)) { + const next = { ...(reasoning as Record) }; + delete next.effort; + if (Object.keys(next).length > 0) body.reasoning = next; + else delete body.reasoning; + } + delete body.reasoning_effort; + delete body.thinking_budget; + delete body.thinking; +} diff --git a/src/config.ts b/src/config.ts index 5e81a7e5f1..932c26d8d8 100644 --- a/src/config.ts +++ b/src/config.ts @@ -1253,6 +1253,8 @@ const configSchema = z.object({ configRebaseProvenance: z.unknown().optional(), // A retry can be billable, so absence and malformed hand edits both stay off. emptyCompletionRetry: z.boolean().optional().catch(false), + // Header suppression changes what Codex sees, so absence and malformed edits stay off. + dropCodexSafetyBuffering: z.boolean().optional().catch(false), // A malformed hand edit must not silently stop opening the browser: fall back // to undefined, which resolves to the historical auto-open behavior. oauthOpenBrowser: z.boolean().optional().catch(undefined), @@ -1269,6 +1271,7 @@ const configSchema = z.object({ contextCapValue: z.number().int().positive().optional(), multiAgentGuidanceEnabled: z.boolean().optional(), // Invalid optional recovery config must not discard unrelated provider/account state. + plaintextV2AgentMessages: z.boolean().optional().catch(undefined), agentTaskRecovery: agentTaskRecoverySchema.optional().catch(undefined), // Same rationale: a bad notify section must not cost the operator their providers. quotaResetNotify: quotaResetNotifySchema.optional().catch(undefined), @@ -2159,6 +2162,17 @@ function warnDegradedUpstreamHostCircuitThreshold(rawParsed: unknown): void { if (warning) console.warn(`⚠️ config.json ${warning}. Other settings were preserved.`); } +function malformedPlaintextV2AgentMessagesWarning(value: unknown): string | null { + const raw = rawConfigRecord(value); + if (!raw || raw.plaintextV2AgentMessages === undefined || typeof raw.plaintextV2AgentMessages === "boolean") return null; + return "plaintextV2AgentMessages ignored: expected a boolean"; +} + +function warnDegradedPlaintextV2AgentMessages(value: unknown): void { + const warning = malformedPlaintextV2AgentMessagesWarning(value); + if (warning) console.warn(`⚠️ config.json ${warning}. Other settings were preserved.`); +} + function malformedAgentTaskRecoveryWarning(rawParsed: unknown): string | null { const raw = rawConfigRecord(rawParsed); if (!raw || !Object.hasOwn(raw, "agentTaskRecovery")) return null; @@ -2416,6 +2430,7 @@ export function loadConfig(): OcxConfig { warnDegradedNativeSubagentConfig(parsed, config); warnDegradedCodexAccountPicker(parsed); warnDegradedUpstreamHostCircuitThreshold(parsed); + warnDegradedPlaintextV2AgentMessages(parsed); warnDegradedAgentTaskRecovery(parsed); warnDegradedRuntimeRole(parsed); warnDegradedOptionalRemoteBlocks(parsed); @@ -2445,6 +2460,7 @@ export function loadConfig(): OcxConfig { warnDegradedNativeSubagentConfig(parsed, config); warnDegradedCodexAccountPicker(parsed); warnDegradedUpstreamHostCircuitThreshold(parsed); + warnDegradedPlaintextV2AgentMessages(parsed); warnDegradedAgentTaskRecovery(parsed); warnDegradedRuntimeRole(parsed); warnDegradedOptionalRemoteBlocks(parsed); @@ -2470,6 +2486,7 @@ export function loadConfig(): OcxConfig { warnDegradedNativeSubagentConfig(parsed, config); warnDegradedCodexAccountPicker(parsed); warnDegradedUpstreamHostCircuitThreshold(parsed); + warnDegradedPlaintextV2AgentMessages(parsed); warnDegradedAgentTaskRecovery(parsed); warnDegradedRuntimeRole(parsed); warnDegradedOptionalRemoteBlocks(parsed); @@ -2621,6 +2638,8 @@ function validFileConfigDiagnostics(config: OcxConfig, rawParsed: unknown): Conf if (notifyWarning) warnings.push(notifyWarning); const codexPoolWarning = malformedCodexPoolWarning(rawParsed); if (codexPoolWarning) warnings.push(codexPoolWarning); + const plaintextWarning = malformedPlaintextV2AgentMessagesWarning(rawParsed); + if (plaintextWarning) warnings.push(plaintextWarning); if (syncDisabledReason) { warnings.push(`syncCodexSubagentDefaults ignored: ${syncDisabledReason}`); } @@ -2702,6 +2721,12 @@ function upstreamHostCircuitThresholdError(value: unknown): string | null { return `schema_invalid: upstreamHostCircuitThreshold: must be an integer from 0 to ${UPSTREAM_HOST_CIRCUIT_MAX_THRESHOLD}`; } +function plaintextV2AgentMessagesError(value: unknown): string | null { + return malformedPlaintextV2AgentMessagesWarning(value) + ? "schema_invalid: plaintextV2AgentMessages: must be a boolean or omitted" + : null; +} + function agentTaskRecoveryError(value: unknown): string | null { const raw = rawConfigRecord(value); if (!raw || !Object.hasOwn(raw, "agentTaskRecovery") || raw.agentTaskRecovery === undefined) return null; @@ -2858,6 +2883,14 @@ function emptyCompletionRetryError(value: unknown): string | null { return "schema_invalid: emptyCompletionRetry: must be a boolean or omitted"; } +function dropCodexSafetyBufferingError(value: unknown): string | null { + const raw = rawConfigRecord(value); + if (!raw || !Object.hasOwn(raw, "dropCodexSafetyBuffering")) return null; + const enabled = raw.dropCodexSafetyBuffering; + if (enabled === undefined || typeof enabled === "boolean") return null; + return "schema_invalid: dropCodexSafetyBuffering: must be a boolean or omitted"; +} + function oauthOpenBrowserError(value: unknown): string | null { const raw = rawConfigRecord(value); if (!raw || !Object.hasOwn(raw, "oauthOpenBrowser")) return null; @@ -2993,6 +3026,7 @@ export function validateConfigCandidate(value: unknown): { ok: true; config: Ocx ?? claudeSubagentEffortError(value) ?? appOwnedMemoryBudgetError(value) ?? upstreamHostCircuitThresholdError(value) + ?? plaintextV2AgentMessagesError(value) ?? agentTaskRecoveryError(value) ?? quotaResetNotifyError(value) ?? codexPoolError(value) @@ -3001,6 +3035,7 @@ export function validateConfigCandidate(value: unknown): { ok: true; config: Ocx ?? codexQuotaAutoRefreshError(value) ?? codexAccountPickerEnabledError(value) ?? emptyCompletionRetryError(value) + ?? dropCodexSafetyBufferingError(value) ?? oauthOpenBrowserError(value) ?? runtimeRoleError(value) ?? remoteGuiConfigError(value) @@ -4022,6 +4057,7 @@ export function getDefaultConfig(): OcxConfig { return { port: 10100, emptyCompletionRetry: false, + dropCodexSafetyBuffering: false, fastRows: true, managementUsageMaxReadBytes: 64 * 1024 * 1024, appOwnedMemoryBudgetMb: DEFAULT_APP_OWNED_MEMORY_BUDGET_BYTES / (1024 * 1024), diff --git a/src/providers/derive.ts b/src/providers/derive.ts index 72a662aee4..3ab1d01500 100644 --- a/src/providers/derive.ts +++ b/src/providers/derive.ts @@ -44,6 +44,7 @@ export interface DerivedKeyLoginProvider { autoToolChoiceOnlyModels?: string[]; preserveReasoningContentModels?: string[]; requiresReasoningPlaceholderModels?: string[]; + showThinkingSummary?: boolean; reasoningSplitModels?: string[]; reasoningDetailsModels?: string[]; thinkingToggleModels?: string[]; @@ -271,6 +272,7 @@ export function providerConfigSeed(entry: ProviderRegistryEntry): OcxProviderCon ...(entry.autoToolChoiceOnlyModels ? { autoToolChoiceOnlyModels: [...entry.autoToolChoiceOnlyModels] } : {}), ...(entry.preserveReasoningContentModels ? { preserveReasoningContentModels: [...entry.preserveReasoningContentModels] } : {}), ...(entry.requiresReasoningPlaceholderModels ? { requiresReasoningPlaceholderModels: [...entry.requiresReasoningPlaceholderModels] } : {}), + ...(entry.showThinkingSummary !== undefined ? { showThinkingSummary: entry.showThinkingSummary } : {}), ...(entry.reasoningSplitModels ? { reasoningSplitModels: [...entry.reasoningSplitModels] } : {}), ...(entry.reasoningDetailsModels ? { reasoningDetailsModels: [...entry.reasoningDetailsModels] } : {}), ...(entry.thinkingToggleModels ? { thinkingToggleModels: [...entry.thinkingToggleModels] } : {}), @@ -320,6 +322,7 @@ export function deriveKeyLoginMap(): Record { ...(entry.autoToolChoiceOnlyModels ? { autoToolChoiceOnlyModels: [...entry.autoToolChoiceOnlyModels] } : {}), ...(entry.preserveReasoningContentModels ? { preserveReasoningContentModels: [...entry.preserveReasoningContentModels] } : {}), ...(entry.requiresReasoningPlaceholderModels ? { requiresReasoningPlaceholderModels: [...entry.requiresReasoningPlaceholderModels] } : {}), + ...(entry.showThinkingSummary !== undefined ? { showThinkingSummary: entry.showThinkingSummary } : {}), ...(entry.reasoningSplitModels ? { reasoningSplitModels: [...entry.reasoningSplitModels] } : {}), ...(entry.reasoningDetailsModels ? { reasoningDetailsModels: [...entry.reasoningDetailsModels] } : {}), ...(entry.thinkingToggleModels ? { thinkingToggleModels: [...entry.thinkingToggleModels] } : {}), @@ -574,6 +577,7 @@ export function enrichProviderFromRegistry(name: string, prov: OcxProviderConfig if (!prov.thinkingToggleModels && seed.thinkingToggleModels) prov.thinkingToggleModels = [...seed.thinkingToggleModels]; if (!prov.thinkingBudgetModels && seed.thinkingBudgetModels) prov.thinkingBudgetModels = [...seed.thinkingBudgetModels]; if (prov.escapeBuiltinToolNames === undefined && seed.escapeBuiltinToolNames !== undefined) prov.escapeBuiltinToolNames = seed.escapeBuiltinToolNames; + if (prov.showThinkingSummary === undefined && seed.showThinkingSummary !== undefined) prov.showThinkingSummary = seed.showThinkingSummary; if (prov.keyOptional === undefined && seed.keyOptional !== undefined) prov.keyOptional = seed.keyOptional; if (prov.freeTier === undefined && seed.freeTier !== undefined) prov.freeTier = seed.freeTier; if (prov.modelSuffixBracketStrip === undefined && seed.modelSuffixBracketStrip !== undefined) prov.modelSuffixBracketStrip = seed.modelSuffixBracketStrip; diff --git a/src/providers/opencode-zen-rate-limit.ts b/src/providers/opencode-zen-rate-limit.ts index c4dbb10319..383c9f70ae 100644 --- a/src/providers/opencode-zen-rate-limit.ts +++ b/src/providers/opencode-zen-rate-limit.ts @@ -175,3 +175,61 @@ export function enrichOpenCodeZenUpstreamMessage( ): string { return enrichOpenCodeZenFreeTierMessage(enrichOpenCodeZenRateLimitMessage(message, opts), opts); } + +/** The effective HTTP endpoint, never the configured row name, identifies Console. */ +export function isConsoleGoDestination(outboundUrl: string | undefined): boolean { + if (!outboundUrl) return false; + try { + const url = new URL(outboundUrl); + return url.protocol === "https:" && url.hostname === "opencode.ai" + && url.port === "" && url.username === "" && url.password === "" + && url.search === "" && url.hash === "" + && /^\/zen\/(?:go\/)?v1\/(?:responses|chat\/completions|messages)$/.test(url.pathname); + } catch { + return false; + } +} + +/** + * The canonical refusal envelope, as served on both Console routes: + * {"error":{"param":null,"type":"invalid_request_error","message":"Error from provider + * (Console Go): Upstream request failed: [invalid_request_error] Invalid upload request."}} + * + * The gateway names itself Console on the Zen key route and Console Go on the Go route, so the + * anchor is the shared product name plus the refusal sentence. The message is matched whole: a + * bare string, a suffix, or any other envelope is a different refusal and must not be replayed. + * Being stricter than necessary is the safe direction: a missed match leaves the turn failing + * exactly as it does today, while a loose match spends an extra request on unrelated 400s. + */ +const CONSOLE_UPLOAD_REFUSALS = new Set([ + "Error from provider (Console Go): Upstream request failed: [invalid_request_error] Invalid upload request.", + "Error from provider (Console): Upstream request failed: [invalid_request_error] Invalid upload request.", +]); + +/** + * True only for the canonical Console upload refusal on a canonical Console destination. + * Route-gated on purpose: the message alone would let any other upstream that happens to answer + * with this English sentence trigger a second send from an unrelated provider. + */ +export function isTransientConsoleGoUploadRejection(opts: { + status: number; + errorBody: string | undefined; + outboundUrl?: string; +}): boolean { + if (opts.status !== 400 || !opts.errorBody) return false; + if (!isConsoleGoDestination(opts.outboundUrl)) return false; + let payload: unknown; + try { + payload = JSON.parse(opts.errorBody); + } catch { + return false; + } + const error = (payload as { error?: unknown } | null)?.error; + if (!error || typeof error !== "object" || Array.isArray(error)) return false; + // The whole envelope, not just the sentence: Console always answers this refusal as + // invalid_request_error with a null param, so a partial envelope is a different error. + const envelope = error as { param?: unknown; type?: unknown; message?: unknown }; + if (envelope.type !== "invalid_request_error" || envelope.param !== null) return false; + // Matched without trimming: padding means the gateway wrapped or appended something. + return typeof envelope.message === "string" && CONSOLE_UPLOAD_REFUSALS.has(envelope.message); +} diff --git a/src/providers/registry.ts b/src/providers/registry.ts index da25a1e7d3..ff065b0561 100644 --- a/src/providers/registry.ts +++ b/src/providers/registry.ts @@ -356,6 +356,10 @@ export interface ProviderRegistryEntry { autoToolChoiceOnlyModels?: string[]; preserveReasoningContentModels?: string[]; requiresReasoningPlaceholderModels?: string[]; + /** + * Opt this provider into visible thinking summaries (see OcxProviderConfig.showThinkingSummary). + */ + showThinkingSummary?: boolean; reasoningSplitModels?: string[]; reasoningDetailsModels?: string[]; thinkingToggleModels?: string[]; @@ -380,7 +384,7 @@ export type ProviderConfigSeed = Pick< | "modelMaxInputTokens" | "defaultMaxOutputTokens" | "modelMaxOutputTokens" | "reasoningEfforts" | "modelReasoningEfforts" | "modelDefaultReasoningEfforts" | "reasoningEffortMap" | "modelReasoningEffortMap" | "reasoningWireFormat" | "noVisionModels" | "noReasoningModels" | "noTemperatureModels" | "noTopPModels" | "noPenaltyModels" - | "autoToolChoiceOnlyModels" | "preserveReasoningContentModels" | "requiresReasoningPlaceholderModels" | "reasoningSplitModels" | "reasoningDetailsModels" | "thinkingToggleModels" | "thinkingBudgetModels" | "escapeBuiltinToolNames" | "openaiChatEofTolerance" + | "autoToolChoiceOnlyModels" | "preserveReasoningContentModels" | "requiresReasoningPlaceholderModels" | "reasoningSplitModels" | "reasoningDetailsModels" | "thinkingToggleModels" | "thinkingBudgetModels" | "escapeBuiltinToolNames" | "openaiChatEofTolerance" | "showThinkingSummary" | "googleMode" | "project" | "location" | "headers" >; @@ -2135,7 +2139,7 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [ // path must stay RELATIVE: this row sets `allowBaseUrlOverride`, and an absolute `url` would // retarget a user's custom base back to Google. A leading `./` is required because a bare // `v1internal:` reads as a URL scheme and `providerModelDiscoverySpecError` rejects it. - { id: "google-antigravity", alias: "agy", label: "Google Antigravity", adapter: "google", baseUrl: "https://daily-cloudcode-pa.googleapis.com", authKind: "oauth", allowBaseUrlOverride: true, dashboardUrl: "https://antigravity.google", models: ANTIGRAVITY_MODELS, liveModels: true, defaultModel: "gemini-3.8-flash", modelContextWindows: ANTIGRAVITY_MODEL_CONTEXT_WINDOWS, modelInputModalities: ANTIGRAVITY_MODEL_INPUT_MODALITIES, modelReasoningEfforts: ANTIGRAVITY_MODEL_EFFORTS, googleMode: "cloud-code-assist", jawcodeBundle: "google", extraMetadataAliases: ["antigravity", "gemini-antigravity"], modelDiscovery: { path: "./v1internal:fetchAvailableModels" } }, + { id: "google-antigravity", alias: "agy", label: "Google Antigravity", adapter: "google", baseUrl: "https://daily-cloudcode-pa.googleapis.com", authKind: "oauth", allowBaseUrlOverride: true, dashboardUrl: "https://antigravity.google", models: ANTIGRAVITY_MODELS, liveModels: true, defaultModel: "gemini-3.8-flash", modelContextWindows: ANTIGRAVITY_MODEL_CONTEXT_WINDOWS, modelInputModalities: ANTIGRAVITY_MODEL_INPUT_MODALITIES, modelReasoningEfforts: ANTIGRAVITY_MODEL_EFFORTS, googleMode: "cloud-code-assist", showThinkingSummary: true, jawcodeBundle: "google", extraMetadataAliases: ["antigravity", "gemini-antigravity"], modelDiscovery: { path: "./v1internal:fetchAvailableModels" } }, { id: "azure-openai", label: "Azure OpenAI", adapter: "azure-openai", baseUrl: "https://{resource}.openai.azure.com/openai", authKind: "key", featured: true, dashboardUrl: "https://portal.azure.com" }, { id: "ollama", label: "Ollama (local)", adapter: "openai-chat", baseUrl: "http://localhost:11434/v1", authKind: "local", allowPrivateNetworkByDefault: true, allowBaseUrlOverride: true, featured: true, note: "Local — key usually blank" }, { id: "vllm", label: "vLLM (local)", adapter: "openai-chat", baseUrl: "http://localhost:8000/v1", authKind: "local", allowPrivateNetworkByDefault: true, allowBaseUrlOverride: true, featured: true, note: "Local — key usually blank" }, diff --git a/src/responses/plaintext-v2-agent-messages.ts b/src/responses/plaintext-v2-agent-messages.ts new file mode 100644 index 0000000000..d1b7fd12dc --- /dev/null +++ b/src/responses/plaintext-v2-agent-messages.ts @@ -0,0 +1,902 @@ +const COLLABORATION_NAMESPACE = "collaboration"; +export const PLAINTEXT_V2_COLLABORATION_NAMESPACE = "collaboration-optimize"; +const COLLABORATION_NAME_PREFIX = `${COLLABORATION_NAMESPACE}__`; +const COLLABORATION_DOTTED_NAME_PREFIX = `${COLLABORATION_NAMESPACE}.`; +const PLAINTEXT_V2_COLLABORATION_NAME_PREFIX = `${PLAINTEXT_V2_COLLABORATION_NAMESPACE}__`; +const PLAINTEXT_V2_COLLABORATION_DOTTED_NAME_PREFIX = `${PLAINTEXT_V2_COLLABORATION_NAMESPACE}.`; + +const PLAINTEXT_V2_AGENT_MESSAGE_TOOLS = new Set([ + "spawn_agent", + "send_message", + "followup_task", +]); + +const PLAINTEXT_V2_AGENT_MESSAGE_TOOL_ALIASES = new Map([ + ["spawn_agent", "start_delegated_task"], + ["send_message", "deliver_delegated_message"], + ["followup_task", "continue_delegated_task"], +]); + +const PLAINTEXT_V2_AGENT_MESSAGE_TOOL_NAMES = new Map( + [...PLAINTEXT_V2_AGENT_MESSAGE_TOOL_ALIASES].map(([name, alias]) => [alias, name]), +); + +export function shouldPreparePlaintextV2AgentMessages(args: { + enabled: boolean; + inboundWire: string; + canonicalChatGpt: boolean; + requestBody: unknown; +}): boolean { + return args.enabled + && args.inboundWire === "responses" + && args.canonicalChatGpt + && hasPlaintextV2CollaborationCatalog(args.requestBody); +} + +function isPlainObject(value: unknown): value is Record { + return value !== null && typeof value === "object" && !Array.isArray(value); +} + +function responseToolCatalogs(body: Record): unknown[][] { + const catalogs: unknown[][] = []; + if (Array.isArray(body.tools)) catalogs.push(body.tools); + if (!Array.isArray(body.input)) return catalogs; + for (const item of body.input) { + if ( + isPlainObject(item) + && item.type === "additional_tools" + && Array.isArray(item.tools) + ) { + catalogs.push(item.tools); + } + } + return catalogs; +} + +function collaborationCatalogInfo(catalogs: readonly unknown[][]): { + hasV2Catalog: boolean; + toolNames: Set; + aliasedAgentMessageToolNames: Set; +} { + let hasV2Catalog = false; + const toolNames = new Set(); + const aliasedAgentMessageToolNames = new Set(); + for (const tools of catalogs) { + for (const tool of tools) { + if ( + !isPlainObject(tool) + || tool.type !== "namespace" + || tool.name !== COLLABORATION_NAMESPACE + || !Array.isArray(tool.tools) + ) { + continue; + } + for (const child of tool.tools) { + if ( + isPlainObject(child) + && (child.type === "function" || child.type === "custom") + && typeof child.name === "string" + ) { + toolNames.add(child.name); + if ( + child.type === "function" + && PLAINTEXT_V2_AGENT_MESSAGE_TOOL_ALIASES.has(child.name) + ) { + aliasedAgentMessageToolNames.add(child.name); + } + if (child.type === "function" && child.name === "spawn_agent") hasV2Catalog = true; + } + } + } + } + return { hasV2Catalog, toolNames, aliasedAgentMessageToolNames }; +} + +export function hasPlaintextV2CollaborationCatalog(body: unknown): boolean { + if (!isPlainObject(body)) return false; + return Array.isArray(body.tools) && collaborationCatalogInfo([body.tools]).hasV2Catalog; +} + +function hasOptimizedNamespaceConflict(catalogs: readonly unknown[][]): boolean { + const pending = [...catalogs]; + while (pending.length > 0) { + const tools = pending.pop()!; + for (const tool of tools) { + if (!isPlainObject(tool)) continue; + if ( + typeof tool.name === "string" + && ( + tool.name === PLAINTEXT_V2_COLLABORATION_NAMESPACE + || tool.name.startsWith(PLAINTEXT_V2_COLLABORATION_NAME_PREFIX) + || tool.name.startsWith(PLAINTEXT_V2_COLLABORATION_DOTTED_NAME_PREFIX) + ) + ) { + return true; + } + if (tool.type === "namespace" && Array.isArray(tool.tools)) pending.push(tool.tools); + } + } + return false; +} + +function isToolIdentity(value: Record): boolean { + return value.type === "function" + || value.type === "custom" + || value.type === "function_call" + || value.type === "custom_tool_call"; +} + +function isOptimizedToolIdentity(value: unknown): boolean { + if (!isPlainObject(value)) return false; + if ( + isToolIdentity(value) + && ( + value.namespace === PLAINTEXT_V2_COLLABORATION_NAMESPACE + || ( + typeof value.name === "string" + && ( + value.name === PLAINTEXT_V2_COLLABORATION_NAMESPACE + || value.name.startsWith(PLAINTEXT_V2_COLLABORATION_NAME_PREFIX) + || value.name.startsWith(PLAINTEXT_V2_COLLABORATION_DOTTED_NAME_PREFIX) + ) + ) + ) + ) { + return true; + } + return value.type === "namespace" && value.name === PLAINTEXT_V2_COLLABORATION_NAMESPACE; +} + +function hasOptimizedReferenceConflict(body: Record): boolean { + if (isOptimizedToolIdentity(body.tool_choice)) return true; + if ( + isPlainObject(body.tool_choice) + && Array.isArray(body.tool_choice.tools) + && body.tool_choice.tools.some(isOptimizedToolIdentity) + ) { + return true; + } + if (!Array.isArray(body.input)) return false; + return body.input.some(item => ( + isPlainObject(item) + && (item.type === "function_call" || item.type === "custom_tool_call") + && isOptimizedToolIdentity(item) + )); +} + +function hasToolSearchCollaborationConflict(body: Record): boolean { + if (!Array.isArray(body.input)) return false; + for (const item of body.input) { + if (!isPlainObject(item) || item.type !== "tool_search_output" || !Array.isArray(item.tools)) { + continue; + } + const pending = [item.tools]; + while (pending.length > 0) { + const tools = pending.pop()!; + for (const tool of tools) { + if (!isPlainObject(tool)) continue; + if ( + typeof tool.name === "string" + && (tool.name === COLLABORATION_NAMESPACE + || tool.name === PLAINTEXT_V2_COLLABORATION_NAMESPACE + || tool.name.startsWith(COLLABORATION_NAME_PREFIX) + || tool.name.startsWith(COLLABORATION_DOTTED_NAME_PREFIX) + || tool.name.startsWith(PLAINTEXT_V2_COLLABORATION_NAME_PREFIX) + || tool.name.startsWith(PLAINTEXT_V2_COLLABORATION_DOTTED_NAME_PREFIX)) + ) { + return true; + } + if (tool.type === "namespace" && Array.isArray(tool.tools)) pending.push(tool.tools); + } + } + } + return false; +} + +function hasAgentMessageEncryptionMarker(tool: Record): boolean { + return tool.type === "function" + && typeof tool.name === "string" + && PLAINTEXT_V2_AGENT_MESSAGE_TOOLS.has(tool.name) + && isPlainObject(tool.parameters) + && isPlainObject(tool.parameters.properties) + && isPlainObject(tool.parameters.properties.message) + && tool.parameters.properties.message.encrypted === true; +} + +function rewriteAgentMessageToolDeclaration(tool: Record): Record { + const alias = tool.type === "function" && typeof tool.name === "string" + ? PLAINTEXT_V2_AGENT_MESSAGE_TOOL_ALIASES.get(tool.name) + : undefined; + if (!alias) return tool; + + let rewritten: Record = { ...tool, name: alias }; + if ( + hasAgentMessageEncryptionMarker(tool) + && isPlainObject(tool.parameters) + && isPlainObject(tool.parameters.properties) + && isPlainObject(tool.parameters.properties.message) + ) { + const { encrypted: _encrypted, ...messageSchema } = tool.parameters.properties.message; + rewritten = { + ...rewritten, + parameters: { + ...tool.parameters, + properties: { + ...tool.parameters.properties, + message: messageSchema, + }, + }, + }; + } + return rewritten; +} + +function hasPrivateToolName(name: string): boolean { + return name === PLAINTEXT_V2_COLLABORATION_NAMESPACE + || name.startsWith(PLAINTEXT_V2_COLLABORATION_NAME_PREFIX) + || name.startsWith(PLAINTEXT_V2_COLLABORATION_DOTTED_NAME_PREFIX) + || PLAINTEXT_V2_AGENT_MESSAGE_TOOL_NAMES.has(name.split(/__|\./).at(-1)!); +} + +function hasAgentMessageToolAliasCatalogConflict( + body: Record, + catalogs: readonly unknown[][], +): boolean { + const pending = [...catalogs]; + if (Array.isArray(body.input)) { + for (const item of body.input) { + if ( + isPlainObject(item) + && item.type === "tool_search_output" + && Array.isArray(item.tools) + ) { + pending.push(item.tools); + } + } + } + while (pending.length > 0) { + const tools = pending.pop()!; + for (const tool of tools) { + if (!isPlainObject(tool)) continue; + if ( + typeof tool.name === "string" + && hasPrivateToolName(tool.name) + ) { + return true; + } + if (tool.type === "namespace" && Array.isArray(tool.tools)) pending.push(tool.tools); + } + } + return false; +} + +function hasAgentMessageToolAliasReference(value: unknown): boolean { + if (!isPlainObject(value) || !isToolIdentity(value) || typeof value.name !== "string") { + return false; + } + if (hasPrivateToolName(value.name)) return true; + for (const prefix of [COLLABORATION_NAME_PREFIX, COLLABORATION_DOTTED_NAME_PREFIX]) { + if ( + value.name.startsWith(prefix) + && PLAINTEXT_V2_AGENT_MESSAGE_TOOL_NAMES.has(value.name.slice(prefix.length)) + ) { + return true; + } + } + return false; +} + +function hasAgentMessageToolAliasReferenceConflict(body: Record): boolean { + if (hasAgentMessageToolAliasReference(body.tool_choice)) return true; + if ( + isPlainObject(body.tool_choice) + && Array.isArray(body.tool_choice.tools) + && body.tool_choice.tools.some(hasAgentMessageToolAliasReference) + ) { + return true; + } + if (!Array.isArray(body.input)) return false; + return body.input.some(item => ( + isPlainObject(item) + && (item.type === "function_call" || item.type === "custom_tool_call") + && hasAgentMessageToolAliasReference(item) + )); +} + +function hasFlattenedCollaborationDeclarationConflict( + catalogs: readonly unknown[][], + collaborationToolNames: ReadonlySet, +): boolean { + const qualifiedNames = new Set( + [...collaborationToolNames].flatMap(name => [ + `${COLLABORATION_NAME_PREFIX}${name}`, + `${COLLABORATION_DOTTED_NAME_PREFIX}${name}`, + ]), + ); + const pending = catalogs.map(tools => ({ tools, collaborationNamespace: false })); + while (pending.length > 0) { + const { tools, collaborationNamespace } = pending.pop()!; + for (const tool of tools) { + if (!isPlainObject(tool)) continue; + if ( + !collaborationNamespace + && (tool.type === "function" || tool.type === "custom") + && typeof tool.name === "string" + && qualifiedNames.has(tool.name) + ) { + return true; + } + if (tool.type === "namespace" && Array.isArray(tool.tools)) { + pending.push({ + tools: tool.tools, + collaborationNamespace: tool.name === COLLABORATION_NAMESPACE, + }); + } + } + } + return false; +} + +function rewriteToolCatalog(tools: unknown[]): { + tools: unknown[]; + namespaceAliased: boolean; +} { + let namespaceAliased = false; + let changed = false; + const rewritten = tools.map(tool => { + if ( + !isPlainObject(tool) + || tool.type !== "namespace" + || tool.name !== COLLABORATION_NAMESPACE + || !Array.isArray(tool.tools) + ) { + return tool; + } + const childTools = tool.tools.map(child => ( + isPlainObject(child) ? rewriteAgentMessageToolDeclaration(child) : child + )); + namespaceAliased = true; + changed = true; + return { + ...tool, + name: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + tools: childTools, + }; + }); + return { tools: changed ? rewritten : tools, namespaceAliased }; +} + +function aliasCollaborationReference( + value: unknown, + collaborationToolNames: ReadonlySet, +): unknown { + if (!isPlainObject(value)) return value; + const type = value.type; + const canCarryNamespace = isToolIdentity(value); + if (canCarryNamespace && value.namespace !== undefined && value.namespace !== COLLABORATION_NAMESPACE) return value; + let rewritten = value; + if (canCarryNamespace && value.namespace === COLLABORATION_NAMESPACE) { + const name = (type === "function" || type === "function_call") && typeof value.name === "string" + ? PLAINTEXT_V2_AGENT_MESSAGE_TOOL_ALIASES.get(value.name) ?? value.name + : value.name; + rewritten = { ...rewritten, namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, name }; + } + if (type === "namespace" && value.name === COLLABORATION_NAMESPACE) { + rewritten = { ...rewritten, name: PLAINTEXT_V2_COLLABORATION_NAMESPACE }; + } else if ( + canCarryNamespace + && typeof value.name === "string" + && value.name.startsWith(COLLABORATION_NAME_PREFIX) + && collaborationToolNames.has(value.name.slice(COLLABORATION_NAME_PREFIX.length)) + ) { + const childName = value.name.slice(COLLABORATION_NAME_PREFIX.length); + rewritten = { + ...rewritten, + name: `${PLAINTEXT_V2_COLLABORATION_NAME_PREFIX}${ + (type === "function" || type === "function_call") + ? PLAINTEXT_V2_AGENT_MESSAGE_TOOL_ALIASES.get(childName) ?? childName + : childName + }`, + }; + } else if ( + canCarryNamespace + && typeof value.name === "string" + && value.name.startsWith(COLLABORATION_DOTTED_NAME_PREFIX) + && collaborationToolNames.has(value.name.slice(COLLABORATION_DOTTED_NAME_PREFIX.length)) + ) { + const childName = value.name.slice(COLLABORATION_DOTTED_NAME_PREFIX.length); + rewritten = { + ...rewritten, + name: `${PLAINTEXT_V2_COLLABORATION_DOTTED_NAME_PREFIX}${ + (type === "function" || type === "function_call") + ? PLAINTEXT_V2_AGENT_MESSAGE_TOOL_ALIASES.get(childName) ?? childName + : childName + }`, + }; + } + return rewritten; +} + +function aliasCollaborationToolChoice( + toolChoice: unknown, + collaborationToolNames: ReadonlySet, +): unknown { + if (!isPlainObject(toolChoice)) return toolChoice; + let rewritten = aliasCollaborationReference( + toolChoice, + collaborationToolNames, + ) as Record; + if (!Array.isArray(toolChoice.tools)) return rewritten; + let toolsChanged = false; + const tools = toolChoice.tools.map(tool => { + const aliased = aliasCollaborationReference(tool, collaborationToolNames); + toolsChanged ||= aliased !== tool; + return aliased; + }); + if (toolsChanged) rewritten = { ...rewritten, tools }; + return rewritten; +} + +/** + * Prepare v2 collaboration tools for plaintext messages on the canonical ChatGPT wire. + * + * ChatGPT reserves both `collaboration` and the three message-tool names. The request therefore + * uses fixed, request-scoped aliases for both, then restores every identity before Codex sees it. + */ +export function preparePlaintextV2AgentMessages(body: unknown): { + body: unknown; + namespaceAliased: boolean; + toolNames: ReadonlySet; + aliasedAgentMessageToolNames: ReadonlySet; +} { + if (!isPlainObject(body)) { + return { + body, + namespaceAliased: false, + toolNames: new Set(), + aliasedAgentMessageToolNames: new Set(), + }; + } + const catalogs = responseToolCatalogs(body); + const catalogInfo = collaborationCatalogInfo(catalogs); + if ( + !hasPlaintextV2CollaborationCatalog(body) + || hasOptimizedNamespaceConflict(catalogs) + || hasOptimizedReferenceConflict(body) + || hasToolSearchCollaborationConflict(body) + || hasFlattenedCollaborationDeclarationConflict(catalogs, catalogInfo.toolNames) + || hasAgentMessageToolAliasCatalogConflict(body, catalogs) + || hasAgentMessageToolAliasReferenceConflict(body) + ) { + return { + body, + namespaceAliased: false, + toolNames: new Set(), + aliasedAgentMessageToolNames: new Set(), + }; + } + + let namespaceAliased = false; + let tools = body.tools; + if (Array.isArray(body.tools)) { + const rewritten = rewriteToolCatalog(body.tools); + tools = rewritten.tools; + namespaceAliased ||= rewritten.namespaceAliased; + } + + let input = body.input; + if (Array.isArray(body.input)) { + let inputChanged = false; + const rewrittenInput = body.input.map(item => { + if ( + !isPlainObject(item) + || item.type !== "additional_tools" + || !Array.isArray(item.tools) + ) { + return item; + } + const rewritten = rewriteToolCatalog(item.tools); + namespaceAliased ||= rewritten.namespaceAliased; + if (rewritten.tools === item.tools) return item; + inputChanged = true; + return { ...item, tools: rewritten.tools }; + }); + if (inputChanged) input = rewrittenInput; + } + + let toolChoice = body.tool_choice; + if (namespaceAliased) { + toolChoice = aliasCollaborationToolChoice(body.tool_choice, catalogInfo.toolNames); + if (Array.isArray(input)) { + let inputChanged = false; + const aliasedInput = input.map(item => { + if (!isPlainObject(item)) return item; + if (item.type !== "function_call" && item.type !== "custom_tool_call") return item; + const aliased = aliasCollaborationReference(item, catalogInfo.toolNames); + inputChanged ||= aliased !== item; + return aliased; + }); + if (inputChanged) input = aliasedInput; + } + } + + if (!namespaceAliased || (tools === body.tools && input === body.input && toolChoice === body.tool_choice)) { + return { + body, + namespaceAliased: false, + toolNames: new Set(), + aliasedAgentMessageToolNames: new Set(), + }; + } + return { + body: { + ...body, + ...(tools !== body.tools ? { tools } : {}), + ...(input !== body.input ? { input } : {}), + ...(toolChoice !== body.tool_choice ? { tool_choice: toolChoice } : {}), + }, + namespaceAliased, + toolNames: new Set(catalogInfo.toolNames), + aliasedAgentMessageToolNames: new Set(catalogInfo.aliasedAgentMessageToolNames), + }; +} + +const MAX_RESTORED_TOOL_IDENTITIES = 10_000; +export const PLAINTEXT_V2_AGENT_MESSAGE_RESTORE_OVERFLOW_MESSAGE = + "plaintext V2 agent-message response could not be restored within safe identity limits"; + +export class PlaintextV2AgentMessageRestoreOverflowError extends Error { + constructor() { + super(PLAINTEXT_V2_AGENT_MESSAGE_RESTORE_OVERFLOW_MESSAGE); + this.name = "PlaintextV2AgentMessageRestoreOverflowError"; + } +} + +type RestoreOutcome = { + value: unknown; + changed: boolean; + overflow: boolean; +}; + +type RestoreContext = { + toolNames: ReadonlySet; + aliasedAgentMessageToolNames: ReadonlySet; + remainingIdentities: number; +}; + +const unchanged = (value: unknown): RestoreOutcome => ({ value, changed: false, overflow: false }); + +function reserveIdentities(context: RestoreContext, count: number): boolean { + if (count > context.remainingIdentities) return false; + context.remainingIdentities -= count; + return true; +} + +function declaredChildName( + name: unknown, + toolNames: ReadonlySet, + aliasedAgentMessageToolNames: ReadonlySet, + allowAgentMessageAlias: boolean, +): string | undefined { + if (typeof name !== "string") return undefined; + if (allowAgentMessageAlias) { + const restoredName = PLAINTEXT_V2_AGENT_MESSAGE_TOOL_NAMES.get(name); + if (restoredName) { + return aliasedAgentMessageToolNames.has(restoredName) && toolNames.has(restoredName) + ? restoredName + : undefined; + } + } + if (toolNames.has(name)) return name; + for (const prefix of [ + PLAINTEXT_V2_COLLABORATION_NAME_PREFIX, + PLAINTEXT_V2_COLLABORATION_DOTTED_NAME_PREFIX, + ]) { + if (!name.startsWith(prefix)) continue; + const childName = name.slice(prefix.length); + const restoredAlias = allowAgentMessageAlias + ? PLAINTEXT_V2_AGENT_MESSAGE_TOOL_NAMES.get(childName) + : undefined; + if (restoredAlias && !aliasedAgentMessageToolNames.has(restoredAlias)) return undefined; + const restoredChildName = restoredAlias ?? childName; + return toolNames.has(restoredChildName) ? restoredChildName : undefined; + } + return undefined; +} + +function restoreToolIdentity( + value: unknown, + context: RestoreContext, + allowNamespaceDeclaration = false, + namespaceMember = false, +): RestoreOutcome { + if (!isPlainObject(value)) return unchanged(value); + if (!reserveIdentities(context, 1)) return { ...unchanged(value), overflow: true }; + + if ( + allowNamespaceDeclaration + && value.type === "namespace" + && value.name === PLAINTEXT_V2_COLLABORATION_NAMESPACE + ) { + if (value.tools !== undefined && !Array.isArray(value.tools)) return { ...unchanged(value), overflow: true }; + const children = restoreIdentityList(value.tools, context, false, true); + if (children.overflow) return { ...unchanged(value), overflow: true }; + return { + value: { + ...value, + name: COLLABORATION_NAMESPACE, + ...(children.changed ? { tools: children.value } : {}), + }, + changed: true, + overflow: false, + }; + } + + const identityType = value.type; + if ( + identityType !== "function" + && identityType !== "custom" + && identityType !== "function_call" + && identityType !== "custom_tool_call" + && identityType !== "response.function_call_arguments.done" + ) { + return unchanged(value); + } + + if (value.namespace !== undefined && value.namespace !== null && typeof value.namespace !== "string") { + return { ...unchanged(value), overflow: true }; + } + const allowAgentMessageAlias = ( + identityType === "function" + || identityType === "function_call" + || identityType === "response.function_call_arguments.done" + ) && ( + value.namespace === undefined + || value.namespace === null + || value.namespace === PLAINTEXT_V2_COLLABORATION_NAMESPACE + ); + if (value.namespace !== undefined && value.namespace !== null && value.namespace !== PLAINTEXT_V2_COLLABORATION_NAMESPACE) { + return unchanged(value); + } + const childName = declaredChildName( + value.name, + context.toolNames, + context.aliasedAgentMessageToolNames, + allowAgentMessageAlias, + ); + if (!childName) { + const privateIdentity = value.namespace === PLAINTEXT_V2_COLLABORATION_NAMESPACE + || (typeof value.name === "string" && hasPrivateToolName(value.name)); + return { ...unchanged(value), overflow: privateIdentity }; + } + + const privateIdentity = value.namespace === PLAINTEXT_V2_COLLABORATION_NAMESPACE + || (typeof value.name === "string" && hasPrivateToolName(value.name)); + if (!privateIdentity) return unchanged(value); + // Codex dispatches by the namespace/name pair; qualified names are literal + // names there. Only namespace member declarations inherit their container. + return { + value: { + ...value, + name: childName, + ...(!namespaceMember || value.namespace !== undefined ? { namespace: COLLABORATION_NAMESPACE } : {}), + }, + changed: true, + overflow: false, + }; +} + +function restoreIdentityList( + values: unknown, + context: RestoreContext, + allowNamespaceDeclaration: boolean, + namespaceMember = false, +): RestoreOutcome { + if (!Array.isArray(values)) return unchanged(values); + if (values.length > context.remainingIdentities) { + return { ...unchanged(values), overflow: true }; + } + let restored: unknown[] | undefined; + for (let index = 0; index < values.length; index += 1) { + const result = restoreToolIdentity(values[index], context, allowNamespaceDeclaration, namespaceMember); + if (result.overflow) return { ...unchanged(values), overflow: true }; + if (!result.changed) continue; + restored ??= values.slice(); + restored[index] = result.value; + } + return restored + ? { value: restored, changed: true, overflow: false } + : unchanged(values); +} + +function restoreToolChoice(value: unknown, context: RestoreContext): RestoreOutcome { + if (isPlainObject(value) && value.tools !== undefined && !Array.isArray(value.tools)) { + return { ...unchanged(value), overflow: true }; + } + const direct = restoreToolIdentity(value, context, true); + if (direct.overflow || !isPlainObject(value) || !Array.isArray(value.tools)) return direct; + const tools = restoreIdentityList(value.tools, context, true); + if (tools.overflow) return { ...unchanged(value), overflow: true }; + if (!tools.changed) return direct; + const base = direct.value as Record; + return { value: { ...base, tools: tools.value }, changed: true, overflow: false }; +} + +function restoreResponseSnapshot(value: unknown, context: RestoreContext): RestoreOutcome { + if (!isPlainObject(value)) return unchanged(value); + if ((value.output !== undefined && !Array.isArray(value.output)) + || (value.tools !== undefined && !Array.isArray(value.tools))) { + return { ...unchanged(value), overflow: true }; + } + const output = restoreIdentityList(value.output, context, false); + if (output.overflow) return { ...unchanged(value), overflow: true }; + const tools = restoreIdentityList(value.tools, context, true); + if (tools.overflow) return { ...unchanged(value), overflow: true }; + const toolChoice = restoreToolChoice(value.tool_choice, context); + if (toolChoice.overflow) return { ...unchanged(value), overflow: true }; + if (!output.changed && !tools.changed && !toolChoice.changed) return unchanged(value); + return { + value: { + ...value, + ...(output.changed ? { output: output.value } : {}), + ...(tools.changed ? { tools: tools.value } : {}), + ...(toolChoice.changed ? { tool_choice: toolChoice.value } : {}), + }, + changed: true, + overflow: false, + }; +} + +/** + * Restore request-scoped collaboration aliases only at documented Responses identity positions. + * Tool arguments, tool results, and extension metadata are deliberately opaque. + */ +export function restorePlaintextV2AgentMessageCalls( + value: unknown, + toolNames: ReadonlySet, + aliasedAgentMessageToolNames: ReadonlySet = toolNames, +): { value: unknown; changed: boolean; overflowed: boolean } { + if (toolNames.size === 0 || !isPlainObject(value)) { + return { value, changed: false, overflowed: false }; + } + const context: RestoreContext = { + toolNames, + aliasedAgentMessageToolNames, + remainingIdentities: MAX_RESTORED_TOOL_IDENTITIES, + }; + + const rootIdentity = restoreToolIdentity(value, context); + if (rootIdentity.overflow) return { value, changed: false, overflowed: true }; + const root = rootIdentity.value as Record; + const item = restoreToolIdentity(root.item, context); + if (item.overflow) return { value, changed: false, overflowed: true }; + const response = restoreResponseSnapshot(root.response, context); + if (response.overflow) return { value, changed: false, overflowed: true }; + const snapshot = restoreResponseSnapshot(root, context); + if (snapshot.overflow) return { value, changed: false, overflowed: true }; + + let restored = snapshot.value as Record; + let changed = rootIdentity.changed || snapshot.changed; + if (item.changed) { + restored = { ...restored, item: item.value }; + changed = true; + } + if (response.changed) { + restored = { ...restored, response: response.value }; + changed = true; + } + return changed + ? { value: restored, changed: true, overflowed: false } + : { value, changed: false, overflowed: false }; +} + +export function restorePlaintextV2AgentMessageCallsInJsonResult( + payload: string, + toolNames: ReadonlySet, + aliasedAgentMessageToolNames: ReadonlySet = toolNames, +): { value: string; changed: boolean; overflowed: boolean } { + if (toolNames.size === 0 || payload === "[DONE]") { + return { value: payload, changed: false, overflowed: false }; + } + let value: unknown; + try { + value = JSON.parse(payload); + } catch { + return { value: payload, changed: false, overflowed: true }; + } + if (!isPlainObject(value)) return { value: payload, changed: false, overflowed: true }; + const restored = restorePlaintextV2AgentMessageCalls( + value, + toolNames, + aliasedAgentMessageToolNames, + ); + if (restored.overflowed) return { value: payload, changed: false, overflowed: true }; + return restored.changed + ? { value: JSON.stringify(restored.value), changed: true, overflowed: false } + : { value: payload, changed: false, overflowed: false }; +} + +export function restorePlaintextV2AgentMessageCallsInJson( + payload: string, + toolNames: ReadonlySet, + aliasedAgentMessageToolNames: ReadonlySet = toolNames, +): string { + const restored = restorePlaintextV2AgentMessageCallsInJsonResult( + payload, + toolNames, + aliasedAgentMessageToolNames, + ); + if (restored.overflowed) throw new PlaintextV2AgentMessageRestoreOverflowError(); + return restored.value; +} + +export function createPlaintextV2AgentMessageCallRestoreRewrite( + toolNames: ReadonlySet, + aliasedAgentMessageToolNames: ReadonlySet = toolNames, +): (payload: string) => string { + type Binding = { namespace: string; name: string; keys: Set }; + const bindings = new Map(); + let refused = false; + return payload => { + if (refused) throw new PlaintextV2AgentMessageRestoreOverflowError(); + try { + const restored = restorePlaintextV2AgentMessageCallsInJson(payload, toolNames, aliasedAgentMessageToolNames); + if (payload === "[DONE]" || toolNames.size === 0) return restored; + const value = JSON.parse(restored) as Record; + const bind = (item: unknown, outputIndex?: unknown): void => { + if (!isPlainObject(item) || typeof item.name !== "string") return; + if (item.type !== "function_call" && item.type !== "response.function_call_arguments.done") return; + let namespace = typeof item.namespace === "string" ? item.namespace : ""; + let name = item.name; + for (const separator of ["__", "."]) { + const prefix = `${COLLABORATION_NAMESPACE}${separator}`; + if (name.startsWith(prefix) && (!namespace || namespace === COLLABORATION_NAMESPACE)) { + namespace = COLLABORATION_NAMESPACE; + name = name.slice(prefix.length); + } + } + const keys = [ + typeof item.id === "string" ? `id:${item.id}` : undefined, + typeof item.item_id === "string" ? `id:${item.item_id}` : undefined, + typeof item.call_id === "string" ? `call:${item.call_id}` : undefined, + typeof outputIndex === "number" ? `index:${outputIndex}` : undefined, + ].filter((key): key is string => key !== undefined); + const groups = [...new Set(keys.flatMap(key => { + const prior = bindings.get(key); + return prior ? [prior] : []; + }))]; + for (const group of groups) { + if (group.name !== name || (group.namespace && namespace && group.namespace !== namespace)) { + throw new PlaintextV2AgentMessageRestoreOverflowError(); + } + namespace ||= group.namespace; + } + // All coordinates for a call share the same refined identity, including + // coordinates omitted by this particular sparse event. Merge smaller groups + // into the largest to bound repeated cross-coordinate refinement work. + groups.sort((left, right) => right.keys.size - left.keys.size); + const binding: Binding = groups[0] ?? { namespace, name, keys: new Set() }; + binding.namespace = namespace; + for (const group of groups.slice(1)) { + for (const key of group.keys) { + binding.keys.add(key); + bindings.set(key, binding); + } + } + for (const key of keys) { + if (!bindings.has(key) && bindings.size >= MAX_RESTORED_TOOL_IDENTITIES) throw new PlaintextV2AgentMessageRestoreOverflowError(); + binding.keys.add(key); + bindings.set(key, binding); + } + }; + bind(value, value.output_index); + bind(value.item, value.output_index); + const response = isPlainObject(value.response) ? value.response : value; + if (Array.isArray(response.output)) response.output.forEach((item, index) => bind(item, index)); + return restored; + } catch (error) { + refused = true; + throw error; + } + }; +} diff --git a/src/router.ts b/src/router.ts index bd8e9dc690..70e427b74b 100644 --- a/src/router.ts +++ b/src/router.ts @@ -413,6 +413,13 @@ export function routedProviderConfig(providerName: string, provider: OcxProvider ...(provider.preserveResponsesReasoningContent === undefined && registryEntry.preserveResponsesReasoningContent !== undefined ? { preserveResponsesReasoningContent: registryEntry.preserveResponsesReasoningContent } : {}), + // The request path resolves through routedProviderConfig() and never calls + // enrichProviderFromRegistry(), so a saved provider row written before the + // registry learned this flag must be backfilled here or route.provider never + // carries it and the showThinkingSummary opt-in stays dead. + ...(provider.showThinkingSummary === undefined && registryEntry.showThinkingSummary !== undefined + ? { showThinkingSummary: registryEntry.showThinkingSummary } + : {}), // Registry-only client-facing repair policy (#938): fill only when the // saved provider has no explicit policy; clone so runtime never aliases // the registry constant. diff --git a/src/server/auth-cors.ts b/src/server/auth-cors.ts index 37ce4b1127..2f7cd66111 100644 --- a/src/server/auth-cors.ts +++ b/src/server/auth-cors.ts @@ -943,6 +943,7 @@ const PROVIDER_CONFIG_FIELD_POLICY = { autoToolChoiceOnlyModels: "editor", preserveReasoningContentModels: "editor", requiresReasoningPlaceholderModels: "editor", + showThinkingSummary: "editor", retryOn429: "editor", transientRetryOn5xx: "editor", reasoningSplitModels: "editor", diff --git a/src/server/index.ts b/src/server/index.ts index 5460f443e9..51d3ab9393 100644 --- a/src/server/index.ts +++ b/src/server/index.ts @@ -150,6 +150,7 @@ import { } from "./relay"; export { consumeForInspection, + codexSafetyBufferingFilterOptions, relaySseWithFailedTail, relaySseWithHeartbeat, relayWithAbort, @@ -661,6 +662,13 @@ export function warnAgentTaskRecoveryStartup(config: { console.warn(" Recovered plaintext assignment data is retained only in a bounded, process-local in-memory cache; exact fidelity is not guaranteed and the path depends on undocumented backend behavior."); } +export function warnPlaintextV2AgentMessagesStartup(config: { plaintextV2AgentMessages?: boolean }): void { + if (config.plaintextV2AgentMessages !== true) return; + console.warn("⚠️ Experimental plaintext V2 agent messages are enabled."); + console.warn(" Eligible ChatGPT collaboration calls may carry plaintext message arguments. HTTPS remains encrypted, but task text may be retained in Codex history, selected providers, and local response/debug state."); + console.warn(" This depends on undocumented ChatGPT and Codex behavior; it does not decrypt existing tasks."); +} + export function startServer(port?: number, deps: StartServerDeps = {}): Server { const localAttestationSecret = deps.localAttestationSecret ?? createLocalAttestationSecret(); // Captured before loadConfig() starts the optional ACL flight so stop() drains the same dir @@ -673,6 +681,7 @@ export function startServer(port?: number, deps: StartServerDeps = {}): Server number; }; @@ -114,7 +117,7 @@ export function relaySseEagerBounded( const terminalEncoder = new TextEncoder(); const adapterEofFrame = adapterEofIncompleteFrame(terminalEncoder); const terminalSentinel = doneFrame(terminalEncoder); - const terminalBoundary = createSseTerminalOutputBoundary(); + const terminalBoundary = createSseTerminalOutputBoundary(opts?.terminalBoundary); const activeRewrite: SseBlockRewrite | undefined = hooks.rewriteBlocks ?? (hooks.rewritePayload ? payloadRewriteAsBlockRewrite(hooks.rewritePayload) : undefined); const encodeFailedTail = (error: unknown): Uint8Array | null => { diff --git a/src/server/relay.ts b/src/server/relay.ts index a483d88f20..f480ace68a 100644 --- a/src/server/relay.ts +++ b/src/server/relay.ts @@ -183,7 +183,10 @@ export type SseTerminalOutputBoundary = { * terminal, and drops every later block/byte. A premature [DONE] is held until * a terminal arrives so clean EOF can synthesize one terminal and one sentinel. */ -export function createSseTerminalOutputBoundary(): SseTerminalOutputBoundary { +export function createSseTerminalOutputBoundary( + options?: CodexSafetyBufferingFilterOptions, +): SseTerminalOutputBoundary { + const dropSafetyBuffering = options?.dropCodexSafetyBuffering === true; const decoder = new TextDecoder(); const encoder = new TextEncoder(); const framer = new BoundedSseFrameBuffer(MAX_INSPECTION_SSE_FRAME_BYTES); @@ -207,15 +210,22 @@ export function createSseTerminalOutputBoundary(): SseTerminalOutputBoundary { // behind EOF, so its log context cannot determine the outgoing terminal. const message = boundedBareUpstreamErrorMessage(parsed); if (message !== undefined) upstreamError = message; + const safetyBuffering = dropSafetyBuffering && parsed !== undefined + ? codexSafetyBufferingBlockAction(parsed) : "keep"; + if (safetyBuffering === "drop") continue; const policyError = parsed !== undefined && isPolicyRewriteType(parsed) ? cyberPolicyTerminalError(parsed) : undefined; - const outboundBlock = policyError - ? encoder.encode(rewritePolicyTerminalBlock( - decoder.decode(frame.block), - policyFailurePayload(policyError, parsed), - )) + const policyPayload = policyError ? policyFailurePayload(policyError, parsed) : undefined; + let outboundBlock = policyPayload !== undefined + ? encoder.encode(rewritePolicyTerminalBlock(decoder.decode(frame.block), policyPayload)) : frame.block; + if (safetyBuffering === "strip") { + outboundBlock = encoder.encode(stripCodexSafetyBufferingField( + decoder.decode(outboundBlock), + policyPayload !== undefined ? parseSsePayload(policyPayload) : parsed, + )); + } if (isDone) { done = true; if (responsesTerminal) { @@ -287,11 +297,11 @@ export function relaySseWithFailedTail( body: ReadableStream, upstream: AbortController, onClientGone?: (reason?: unknown) => void, - opts?: { upstreamError?: string }, + opts?: { upstreamError?: string; terminalBoundary?: CodexSafetyBufferingFilterOptions }, ): ReadableStream { const reader = body.getReader(); const encoder = new TextEncoder(); - const terminalBoundary = createSseTerminalOutputBoundary(); + const terminalBoundary = createSseTerminalOutputBoundary(opts?.terminalBoundary); let closed = false; const relayChunk = ( controller: ReadableStreamDefaultController, @@ -468,6 +478,29 @@ function isPolicyRewriteType(parsed: unknown): boolean { return type === "response.failed" || type === "response.incomplete" || type === "error"; } +/** + * Codex emits its safety-buffering hint in the SSE body as well as in headers: + * a `response.metadata` event whose `metadata.type` is `safety_buffering`, or a + * `safety_buffering` field on another event. The metadata event is dropped whole; + * the field is stripped so the carrying event is otherwise relayed unchanged. + */ +function codexSafetyBufferingBlockAction(parsed: unknown): "keep" | "drop" | "strip" { + const root = asJsonRecord(parsed); + if (!root) return "keep"; + if (root.type === "response.metadata") { + const metadata = asJsonRecord(root.metadata); + if (metadata?.type === "safety_buffering") return "drop"; + } + return Object.hasOwn(root, "safety_buffering") ? "strip" : "keep"; +} + +function stripCodexSafetyBufferingField(block: string, parsed: unknown): string { + const root = asJsonRecord(parsed); + if (!root) return block; + const { safety_buffering: _safetyBuffering, ...rest } = root; + return replaceSseDataPayload(block, JSON.stringify(rest)); +} + function rewritePolicyTerminalBlock(block: string, payload: string): string { const newline = block.includes("\r\n") ? "\r\n" : "\n"; const rewritten = replaceSseDataPayload(block, payload); @@ -1461,7 +1494,31 @@ export function consumeForResponseLogMetadata( * body makes the caller (Codex) double-decode / truncate → "stream error" on every gpt passthrough. * Drop encoding + hop-by-hop headers; relay everything else (content-type, etc.) verbatim. */ -export function sanitizePassthroughHeaders(upstream: Headers): Headers { +export const CODEX_SAFETY_BUFFERING_HEADERS = [ + "x-codex-safety-buffering-enabled", + "x-codex-safety-buffering-faster-model", +] as const; + +const CODEX_SAFETY_BUFFERING_HEADER_SET: ReadonlySet = new Set(CODEX_SAFETY_BUFFERING_HEADERS); + +export interface CodexSafetyBufferingFilterOptions { + /** + * Drop Codex safety-buffering hints: the `x-codex-safety-buffering-*` response + * headers and the `safety_buffering` SSE metadata event / field. Absent and + * `false` relay everything unchanged. + */ + dropCodexSafetyBuffering?: boolean; +} + +/** Resolve the passthrough header policy from the loaded config (absent means "forward everything"). */ +export function codexSafetyBufferingFilterOptions( + config: { dropCodexSafetyBuffering?: boolean }, +): CodexSafetyBufferingFilterOptions { + return { dropCodexSafetyBuffering: config.dropCodexSafetyBuffering === true }; +} + +export function sanitizePassthroughHeaders(upstream: Headers, options?: CodexSafetyBufferingFilterOptions): Headers { + const dropSafetyBuffering = options?.dropCodexSafetyBuffering === true; const DROP = new Set([ "content-encoding", "content-length", @@ -1478,7 +1535,10 @@ export function sanitizePassthroughHeaders(upstream: Headers): Headers { ]); const out = new Headers(); upstream.forEach((value, key) => { - if (!DROP.has(key.toLowerCase())) out.set(key, value); + const lower = key.toLowerCase(); + if (DROP.has(lower)) return; + if (dropSafetyBuffering && CODEX_SAFETY_BUFFERING_HEADER_SET.has(lower)) return; + out.set(key, value); }); return out; } diff --git a/src/server/responses-reasoning-summary-rewrite.ts b/src/server/responses-reasoning-summary-rewrite.ts deleted file mode 100644 index 55c6d8ae7b..0000000000 --- a/src/server/responses-reasoning-summary-rewrite.ts +++ /dev/null @@ -1,178 +0,0 @@ -import type { SsePayloadRewrite } from "./sse-payload-rewrite"; - -/** - * Route content-channel reasoning from native-Responses upstreams through the - * expandable summary channel (issue #45). - * - * Codex renders the expandable reasoning trace from the Responses reasoning - * item's `summary[]` channel. DeepSeek's native `/responses` endpoint emits - * raw thinking on the content channel instead (`response.reasoning_text.delta` - * plus items with `content: [{type: "reasoning_text", text}]` and an empty - * `summary`), so routed DeepSeek turns showed the "Worked for Xs" timer with - * nothing to expand. Native OpenAI upstreams already emit summary-channel - * events; this rewrite is a no-op for them (no reasoning_text events to - * rewrite) and only engages when the upstream produces content-channel - * reasoning. - * - * Replay compatibility: Codex echoes the reasoning item it received back into - * the next request's input. DeepSeek's Responses API accepts summary-shaped - * reasoning input items (verified live), so the rewrite round-trips. - */ - -function isPlainObject(value: unknown): value is Record { - return !!value && typeof value === "object" && !Array.isArray(value); -} - -function reasoningTextOf(item: Record): string { - if (!Array.isArray(item.content)) return ""; - return item.content - .filter((part): part is Record => isPlainObject(part) && part.type === "reasoning_text") - .map(part => (typeof part.text === "string" ? part.text : "")) - .join(""); -} - -/** Move a reasoning item's content channel into the summary channel. */ -function reasoningItemToSummaryShape(item: Record): Record { - if (item.type !== "reasoning") return item; - // `encrypted_content` is opaque, state-bearing provider data, so the entire item must retain its - // upstream shape unless that backend has an explicit replay contract permitting a rewrite. This - // defensively protects content-channel backends that do issue blobs when the client replays the - // stored item. The delta rewrite can still provide the expandable trace for the live turn. - // DeepSeek — the provider this rewrite was verified against — is `statelessResponses` and issues - // no blob, so it is unaffected. - if (typeof item.encrypted_content === "string" && item.encrypted_content.length > 0) return item; - const text = reasoningTextOf(item); - // Items that already use the summary channel (or carry no content text at - // all) are left untouched: rewriting them could clear a valid summary. - if (text.length === 0) return item; - const next: Record = { ...item }; - delete next.content; - next.summary = [{ type: "summary_text", text }]; - return next; -} - -/** - * Rewrite one parsed SSE payload in place of the content channel, or return - * `null` when nothing changed (caller keeps the original payload). - */ -function rewritePayload(payload: Record): Record | null { - switch (payload.type) { - case "response.reasoning_text.delta": { - const next: Record = { - type: "response.reasoning_summary_text.delta", - item_id: payload.item_id, - output_index: payload.output_index, - summary_index: 0, - delta: payload.delta, - }; - if (payload.sequence_number !== undefined) next.sequence_number = payload.sequence_number; - return next; - } - case "response.reasoning_text.done": { - const next: Record = { - type: "response.reasoning_summary_text.done", - item_id: payload.item_id, - output_index: payload.output_index, - summary_index: 0, - text: payload.text, - }; - if (payload.sequence_number !== undefined) next.sequence_number = payload.sequence_number; - return next; - } - default: { - let changed = false; - const next: Record = { ...payload }; - if (isPlainObject(next.item) && next.item.type === "reasoning") { - const rewritten = reasoningItemToSummaryShape(next.item); - if (rewritten !== next.item) { - next.item = rewritten; - changed = true; - } - } - // SSE event shape: {type: "response.completed", response: {output}}. - const response = isPlainObject(next.response) ? { ...next.response } : null; - if (response && Array.isArray(response.output)) { - const output = response.output.map(item => { - if (!isPlainObject(item) || item.type !== "reasoning") return item; - const rewritten = reasoningItemToSummaryShape(item); - if (rewritten !== item) changed = true; - return rewritten; - }); - if (changed) { - response.output = output; - next.response = response; - } - } - // Bare response document shape (non-streaming passthrough): - // {object: "response", output: [...]}. - if (Array.isArray(next.output)) { - const output = next.output.map(item => { - if (!isPlainObject(item) || item.type !== "reasoning") return item; - const rewritten = reasoningItemToSummaryShape(item); - if (rewritten !== item) changed = true; - return rewritten; - }); - if (changed) next.output = output; - } - return changed ? next : null; - } - } -} - -/** Payload rewrite for passthrough relays whose upstream emits content-channel reasoning. */ -export function createReasoningSummaryChannelPayloadRewrite(): SsePayloadRewrite { - return (payload: string): string => { - let parsed: unknown; - try { - parsed = JSON.parse(payload); - } catch { - return payload; - } - if (!isPlainObject(parsed)) return payload; - const rewritten = rewritePayload(parsed); - return rewritten !== null ? JSON.stringify(rewritten) : payload; - }; -} - -/** - * Object-level variant for the non-streaming passthrough: the bounded-JSON - * relay bypasses the SSE payload rewrite, so reasoning items inside a full - * Responses JSON document need the same normalization before plain JSON - * serialization or forced JSON-to-SSE reframing. Returns the same reference - * when nothing changed. - */ -export function rewriteReasoningSummaryInJson(value: unknown): unknown { - if (!isPlainObject(value)) return value; - const rewritten = rewritePayload(value); - return rewritten !== null ? rewritten : value; -} - -/** String-level variant of {@link rewriteReasoningSummaryInJson}. */ -export function rewriteReasoningSummaryInJsonString(json: string): string { - let parsed: unknown; - try { - parsed = JSON.parse(json); - } catch { - return json; - } - const rewritten = rewriteReasoningSummaryInJson(parsed); - return rewritten === parsed ? json : JSON.stringify(rewritten); -} - -/** - * True when a routed native-Responses provider emits content-channel reasoning - * (raw `reasoning_text`) instead of the summary channel. DeepSeek's - * `/responses` endpoint is the current example: it ships raw thinking with an - * empty `summary` and keeps `preserveReasoningContentModels` so multi-turn - * replays round-trip. - */ -export function routeUsesContentChannelReasoning( - provider: { statelessResponses?: boolean; preserveReasoningContentModels?: string[] }, - modelId: string, -): boolean { - if (provider.statelessResponses === true) return true; - const preserved = provider.preserveReasoningContentModels; - const normalizedModelId = modelId.toLowerCase(); - return Array.isArray(preserved) - && preserved.some(id => id.toLowerCase() === normalizedModelId); -} diff --git a/src/server/responses/core.ts b/src/server/responses/core.ts index 3952db7afc..cb201153a1 100644 --- a/src/server/responses/core.ts +++ b/src/server/responses/core.ts @@ -104,7 +104,10 @@ import { } from "../../lib/errors"; import { injectionDebugLog } from "../../lib/injection-debug-log"; import { resolveClientRetryAfter } from "../../lib/retry-after"; -import { enrichOpenCodeZenUpstreamMessage } from "../../providers/opencode-zen-rate-limit"; +import { + enrichOpenCodeZenUpstreamMessage, + isTransientConsoleGoUploadRejection, +} from "../../providers/opencode-zen-rate-limit"; import { CODE_MODE_EXEC_TOOL_NAME, modelInList, namespacedToolName } from "../../types"; import type { AdapterEvent, @@ -216,6 +219,7 @@ import { isNonReplayableResponse, isTransientUpstreamStatus, prepareSameTarget429Wait, + sleepWithAbort, } from "../../lib/upstream-retry"; import { ForwardAdmissionCredentialError, @@ -346,6 +350,7 @@ import { markEagerRelaySseResponse, markNativePassthroughSseResponse, relaySseWithFailedTail, + codexSafetyBufferingFilterOptions, relayWithAbort, sanitizePassthroughHeaders, } from "../relay"; @@ -369,12 +374,6 @@ import { hasResponsesItemIdRepair, repairResponsesJsonItemIds, } from "../responses-item-id-repair"; -import { - createReasoningSummaryChannelPayloadRewrite, - rewriteReasoningSummaryInJson, - rewriteReasoningSummaryInJsonString, - routeUsesContentChannelReasoning, -} from "../responses-reasoning-summary-rewrite"; import { createImageGenCallRestoreRewrite, imageGenToolCallAliases, @@ -431,6 +430,13 @@ import { restoreRoutedNamespaceCallsInJson, type RoutedNamespaceToolAliases, } from "../../responses/namespace-tool-compat"; +import { + createPlaintextV2AgentMessageCallRestoreRewrite, + PLAINTEXT_V2_AGENT_MESSAGE_RESTORE_OVERFLOW_MESSAGE, + restorePlaintextV2AgentMessageCalls, + restorePlaintextV2AgentMessageCallsInJsonResult, + shouldPreparePlaintextV2AgentMessages, +} from "../../responses/plaintext-v2-agent-messages"; import { createMuseToolNameRestoreRewrite, restoreMuseToolNames, @@ -859,6 +865,31 @@ async function opaqueBlobRejectionBodyForRecovery( } } +/** + * Backoff for the single exact-request replay after a canonical Console upload rejection. + */ +const CONSOLE_GO_UPLOAD_RETRY_DELAY_MS = 800; + +/** + * Peek the upstream error body for the Console Go transient-400 recovery. Only a complete, + * display-safe body may drive a retry decision (same contract as + * opaqueBlobRejectionBodyForRecovery), and reading a clone leaves the original response intact + * for the caller's own error surface when no retry is taken. + */ +async function consoleGoUploadRejectionBody( + response: Response, + alreadyAttempted: boolean, + signal: AbortSignal, +): Promise { + if (isNonReplayableResponse(response) || response.status !== 400 || alreadyAttempted) return undefined; + try { + const body = await readBoundedResponseBody(response.clone(), { signal }); + return body.displaySafe && !body.truncated ? body.text : undefined; + } catch { + return undefined; + } +} + /** * Materialize an upstream error body only when the bounded reader observed a complete, * display-safe payload. Partial timeout and over-limit prefixes are attacker-controlled, @@ -2578,6 +2609,19 @@ async function applyFinalRouteRequestNormalization(args: { route.provider = resolveOpenCodeGoTransport(route.provider, args.claudeGoAffinity ? args.claudeGoAffinity.sessionLane : getOrAllocateRequestSessionLane(req)); route.provider = resolveWireProtocolOverride(route.providerName, route.modelId, route.provider, inboundWire); + parsed._plaintextV2AgentMessages = shouldPreparePlaintextV2AgentMessages({ + enabled: config.plaintextV2AgentMessages === true, + inboundWire, + canonicalChatGpt: isCanonicalOpenAiForwardProvider(route.provider), + requestBody: parsed._rawBody, + }); + // Recompute from the original wire preference on every route, including fallback. + // A provider default never converts raw reasoning into a summary. + if (inboundWire === "responses" && parsed._rawBody) { + const summary = (parsed._rawBody as { reasoning?: { summary?: unknown } }).reasoning?.summary; + parsed.options.hideThinkingSummary = summary === "none" + || (!summary && route.provider.showThinkingSummary !== true); + } if (preserveAnthropicResponseModel) parsed._responseModelId = responseModelId; logCtx.model = route.modelId; logCtx.provider = route.providerName; @@ -2933,6 +2977,7 @@ export async function handleComboResponses( pick.target, comboDefaultEffort(config, comboId), supportedLadderFor({ provider: targetRoute.provider, modelId: targetRoute.modelId }), + combo.reasoningEffortMode, ); const childHeaders = buildComboChildHeaders(req.headers); const childRequest = new Request(req.url, { @@ -4810,9 +4855,13 @@ async function handleResponsesInner( } let routedNamespaceToolAliases: RoutedNamespaceToolAliases = new Map(); + let plaintextV2AgentMessageToolNames: ReadonlySet = new Set(); + let plaintextV2AgentMessageAliasedToolNames: ReadonlySet = new Set(); let routedMuseToolNameAliases: MuseToolNameAliases = new Map(); - const refreshRoutedNamespaceToolAliases = (builtRequest: AdapterRequest): void => { + const refreshRequestToolAliases = (builtRequest: AdapterRequest): void => { routedNamespaceToolAliases = builtRequest.convertedRoutedNamespaceToolAliases ?? new Map(); + plaintextV2AgentMessageToolNames = builtRequest.plaintextV2AgentMessageToolNames ?? new Set(); + plaintextV2AgentMessageAliasedToolNames = builtRequest.plaintextV2AgentMessageAliasedToolNames ?? new Set(); routedMuseToolNameAliases = builtRequest.convertedMuseToolNameAliases ?? new Map(); }; @@ -4820,6 +4869,9 @@ async function handleResponsesInner( let hostAdmissionLease = pendingHostAdmissionLease; pendingHostAdmissionLease = null; try { + const codexSafetyBufferingOptions = isCanonicalOpenAiForwardProvider(route.provider) + ? codexSafetyBufferingFilterOptions(config) + : undefined; const imageGenCallAliases = route.provider.authMode === "forward" ? new Map() : imageGenToolCallAliases(toolBridgeMaps.toolNsMap, parsed._rawBody, translatorBudget); @@ -4910,7 +4962,7 @@ async function handleResponsesInner( // would incorrectly disable restoration for the exact ambiguous-name case the alias fixes. routedToolSearchNames.add(name); } - refreshRoutedNamespaceToolAliases(request); + refreshRequestToolAliases(request); // #1700: the bridged paths refuse a call to a tool the request never declared // (`declaredToolNames`, src/bridge.ts). The passthrough had no equivalent, so a routed // provider's top-level `apply_patch` — which under Codex code mode exists only as a nested @@ -5098,7 +5150,7 @@ async function handleResponsesInner( } // The snapshot callback opts the inspector into output reconstruction. Compaction // has no continuation cache, so use the parsed terminal here without adding retention. - if (!rememberPassthroughResponse && payload && typeof payload === "object" + if (plaintextV2AgentMessageToolNames.size === 0 && !rememberPassthroughResponse && payload && typeof payload === "object" && "type" in payload && payload.type === "response.completed" && "response" in payload && payload.response && typeof payload.response === "object" && !Array.isArray(payload.response)) { @@ -5120,14 +5172,16 @@ async function handleResponsesInner( routedCustomToolRepairNames, declaredWireToolNames, ).value; - const restoredResponse = (functionRepairSchemas.size > 0 + const normalizedResponse = (functionRepairSchemas.size > 0 ? JSON.parse(normalizeFunctionCompletionJson(JSON.stringify(restored))) : restored) as { id?: unknown; output?: unknown; status?: unknown }; + const plaintextRestore = restorePlaintextV2AgentMessageCalls( + normalizedResponse, plaintextV2AgentMessageToolNames, plaintextV2AgentMessageAliasedToolNames, + ); + if (plaintextRestore.overflowed) return; + const restoredResponse = plaintextRestore.value as typeof normalizedResponse; // Replay overlap compares the items the client echoes, including visible reasoning shape. - const replayResponse = parsed.options.hideThinkingSummary !== true - && routeUsesContentChannelReasoning(route.provider, route.modelId) - ? rewriteReasoningSummaryInJson(restoredResponse) as typeof restoredResponse - : restoredResponse; + const replayResponse = restoredResponse; if ( undeclaredToolGuardActive && undeclaredToolCallNameInResponse( @@ -5351,6 +5405,9 @@ async function handleResponsesInner( const opaqueBlobRecoveryGuard: OpaqueBlobRecoveryGuard = { attempted: false }; let oauth401ReplayAttempted = false; let codex401ReplayKind: "main" | "stored" | null = null; + // Console Go answers a transient 400 "Invalid upload request." for bodies it accepts + // moments later; at most one byte-identical replay is allowed per request. + const consoleGoUploadRetryGuard: { attempted: boolean } = { attempted: false }; const rateLimitPolicy = rateLimitRetryPolicyFor(route.provider); let rateLimitRetries = 0; const rebuildAndRefetch = async ( @@ -5365,11 +5422,13 @@ async function handleResponsesInner( return { failed: formatErrorResponse(502, "upstream_error", "Recovery changed the provider wire unexpectedly") }; } try { - request = await retryAdapter.buildRequest(parsed, { - headers: selectedForwardHeaders, - translatorBudget, - }); - refreshRoutedNamespaceToolAliases(request); + if (recovery !== "console-go-upload-retry") { + request = await retryAdapter.buildRequest(parsed, { + headers: selectedForwardHeaders, + translatorBudget, + }); + } + refreshRequestToolAliases(request); recordAdapterReasoning(logCtx, request); recordAdapterTier(logCtx, request); } catch (err) { @@ -5486,7 +5545,7 @@ async function handleResponsesInner( headers: selectedForwardHeaders, translatorBudget, }); - refreshRoutedNamespaceToolAliases(request); + refreshRequestToolAliases(request); recordAdapterReasoning(logCtx, request); recordAdapterTier(logCtx, request); refreshUndeclaredToolGuard(request); @@ -5607,7 +5666,7 @@ async function handleResponsesInner( headers: selectedForwardHeaders, translatorBudget, }); - refreshRoutedNamespaceToolAliases(request); + refreshRequestToolAliases(request); recordAdapterReasoning(logCtx, request); recordAdapterTier(logCtx, request); } catch (err) { @@ -5834,7 +5893,7 @@ async function handleResponsesInner( if (retry.kind === "retried") { authCtx = retry.authCtx; request = retry.request; - refreshRoutedNamespaceToolAliases(request); + refreshRequestToolAliases(request); refreshUndeclaredToolGuard(request); upstreamResponse = retry.upstreamResponse; selectedForwardHeaders = retry.selectedForwardHeaders; @@ -5902,9 +5961,39 @@ async function handleResponsesInner( logCtx.terminalIncompleteReason = preflightLog.terminalIncompleteReason; } } + // Console Go (opencode-zen / opencode-go) intermittently rejects a body it accepts seconds + // later with 400 invalid_request_error / "Invalid upload request." Replay the byte-identical + // request once after the exact gateway rejection. Single-shot guard. + // This recovery reuses the captured request; other recovery kinds still rebuild. + if (!consoleGoUploadRetryGuard.attempted) { + const uploadRejectionBody = await consoleGoUploadRejectionBody( + upstreamResponse, + consoleGoUploadRetryGuard.attempted, + upstream.signal, + ); + if (uploadRejectionBody !== undefined + && isTransientConsoleGoUploadRejection({ + status: upstreamResponse.status, + errorBody: uploadRejectionBody, + outboundUrl: request.url, + })) { + consoleGoUploadRetryGuard.attempted = true; + try { void upstreamResponse.body?.cancel().catch(() => {}); } catch { /* already consumed/closed */ } + if (!upstream.signal.aborted) { + try { + await sleepWithAbort(CONSOLE_GO_UPLOAD_RETRY_DELAY_MS, upstream.signal); + } catch { return clientCancelledResponse(); } + } + if (upstream.signal.aborted) return clientCancelledResponse(); + const result = await rebuildAndRefetch("console-go-upload-retry"); + if ("failed" in result) return result.failed; + upstreamResponse = result; + continue passthroughRecovery; + } + } break; } - const headers = sanitizePassthroughHeaders(upstreamResponse.headers); + const headers = sanitizePassthroughHeaders(upstreamResponse.headers, codexSafetyBufferingOptions); const resolvedModel = headers.get("openai-model")?.trim(); if (resolvedModel && !logCtx.preserveResolvedModelFromRoute) logCtx.resolvedModel = resolvedModel; if (isUsageDebugEnabled()) { @@ -5915,7 +6004,7 @@ async function handleResponsesInner( // treating a successful body as SSE when the caller requested streaming. const passthroughCt = headers.get("content-type")?.toLowerCase(); const isEventStream = passthroughCt?.includes("text/event-stream") - || (upstreamResponse.ok && !!upstreamResponse.body && !passthroughCt && parsed.stream); + || (plaintextV2AgentMessageToolNames.size === 0 && upstreamResponse.ok && !!upstreamResponse.body && !passthroughCt && parsed.stream); const recordTerminalOutcome = codexForwardTerminalOutcomeRecorder( config, authCtx, @@ -5993,7 +6082,7 @@ async function handleResponsesInner( return new Response(upstreamResponse.body, { status: upstreamResponse.status, statusText: upstreamResponse.statusText, - headers: sanitizePassthroughHeaders(upstreamResponse.headers), + headers: sanitizePassthroughHeaders(upstreamResponse.headers, codexSafetyBufferingOptions), }); } if (!upstreamResponse.ok) { @@ -6128,15 +6217,23 @@ async function handleResponsesInner( ? createResponsesItemIdPayloadRewrite(repairConfig!, translatorBudget) : undefined, responseModelRewrite, - parsed.options.hideThinkingSummary !== true - && routeUsesContentChannelReasoning(route.provider, route.modelId) - ? createReasoningSummaryChannelPayloadRewrite() - : undefined, ].filter((rewrite): rewrite is NonNullable => rewrite !== undefined); // #893: sparse-snapshot gateways get field backfills AND lifecycle event // injection at the block level, after payload rewrites. Defaults come // from the finalized OUTBOUND body — the normalized internal tool shapes // are not the Responses wire shapes the snapshot must mirror. + // Only validated client blocks may publish plaintext continuation state. + // Raw inspection precedes rewriting on eager relays, so it cannot own this write. + const plaintextInspector = plaintextV2AgentMessageToolNames.size > 0 + ? createSseInspector({ onCompletedResponse: rememberPassthroughResponseChecked }) + : undefined; + const plaintextEncoder = plaintextInspector ? new TextEncoder() : undefined; + const rememberPlaintextBlock = plaintextInspector + ? Object.assign((block: string): readonly string[] => { + plaintextInspector.feed(plaintextEncoder!.encode(`${block}\n\n`)); + return [block]; + }, { dispose: () => plaintextInspector.dispose() }) + : undefined; const blockRewrites = [ payloadRewrites.length > 0 ? payloadRewriteAsBlockRewrite(composeSsePayloadRewrites(...payloadRewrites)) @@ -6164,6 +6261,11 @@ async function handleResponsesInner( snapshotRepairEnabled ? createResponsesSnapshotBlockRewrite(outboundRequestBody, translatorBudget) : undefined, + plaintextV2AgentMessageToolNames.size > 0 + ? payloadRewriteAsBlockRewrite(createPlaintextV2AgentMessageCallRestoreRewrite( + plaintextV2AgentMessageToolNames, plaintextV2AgentMessageAliasedToolNames, + )) + : undefined, createResponsesFieldBackfillBlockRewrite(), functionRepairSchemas.size > 0 ? createResponsesFunctionToolRepairBlockRewrite(functionRepairSchemas, translatorBudget) @@ -6178,6 +6280,7 @@ async function handleResponsesInner( declaredBareWireToolNames, ) : undefined, + rememberPlaintextBlock, ].filter((rewrite): rewrite is NonNullable => rewrite !== undefined); const clientBlockRewrite = blockRewrites.length > 0 ? composeSseBlockRewrites(...blockRewrites) @@ -6225,7 +6328,7 @@ async function handleResponsesInner( const inspector = createSseInspector({ onTerminal: reportNativeTerminal, logCtx, - onCompletedResponse: rememberPassthroughResponse ? rememberPassthroughResponseChecked : undefined, + onCompletedResponse: rememberPassthroughResponse && plaintextV2AgentMessageToolNames.size === 0 ? rememberPassthroughResponseChecked : undefined, onParsedPayload: noteInspectedPayload, onFirstOutput: options.onFirstOutput, pinCompletedResponseIdToFirstSeen: githubCopilotRepairEnabled, @@ -6262,6 +6365,7 @@ async function handleResponsesInner( onDone: () => unregisterTurn(turnAc), }, { clientGoneSignal: options.abortSignal, + terminalBoundary: codexSafetyBufferingOptions, ...(inlineEagerRewrite ? { rewriteBudget: translatorBudget } : {}), ...(logCtx.upstreamError === undefined ? {} : { upstreamError: logCtx.upstreamError }), }); @@ -6324,7 +6428,7 @@ async function handleResponsesInner( responseCompletionCancelled = true; options.onNativePassthroughCancel?.(); }, - rememberPassthroughResponse ? rememberPassthroughResponseChecked : undefined, + rememberPassthroughResponse && plaintextV2AgentMessageToolNames.size === 0 ? rememberPassthroughResponseChecked : undefined, options.onFirstOutput, inspectionConsumerOptions, ); @@ -6334,7 +6438,7 @@ async function handleResponsesInner( logCtx, turnAc.signal, () => unregisterTurn(turnAc), - rememberPassthroughResponse ? rememberPassthroughResponseChecked : undefined, + rememberPassthroughResponse && plaintextV2AgentMessageToolNames.size === 0 ? rememberPassthroughResponseChecked : undefined, options.onFirstOutput, inspectionConsumerOptions, ); @@ -6353,7 +6457,7 @@ async function handleResponsesInner( responseCompletionCancelled = true; clientGone.abort(reason); }, - { upstreamError: logCtx.upstreamError }, + { upstreamError: logCtx.upstreamError, terminalBoundary: codexSafetyBufferingOptions }, ); return markNativePassthroughSseResponse(new Response(clientBody, { status: upstreamResponse.status, @@ -6376,6 +6480,7 @@ async function handleResponsesInner( } const text = bounded.text; inspectResponseLogJson(logCtx, text); + let plaintextV2RestoreFailed = false; let clientJson = (() => { const restoredNamespace = restoreRoutedNamespaceCallsInJson( scrubSelfNamedToolCallNamespaceInJson( @@ -6401,18 +6506,20 @@ async function handleResponsesInner( restored, routedToolSearchNames, ); - const repaired = normalizeFunctionCompletionJson(restoredToolSearch); + const normalizedJson = normalizeFunctionCompletionJson(restoredToolSearch); + const plaintextRestore = restorePlaintextV2AgentMessageCallsInJsonResult( + normalizedJson, plaintextV2AgentMessageToolNames, plaintextV2AgentMessageAliasedToolNames, + ); + plaintextV2RestoreFailed = plaintextRestore.overflowed; + const repaired = plaintextRestore.value; const modelRewritten = parsed._responseModelId !== undefined && parsed._responseModelId !== parsed.modelId ? rewriteResponsesModelJson(repaired, parsed._responseModelId) : repaired; - // The bounded-JSON answer bypasses the SSE payload rewrite, so content- - // channel reasoning needs the same normalization here for the plain - // JSON answer and every reframed-SSE variant built from clientJson. - return parsed.options.hideThinkingSummary !== true - && routeUsesContentChannelReasoning(route.provider, route.modelId) - ? rewriteReasoningSummaryInJsonString(modelRewritten) - : modelRewritten; + return modelRewritten; })(); + if (plaintextV2RestoreFailed) { + return formatErrorResponse(502, "upstream_error", PLAINTEXT_V2_AGENT_MESSAGE_RESTORE_OVERFLOW_MESSAGE); + } // #1700: same fail-closed policy as the SSE relay above. Both the plain JSON answer and // the reframed-SSE branch below are built from this body, so one check covers them. This // runs BEFORE the continuation cache write below: a refused turn must not become state a @@ -6487,7 +6594,7 @@ async function handleResponsesInner( } throw error; } - const sseHeaders = sanitizePassthroughHeaders(headers); + const sseHeaders = sanitizePassthroughHeaders(headers, codexSafetyBufferingOptions); sseHeaders.set("content-type", "text/event-stream"); sseHeaders.set("cache-control", "no-store"); return new Response(stream, { @@ -6520,6 +6627,10 @@ async function handleResponsesInner( headers, }); } + if (plaintextV2AgentMessageToolNames.size > 0) { + try { void upstreamResponse.body?.cancel().catch(() => {}); } catch { /* already closed */ } + return formatErrorResponse(502, "upstream_error", "plaintext V2 agent-message response used an unsupported content type"); + } // An unclassified passthrough body is relayed directly and has no bounded completion observer; // use the same non-error-status success boundary as SSE instead of retaining per-stream state. commitReasoningReplayServingRoute(request.headers); @@ -7288,7 +7399,7 @@ async function handleResponsesInner( Math.max(1, budget - transientSendsUsed); try { initialRequest = await activeAdapter.buildRequest(parsed, { headers: selectedForwardHeaders, translatorBudget }); - refreshRoutedNamespaceToolAliases(initialRequest); + refreshRequestToolAliases(initialRequest); recordAdapterReasoning(logCtx, initialRequest); recordAdapterTier(logCtx, initialRequest); inputTokenEstimate = typeof initialRequest.usageLog?.inputTokens === "number" @@ -7397,6 +7508,9 @@ async function handleResponsesInner( // 413→429 rotation cannot silently undo the tightening. let imageRetryAttempted = false; const opaqueBlobRecoveryGuard: OpaqueBlobRecoveryGuard = { attempted: false }; + // Console Go answers a transient 400 "Invalid upload request." for bodies it accepts + // moments later; at most one byte-identical replay is allowed per request. + const consoleGoUploadRetryGuard: { attempted: boolean } = { attempted: false }; let oauth401ReplayAttempted = false; /** * Rebuild the request from the current parsed input (and any image-tier bias) and refetch @@ -7433,7 +7547,7 @@ async function handleResponsesInner( sameTargetParsed = parsed; sameTargetToken = transportToken; } - refreshRoutedNamespaceToolAliases(retryRequest); + refreshRequestToolAliases(retryRequest); const retryEstimate = typeof retryRequest.usageLog?.inputTokens === "number" ? retryRequest.usageLog.inputTokens : undefined; @@ -7781,6 +7895,35 @@ async function handleResponsesInner( upstreamResponse = result; continue recovery; } + // Console Go (opencode-zen / opencode-go) intermittently rejects a body it accepts seconds + // later with 400 invalid_request_error / "Invalid upload request." Replay the + // byte-identical request once after the exact gateway rejection. + if (!consoleGoUploadRetryGuard.attempted) { + const uploadRejectionBody = await consoleGoUploadRejectionBody( + upstreamResponse, + consoleGoUploadRetryGuard.attempted, + upstream.signal, + ); + if (uploadRejectionBody !== undefined + && isTransientConsoleGoUploadRejection({ + status: upstreamResponse.status, + errorBody: uploadRejectionBody, + outboundUrl: sameTargetRequest?.url, + })) { + consoleGoUploadRetryGuard.attempted = true; + try { void upstreamResponse.body?.cancel().catch(() => {}); } catch { /* already consumed/closed */ } + if (!upstream.signal.aborted) { + try { + await sleepWithAbort(CONSOLE_GO_UPLOAD_RETRY_DELAY_MS, upstream.signal); + } catch { cleanupUpstreamAbort(); return clientCancelledResponse(); } + } + if (upstream.signal.aborted) { cleanupUpstreamAbort(); return clientCancelledResponse(); } + const result = await rebuildAndRefetch("console-go-upload-retry"); + if ("failed" in result) return result.failed; + upstreamResponse = result; + continue recovery; + } + } break; } if (!upstreamResponse.ok) { diff --git a/src/types/config.ts b/src/types/config.ts index 52c4ed8b0e..372c5b82d0 100644 --- a/src/types/config.ts +++ b/src/types/config.ts @@ -380,6 +380,8 @@ export interface OcxConfig { privacy?: OcxPrivacyConfig; /** Opt in to one identical-turn retry when a Responses completion has no text or tool call. */ emptyCompletionRetry?: boolean; + /** Suppress allowlisted client-facing Codex transport hints; provider enforcement is unchanged. */ + dropCodexSafetyBuffering?: boolean; /** * Whether a login may open a browser on the machine running the proxy. * @@ -642,6 +644,8 @@ export interface OcxConfig { * Routed parents get v2 tools; Sol/Terra can still spawn Grok/Claude (issue #92). */ keepNativeChatGptOnV1?: boolean; + /** Experimental plaintext delivery for native v2 collaboration messages; disabled unless true. */ + plaintextV2AgentMessages?: boolean; /** Experimental, default-off ChatGPT recovery for encrypted V2 routed tasks. */ agentTaskRecovery?: { enabled?: boolean; @@ -967,8 +971,9 @@ export type OcxComboDefaultEffort = "low" | "medium" | "high" | "xhigh" | "max" * advertises no effort control (`reasoningEfforts: []`) empties the combo's picker. * `adaptive` excludes those empty ladders from the published intersection, keeping the * control usable for a mixed-capability group. Unknown (`undefined`) ladders stay - * wildcards in both modes. Dispatch is unchanged: each concrete target still resolves - * its own effort at request time. + * wildcards in both modes. An explicit empty ladder removes unsupported effort controls + * in either mode; adaptive dispatch also removes them before sending to an unknown target, + * while each known target still resolves its own effort. */ export type OcxComboReasoningEffortMode = "strict" | "adaptive"; diff --git a/src/types/provider.ts b/src/types/provider.ts index d559ffef8f..bf37ac708a 100644 --- a/src/types/provider.ts +++ b/src/types/provider.ts @@ -773,6 +773,12 @@ export interface OcxProviderConfig { * out explicitly (e.g. MiniMax, where low effort disables thinking). */ requiresReasoningPlaceholderModels?: string[]; + /** + * Default to displaying provider-authored summaries when Responses summary is omitted. + * Explicit wire summary:"none" wins; false disables a seeded provider default. + * Raw reasoning is never relabeled as a summary. + */ + showThinkingSummary?: boolean; /** * Opt-in same-target 429 retry policy. Codex itself never retries 429 (it retries 5xx only, * openai/codex#30471), and single-key pools have no failover, so the proxy waits and replays diff --git a/src/types/request.ts b/src/types/request.ts index 480a0509ad..cec8294a50 100644 --- a/src/types/request.ts +++ b/src/types/request.ts @@ -79,6 +79,8 @@ export interface OcxParsedRequest { * prepareOpaqueBlobRecovery after an authoritative rejection; consumers strip replayed blobs. */ _stripReasoningEncryptedContent?: boolean; + /** Final-route opt-in: emit v2 collaboration message arguments as plaintext on ChatGPT. */ + _plaintextV2AgentMessages?: boolean; /** * Optional authenticated tenant/operator namespace for Cursor thread→conversation derivation. * When absent (single-operator local proxy), derivation stays local-scoped. diff --git a/src/usage/log.ts b/src/usage/log.ts index 2944c22f9a..953195b7ca 100644 --- a/src/usage/log.ts +++ b/src/usage/log.ts @@ -70,6 +70,7 @@ export type AttemptRecoveryKind = | "anthropic-oauth-429" | "oauth-account-429" | "image-413" + | "console-go-upload-retry" | "opaque-blob-rejection" | "empty-completion"; @@ -309,6 +310,7 @@ const ATTEMPT_RECOVERY_KINDS = new Set([ "anthropic-oauth-429", "oauth-account-429", "image-413", + "console-go-upload-retry", "opaque-blob-rejection", "empty-completion", ]); diff --git a/structure/adapters/registry.md b/structure/adapters/registry.md index aa0bc914cb..fe6f7cb994 100644 --- a/structure/adapters/registry.md +++ b/structure/adapters/registry.md @@ -1,5 +1,8 @@ # Adapter Registry Authority +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Decision Runtime adapter construction has one authority: `src/adapters/registry.ts`. @@ -72,6 +75,15 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). + +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Adapter events distinguish raw reasoning content from summary-channel thinking; CCA Gemini classification is request-local. See [Google provenance](../providers/google.md). + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/catalog.md b/structure/catalog.md index 49f0e353e1..a6a6250275 100644 --- a/structure/catalog.md +++ b/structure/catalog.md @@ -1,5 +1,8 @@ # Model Catalog +The configuration-only [plaintext V2 contract](subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Shared catalog `src/codex/catalog.ts` builds a shared Codex-shaped catalog for CLI, TUI, App, and SDK. It: @@ -240,6 +243,11 @@ wire-clamps ultra/max to each model's real top rung (e.g. gpt-5.5 ultra → xhig (`src/server/effort-policy.ts`): they lower or preserve the requested effort rather than rejecting the request, and they never raise it. +Combo dispatch reads the final target ladder through the same `supportedLadderFor` authority. An +explicit empty ladder means that target receives no effort control; an unknown ladder receives no +parent effort controls only when the combo opts into `reasoningEffortMode: "adaptive"`. Known +non-empty ladders continue through the existing per-target resolution. + The `ocx effort` CLI accepts only the same canonical cap ladder before live probing or persistence. Its status output preserves unsupported legacy cap values and reports that those fields are ignored; the read does not normalize or migrate them, and an ignored subagent field does not disable a valid @@ -274,6 +282,11 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](transports/responses.md). + +Provider `showThinkingSummary` is a Responses request default; it does not rewrite catalog summary defaults or client configuration. See [Google summaries](providers/google.md). + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. diff --git a/structure/clients/claude-desktop.md b/structure/clients/claude-desktop.md index 54689b36a1..e512b469fc 100644 --- a/structure/clients/claude-desktop.md +++ b/structure/clients/claude-desktop.md @@ -1,5 +1,8 @@ # Claude Desktop Integration +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Connected Claude Desktop profiles Connected `ocx claude desktop apply` reads the hub's Desktop snapshot and writes the hub origin @@ -83,6 +86,11 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Provider summary defaults are Responses-specific and do not rewrite connected Claude Desktop profiles. See [inbound compatibility](../data-planes/inbound-compat.md). + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. @@ -95,3 +103,5 @@ Config JSON preserves the boolean; only literal true activates the role-changing The lightweight top-level CLI help counts Cline CLI among the fifteen registered export clients; registry parity remains covered by the client help and integration tests. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/config.md b/structure/config.md index b37d315fce..7001e54c92 100644 --- a/structure/config.md +++ b/structure/config.md @@ -1,5 +1,8 @@ # Config Surface +The configuration-only [plaintext V2 contract](subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Config surface ### OpenCodex home and live process state @@ -45,7 +48,7 @@ matters for maintainers is which groups exist and who resolves them: | Group | Keys | Resolution rule | | --- | --- | --- | | Listener | `port`, `hostname` | The listener owns the port; `runtime-port.json` reports where it actually landed. | -| Routing | `defaultProvider`, `providers`, per-provider `selectedModels` | Explicit `provider/model` wins over `defaultProvider`. | +| Routing | `defaultProvider`, `providers`, per-provider `selectedModels`, `combos` | Explicit `provider/model` wins over `defaultProvider`; combo dispatch uses the selected target's existing capability ladder and does not create a second catalog authority. | | Catalog | `disabledModels`, `customModels`, `modelCacheTtlMs`, `providerContextCaps`, `contextCapValue`, per-provider `modelDisplayNames`, `codexAccountNamespaces`, `codexAccountPickerEnabled` | Catalog state is derived; config only records intent. Exact provider model display names are durable display only overlays. The picker flag is an explicit visibility override, while selector mappings remain the durable exact-routing contract. | | Retained state | `appOwnedMemoryBudgetMb` | Process-wide eviction target for app-owned logs, caches, blobs, and continuation payloads. Default 256 MiB, valid 64..4096; pinned state may temporarily exceed the target, but every pin-capable store has a finite local cap and their documented aggregate stays below `APP_OWNED_WORST_CASE_PINNED_BYTES` (512 MiB). Neither value caps RSS or native runtime memory. | | Transport | stream mode, timeouts, proxy settings, `websockets`, `emptyCompletionRetry` | `streamMode` persists in config.json; Windows services need a persisted input, and macOS uses it for explicit eager-relay opt-in. Empty-completion replay is an explicit top-level opt-in because its second upstream request may be billable. | @@ -203,6 +206,10 @@ Client connection metadata stores a stable `apiKeyId` and a non-secret rotation Codex display-cache expiry, retained main-policy evidence, and reset history follow the [quota cache contract](providers/openai-tiers.md#quota-cache-and-short-window-history). +`dropCodexSafetyBuffering` is an optional boolean, default false. Invalid API candidates reject; +malformed persisted values stay disabled. It controls only the allowlisted client-output hints +described in [Responses transport](transports/responses.md), not upstream policy or model selection. + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. diff --git a/structure/data-planes/images.md b/structure/data-planes/images.md index d1d0193048..0799866f21 100644 --- a/structure/data-planes/images.md +++ b/structure/data-planes/images.md @@ -1,5 +1,8 @@ # Images Data Plane +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Standalone Images Codex's local `image_gen.imagegen` tool makes a second Images request after the model calls it: @@ -77,6 +80,13 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +CCA image-capable requests do not acquire the text-summary includeThoughts opt-in. See [Google summary boundary](../providers/google.md). + Claude replay carries [Go conversation affinity](inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/data-planes/inbound-compat.md b/structure/data-planes/inbound-compat.md index 01b3db72a8..bde85d4a32 100644 --- a/structure/data-planes/inbound-compat.md +++ b/structure/data-planes/inbound-compat.md @@ -1,5 +1,8 @@ # Inbound Compatibility Surfaces +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Standalone file transcription `src/server/audio-transcriptions.ts` owns `POST /v1/audio/transcriptions`, independently of @@ -36,6 +39,9 @@ the native passthrough there is no canonical Fast injection and no wire mapping: and `fastMode` injects nothing here. Resolved-Fast-policy injection applies only to routes that take the Chat -> Responses -> Chat bridge below. `parallel_tool_calls` is emitted only for providers opted into parallel tools (or pinned false by the existing provider opt-out contract). +The native passthrough still applies the existing model capability authority to reasoning: an +explicit empty ladder removes caller `reasoning_effort`, while an unknown ladder remains +unclassified. This guard does not alter the separate raw service-tier contract. On the response side, the upstream `service_tier` echo (xAI Priority Processing, OpenAI fast tier) relays to the Chat Completions caller on every delivery shape: the non-streaming body @@ -135,6 +141,13 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). + +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +The provider summary default applies at Responses ingress; native Chat and Anthropic inbound preferences keep their existing handling. Raw content is never renamed to a summary. See [bridge contract](../providers/chat-compat.md). + ## Claude affinity at final Go dispatch `src/server/claude-messages.ts` carries validated conversation affinity privately through diff --git a/structure/gui-and-management-api.md b/structure/gui-and-management-api.md index 9fde77fc9e..76ccd409b8 100644 --- a/structure/gui-and-management-api.md +++ b/structure/gui-and-management-api.md @@ -1,5 +1,8 @@ # GUI And Management API +The configuration-only [plaintext V2 contract](subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Dashboard serving The bundled React dashboard is built into `gui/dist` and served by the same Bun proxy. `ocx gui` @@ -531,6 +534,11 @@ expiry, including a deadline crossed before effects run, rechecks activation and refreshes quota with Combo data while preserving drafts. Each successful quota snapshot also advances the observation clock, so a retained older row cannot defer evaluation of a fresh row. +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](transports/responses.md). + +The provider editor field policy exposes `showThinkingSummary` as a boolean provider option; it controls Responses summary defaults without a dashboard rendering change. See [Google provider](providers/google.md). + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. @@ -544,3 +552,5 @@ integration IO adapter. Its snapshot fingerprint cannot be checked against provi The existing dashboard file-client maps include Cline CLI and reuse its committed color mark. The export panel labels its download as a settings/catalog bundle; all locales explain that Undo restores both original files. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/ops/docs-and-release.md b/structure/ops/docs-and-release.md index 9570086e21..625fec16b2 100644 --- a/structure/ops/docs-and-release.md +++ b/structure/ops/docs-and-release.md @@ -1,5 +1,8 @@ # Docs And Release +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Public docs The public documentation site lives in `docs-site/` and is built with Astro + Starlight. English is @@ -309,6 +312,13 @@ The Combo guides describe the distinction between display quota and single-crede The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). + +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Provider configuration documents distinguish actual summaries from raw reasoning content. The test layout registers the summary-default contract cases and removes the obsolete content-rewrite test with its implementation. + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](../codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. @@ -318,3 +328,5 @@ The integrations guide documents Cline CLI as a two-file, loopback-only integrat The lightweight top-level CLI help counts Cline CLI among the fifteen registered export clients; registry parity remains covered by the client help and integration tests. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/ops/service-and-sidecars.md b/structure/ops/service-and-sidecars.md index c8ef069110..8b7a745165 100644 --- a/structure/ops/service-and-sidecars.md +++ b/structure/ops/service-and-sidecars.md @@ -1,5 +1,8 @@ # Background Service And Sidecars +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Background service command selection A bare `ocx service` is an idempotent install-or-repair command. Argument validation happens before @@ -140,6 +143,13 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Provider summary defaults are evaluated per routed Responses request without changing service lifecycle or sidecar activation. See [runtime](../runtime.md). + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/overview.md b/structure/overview.md index d5d1a2207a..dbc6c88360 100644 --- a/structure/overview.md +++ b/structure/overview.md @@ -1,5 +1,8 @@ # Overview +The configuration-only [plaintext V2 contract](subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Product boundary opencodex is a local proxy for Codex. It does not patch Codex binaries. It changes local Codex @@ -107,4 +110,9 @@ would pass while the rule was violated. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](transports/responses.md). + +Raw reasoning content and provider-authored summaries remain distinct on the Responses wire. See [reasoning presentation](providers/chat-compat.md). + Cline CLI is a managed file integration: its provider settings and catalog share one recoverable journal operation. The [paired-file contract](clients/integrations.md#cline-paired-files) defines its stop/restart requirement. diff --git a/structure/providers/chat-compat.md b/structure/providers/chat-compat.md index 7cf84ef7ec..6aefae99f7 100644 --- a/structure/providers/chat-compat.md +++ b/structure/providers/chat-compat.md @@ -1,5 +1,8 @@ # Chat Provider Compatibility +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Reasoning and tool-result compatibility Kiro groups only consecutive original-message tool results whose raw call ID exactly matches @@ -22,7 +25,9 @@ content converter after validation and imports no optional subsystem. Stateful developer-guidance injection reuses that validator for its raw insertion boundary, so parsed messages and stored raw history retain the same task/guidance order. -Native OpenAI passthrough sanitizes routed reasoning history so `reasoning` input items do not send +Native OpenAI passthrough consults the existing configured capability ladder before forwarding +`reasoning_effort`; an explicitly empty ladder removes that unsupported control while an unknown +ladder remains unclassified. It also sanitizes routed reasoning history so `reasoning` input items do not send non-empty `content` arrays to upstream models that reject them. Chat Completions bridging repairs orphan `toolResult` messages by inserting a synthetic assistant `tool_call` before tool messages. It also repairs the opposite direction (260718): an assistant `tool_calls` round left dangling — @@ -220,18 +225,13 @@ honored by BOTH reasoning paths: anthropic `thinking_delta` AND raw `reasoning_r item (`summary: []`, txt-only `ocxr1:` `encrypted_content`, no text deltas) — invisible in the Codex app, so tool cells group like native models — while the text still round-trips for `preserveReasoningContentModels` replay. Visible mode (summary "auto") keeps the raw -`content[reasoning_text]` shape. Diagnosis and codex-rs grouping evidence: -`devlog/_fin/260709_native_response_pattern/`. - -The content-to-summary channel rewrite skips any reasoning item that carries a native -`encrypted_content` blob. The blob is opaque, state-bearing provider data, so the item must -round-trip unchanged unless that backend has an explicit replay contract permitting a rewrite. -This defensively protects providers that issue blobs and later join the route through -`preserveReasoningContentModels`. The rewrite's round trip was verified against DeepSeek, which is -`statelessResponses` and issues no blob. Grok is unaffected in practice because it natively emits -summary-channel reasoning and no `reasoning_text` events, so this content-to-summary item rewrite -does not engage on its route. Only the stored item is exempt — `reasoning_text` delta events carry -no blob and still route to the summary channel, so the live expandable trace is unchanged. +`content[reasoning_text]` shape: raw deltas stream as `response.reasoning_text.delta` and the final +item carries `content: [{type: "reasoning_text", text}]`, so Codex applies its own display policy — +the desktop thinking band shows the "Thinking…" placeholder, and raw text appears only when +`show_raw_agent_reasoning` is enabled. Routing raw CoT through the summary channel instead (the +#45 display intent, intentionally reverted 260911) put unsummarized thinking in the desktop band, +which only fits native OpenAI providers that author real summaries. Diagnosis and codex-rs +grouping evidence: `devlog/_fin/260709_native_response_pattern/`. The process-local raw-reasoning fallback is fail-closed unless a request has an explicit client thread plus an exact provider destination, wire adapter, final model, and physical credential @@ -269,3 +269,5 @@ fragments are not guessed onto pending ID-only calls. parallel/colliding identities, distinct unsafe raw JSON index literals, the maximum safe-integer boundary, invalid index types, missing/null continuations and UTF-8 byte-limit boundaries. + +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). diff --git a/structure/providers/cursor.md b/structure/providers/cursor.md index be42793e0b..c79b8085c7 100644 --- a/structure/providers/cursor.md +++ b/structure/providers/cursor.md @@ -1,5 +1,8 @@ # Cursor Provider +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Cursor Native Exec Cursor's experimental live transport can receive server-driven local read/write/delete/ls/grep, @@ -82,3 +85,9 @@ constraints cannot widen the canonical shape. Bare shell bridge names are reject on the freeform path. Namespaced tools do not acquire bare-shell behavior. Regression coverage lives in `tests/providers/cursor/cursor-tool-definitions.test.ts`. + +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). + +Shared raw-reasoning events retain content-channel presentation; provider-authored thinking keeps its existing summary path. See [bridge contract](chat-compat.md). + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/providers/google.md b/structure/providers/google.md index 26187aeebf..1c1d5648f8 100644 --- a/structure/providers/google.md +++ b/structure/providers/google.md @@ -2,12 +2,14 @@ ## Google thought-text visibility boundary -Google-family responses may represent model-internal reasoning as a text-bearing part with -`thought: true`. The Google adapter maps that text to the internal `reasoning_raw_delta` event; -only text without the marker becomes visible `text_delta`. Streaming SSE and buffered JSON share -one classifier so transport selection cannot change whether provider-declared reasoning is shown -as assistant output. Thought-signature observation still runs on the original parts before text -classification, preserving the opaque continuation state independently of display semantics. +Google-family parts with `thought: true` stay separate from assistant output. After a CCA +Gemini request is built, the shared streaming/buffered classifier emits `thinking_delta` for +these provider-authored summaries. Other Google wires, non-Gemini CCA models and uninitialized +adapters retain `reasoning_raw_delta`. Model provenance is refreshed on every build. +`showThinkingSummary` defaults on only for the Antigravity preset; explicit provider false and +explicit wire summary none win. Eligible CCA Gemini requests use `includeThoughts: true` only +when provider opt-in and per-request display both allow it. Thought signatures remain attached +to their tool calls independently; they never become Anthropic thinking signatures. > Decision record: [ADR-0055](../decisions/ADR-0055-google-thought-text-visibility-boundary.md) diff --git a/structure/providers/kiro.md b/structure/providers/kiro.md index 4881f628f4..c499d44428 100644 --- a/structure/providers/kiro.md +++ b/structure/providers/kiro.md @@ -1,5 +1,8 @@ # Kiro Provider +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Kiro client parallel-tool hint Kiro's wire remains serialized even when an OpenAI Responses client sends diff --git a/structure/providers/openai-tiers.md b/structure/providers/openai-tiers.md index 8080c8bf04..068f43dbc8 100644 --- a/structure/providers/openai-tiers.md +++ b/structure/providers/openai-tiers.md @@ -1,5 +1,8 @@ # OpenAI Provider Account Modes +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + This current contract supersedes the provider-identity and account-selection sections of `devlog/_fin/260717_openai_hardening`; that archived unit remains historical evidence for the earlier three-tier implementation. The replacement contract and its verification evidence live in @@ -424,6 +427,9 @@ model settings, and noncanonical `openai` rows never receive that recovery path. terminal 403 codes as `needsReauth`; generic permission failures remain non-terminal, and a successful main usage refresh clears the runtime mark. +Canonical forwarding alone can apply the optional client-output safety-buffering hint filter; +API-key and custom forward destinations preserve their metadata. See [Responses transport](../transports/responses.md). + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](../codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. diff --git a/structure/providers/xai-grok.md b/structure/providers/xai-grok.md index fc6cf7c63f..8280356d3a 100644 --- a/structure/providers/xai-grok.md +++ b/structure/providers/xai-grok.md @@ -1,5 +1,8 @@ # xAI Grok Provider +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## xAI Grok hardening (official Grok Build contract parity) Grounded in the open-sourced official client (xai-org/grok-build); unit + evidence: @@ -63,12 +66,20 @@ Account-scoped OAuth quota remains display evidence for provider-level Combo sel The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Grok chat raw reasoning uses content-channel output with an empty summary; hidden replay envelopes retain continuation text. Native Responses content is not promoted to summaries. See [chat compatibility](chat-compat.md). + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Devin CLI credential path composition in `src/oauth/devin-cli.ts` follows the selected platform: Windows uses Win32 APPDATA paths, other platforms use POSIX XDG-data paths. The explicit absolute override remains verbatim; credential parsing and login behavior are unchanged. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](../transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. + ### OAuth Fast Tier (Priority Processing) xAI's Priority Processing (`service_tier: "priority"` on Chat Completions and Responses, diff --git a/structure/runtime.md b/structure/runtime.md index 94541ef247..4fc022c552 100644 --- a/structure/runtime.md +++ b/structure/runtime.md @@ -1,5 +1,8 @@ # Runtime +The configuration-only [plaintext V2 contract](subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Entrypoints | Path | Responsibility | @@ -147,6 +150,7 @@ The server exposes `POST /api/stop` which restores native Codex config, stops an | `src/providers/registry.ts` | Canonical provider presets for CLI, dashboard, OAuth, key providers, and metadata. | | `src/providers/derive.ts` | Enrichment from provider presets into user config. | | `src/oauth/` | OAuth providers, token storage, refresh, and auth-token resolution. The login callback listener binds a per-provider FIXED loopback port, so consecutive logins reuse the same number; every response it sends ends its connection (`Connection: close`, including non-callback paths such as a stray `/favicon.ico` 404). Stopping the listener does not close an established socket, so without that a pooled client would deliver the next login's callback to the retired flow, which rejects the unknown state as a CSRF mismatch while the live flow waits. | +| `src/combos/request.ts` | Clones each selected combo target request and applies the existing target capability ladder: adaptive unknown targets and explicit empty ladders receive no unsupported reasoning/thinking controls, while known ladders retain per-target resolution. | | `src/adapters/openai-responses.ts` | Native OpenAI/ChatGPT Responses passthrough. | | `src/responses/muse-tool-name-alias.ts` | Host-gated Meta Muse 64-char tool-name alias/restore used by the Responses passthrough. | | `src/adapters/openai-chat.ts` | OpenAI-compatible Chat Completions bridge. Its client delivery shapes in `src/chat/outbound.ts` and `src/server/chat-native-sse.ts` relay the upstream `service_tier` echo on non-stream, folded-stream, and synthesized-SSE bodies, never inventing the key when the upstream omits it. | @@ -221,6 +225,13 @@ cooldowns and response-driven retry remain authoritative. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](gui-and-management-api.md#combo-editor-routing-quota). +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](transports/responses.md). + +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](transports/responses.md). + +Responses route normalization resolves provider summary defaults from the original wire preference on every final route. See [reasoning presentation](providers/chat-compat.md) and [CCA summary provenance](providers/google.md). + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. diff --git a/structure/subagents.md b/structure/subagents.md index 24d37af883..7560c235ff 100644 --- a/structure/subagents.md +++ b/structure/subagents.md @@ -1,5 +1,30 @@ # Subagents And Multi-Agent Surface +## Plaintext V2 agent messages + +`src/responses/plaintext-v2-agent-messages.ts` owns the experimental, configuration-only +`plaintextV2AgentMessages` request compiler and response restoration. The default is unset; +only explicit true on Responses ingress to the final canonical ChatGPT forward route activates it. +A default top-level collaboration catalog is required. The compiler preserves caller objects, +aliases the namespace and three message functions, and removes only their true encryption marker. +Declaration/reference collisions refuse the whole rewrite without changing the request. + +`src/adapters/openai-responses.ts` returns request-local alias capabilities. The Responses core +refreshes them after every request rebuild and restores JSON, SSE and WebSocket identities after +snapshot repair. Malformed, conflicting, unsupported or over-limit responses fail closed without +retrying the model. Raw stream inspection cannot publish plaintext continuation state: only +restored client blocks reach its dedicated bounded collector. Foreign namespaces and opaque +argument/metadata values remain unchanged; the empty encrypted-function-args marker is preserved. + +Startup warns that task text can remain in Codex history, selected-provider requests and local +response/debug state. This is application-level plaintext over HTTPS, depends on undocumented +upstream behavior, and does not decrypt existing tasks or replace authenticated recovery. + +Restored calls and selectors carry an explicit collaboration namespace and unqualified child name. +Codex treats qualified names literally and defaults absent namespaces to functions. Only child +declarations inherit their restored namespace container; the compiler never invents an empty +encryption marker when the upstream omitted it or returned a nonempty marker. + ## Multi-agent surface mode (3-state) `OcxConfig.multiAgentMode` controls the `multi_agent_version` field stamped on catalog entries: @@ -208,6 +233,11 @@ Provider-level Combo eligibility uses explicit inference evidence for the curren The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](transports/responses.md). + +Final-route summary visibility is recomputed after fallback from the original Responses preference; an earlier provider opt-in does not carry into a later provider. See [reasoning presentation](providers/chat-compat.md). + ## Paginated history writer boundary `src/codex/history-provider.ts` refuses external writes to paginated or migration-capable history. `src/codex/inject.ts` checks affected rows and manifest-owned restore targets before and after config/profile/journal changes, including successful journal and fallback restores, and compensates detected migration. Failed config restore stops later catalog/history work and rolls back a coordinated remove transition. See the [history writer contract](codex-home.md#paginated-history-writer-boundary) for guarantees and concurrent-writer limits. @@ -215,3 +245,5 @@ see [Combo editor routing quota](gui-and-management-api.md#combo-editor-routing- Claude replay carries [Go conversation affinity](data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](transports/responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/structure/transports/inventory.md b/structure/transports/inventory.md index 623abe10a8..0cdb54b0a1 100644 --- a/structure/transports/inventory.md +++ b/structure/transports/inventory.md @@ -1,5 +1,8 @@ # Transport Inventory +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Transport inventory The sections above cover the transports with load-bearing invariants. The rest of the transport @@ -15,7 +18,7 @@ surface is listed here so a maintainer can find the owner without grepping: | Adapter execution support | `src/adapters/run-turn-queue.ts`, `src/adapters/tool-catalog-nudge.ts`, `src/adapters/identity.ts`, `src/adapters/image.ts`, `src/adapters/upstream-http-error.ts` | Shared machinery: turn ordering, tool-catalog nudging, client fingerprinting, image conversion, upstream error normalization. | | Cursor (beyond the sections above) | `src/adapters/cursor/live-transport.ts`, `src/adapters/cursor/http1-bidi.ts`, `src/adapters/cursor/live-models.ts`, `src/adapters/cursor/transport-retry.ts`, `src/adapters/cursor/mcp-manager.ts`, `src/adapters/cursor/thread-continuity.ts`, `src/adapters/cursor/checkpoint-store.ts` | Thread continuity is the point: a retry must not start a new Cursor thread, and a validated checkpoint must not rebuild the full root history. HTTP/2 remains the default; an explicit `http1.1`/`h1` pin maps the bidi run onto Cursor's `RunSSE` receive stream plus sequenced `BidiAppend` sends, and applies to live discovery too. | | Claude Messages | `src/server/claude-messages.ts` | Routed translation, a native Anthropic passthrough branch, and `count_tokens`. | -| Chat Completions inbound | `src/server/chat-completions.ts`, `src/chat/` | Inbound translation onto the same routing pipeline. The content mapper preserves image URLs and supported detail, including screenshot-bearing tool results; target adapters own image placement on their wire. Image-free tool results stay strings. On the response side, the upstream `service_tier` echo relays on every delivery shape (`src/chat/outbound.ts` projections, `src/server/chat-native-sse.ts` chunks); an upstream without the field gets no injected key. | +| Chat Completions inbound | `src/server/chat-completions.ts`, `src/server/chat-native.ts`, `src/chat/`, `src/adapters/openai-chat.ts` | Inbound translation onto the same routing pipeline. The content mapper preserves image URLs and supported detail, including screenshot-bearing tool results; target adapters own image placement on their wire. Image-free tool results stay strings. The native handler owns pin/cap normalization; the adapter wire builder removes effort only for explicit empty declarations or no-reasoning models, preserving unknown raw declarations. On the response side, the upstream `service_tier` echo relays on every delivery shape (`src/chat/outbound.ts` projections, `src/server/chat-native-sse.ts` chunks); an upstream without the field gets no injected key. | | Hosted search relay | `src/server/search.ts` | Direct relay; distinct from the web-search sidecar loop below. | | Image/video generation loop | `src/images/loop.ts`, `src/images/plan.ts`, `src/images/fulfill.ts`, `src/images/xai-client.ts`, `src/images/xai-video-client.ts`, `src/images/artifacts.ts` | A provider-returned image URL is downloaded into a local artifact once, then served locally; warnings stay URL-free because provider CDN URLs may embed credentials. | | GitHub Copilot | `src/providers/xai-transport.ts` (`resolveProviderTransport`), `src/providers/github-copilot-transport.ts` | `resolveProviderTransport` selects the Copilot transport when the routed provider name is `github-copilot`; the Copilot module then resolves its headers and base URL, and the registry seeds the provider row and model fallback. | @@ -69,6 +72,13 @@ Quota publication distinguishes display reports from explicitly supplied inferen The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Canonical Spark Lite metadata follows the final serialized model and surviving nonempty Lite tool catalog; see [Responses transport](../transports/responses.md). + +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +CCA Gemini summary provenance and request opt-in are specified in [Google provider](../providers/google.md); raw Responses content retains its wire channel. + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. diff --git a/structure/transports/responses.md b/structure/transports/responses.md index 9b2c0f883e..a2860c09c0 100644 --- a/structure/transports/responses.md +++ b/structure/transports/responses.md @@ -1,5 +1,10 @@ # Responses Transport +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + +Plaintext collaboration restoration treats a null namespace as absent, rejects non-string namespace types, and restores the native namespace/name pair before HTTP/WS delivery and continuation publication. + ## Responses HTTP/SSE `/v1/responses` is the main Codex-facing endpoint. The server parses Responses input, routes to a @@ -316,14 +321,11 @@ custom result has no local call, because its original wire type cannot be establ would send an unmatched result upstream. The check resolves the selected wire protocol and the request's own tool declarations after final route selection, so stateful destinations keep their upstream-owned native function and native-only custom continuations. Explicit input still receives -orphan repair; this path asks the client to replay rather than reconstructing history. This flag also enables the existing -visible content-to-summary rewrite for SSE and JSON; summary-channel items and opaque reasoning -blobs keep their existing response handling. The shared recording callback applies the same -reasoning rewrite under the exact client-visible predicate before caching output, after tool -restoration and function normalization. This keeps full-content replay fingerprints comparable -for both full-history-plus-ID and delta continuations without weakening identity checks. Hidden -summaries and opaque blobs keep their existing cache representation. It does not change streaming selection or Chat -model routes. Go fixtures cover Luna, Grok and Muse against both response formats. +orphan repair; this path asks the client to replay rather than reconstructing history. Content-channel reasoning stays content in SSE, JSON and stored replay output; native +summary items and opaque blobs retain their upstream representation. Full-content replay +fingerprints compare the same client-visible items without content-to-summary conversion. +It does not change streaming selection or Chat model routes. Go fixtures cover Luna, Grok +and Muse against both response formats. The canonical OpenCode Go transport also derives `x-opencode-session` from the existing hashed session lane before per-model wire selection. One conversation keeps one opaque affinity value @@ -424,6 +426,16 @@ final outgoing model/tier. No caller identity is synthesized. Noncanonical opt-in gateways keep their own metadata policy. Oversized/unsupported-runtime HTTP fallback preserves the original HTTP body and Lite header. +For the final wire model `gpt-5.3-codex-spark`, the canonical forward adapter normalizes the +Lite header from the BODY, overriding caller/configured headers and stale native WS Lite +metadata in both directions. A body carrying a nonempty `additional_tools` input item is +pinned to `true`: that item IS the Lite tool-delivery format and the non-Lite wire shape +expects top-level `tools`, so an inherited `false` would advertise non-Lite while the tools +exist only in the Lite shape and hide the client tool surface. Any other Spark body is set to +`false`, selecting the non-Lite framing policy. A changed Lite identity retires the previous socket; +subsequent eligible Spark requests with the same identity can reuse the new socket. Malformed +native metadata retains HTTP fallback eligibility without rewriting its body. + Canonical WS quota and response metadata preceding the first Responses event are projected into bounded, allowlisted HTTP headers before the response is committed. Later quota observations update only the captured serving account; @@ -499,6 +511,10 @@ retried. Guarded paths: the ChatGPT passthrough and generic adapter fetch in fallback. Adapters with their own `fetchResponse` (kiro, cursor, google) keep their own retry policies; kiro imports the shared abort/sleep helpers from this module. +## Console upload rejection recovery + +`src/providers/opencode-zen-rate-limit.ts` recognizes the complete Console upload-rejection envelope only at the effective HTTPS opencode.ai Zen/Go generation endpoint. A provider row name cannot authorize another destination. The two recovery loops in `src/server/responses/core.ts` wait 800 ms and replay the captured serialized request once; cancellation, nonreplayable responses, other errors and a second upload rejection keep their failure semantics. The recovery kind is persisted as `console-go-upload-retry` and has a localized Logs label. + ## Same-provider combo quota fallback For a failover combo with multiple models on the same Codex-login OpenAI provider, a pre-stream @@ -511,6 +527,15 @@ combo whose remaining eligible targets use other providers. > Decision record: [ADR-0070](../decisions/ADR-0070-same-provider-combo-quota-fallback.md) +## Combo per-target reasoning controls + +`src/server/responses/core.ts` passes the combo's `reasoningEffortMode` and the final target's +`supportedLadderFor` result to `src/combos/request.ts` before adapter parsing. Explicit empty +capability ladders remove effort and thinking controls in every combo mode; adaptive mode also +removes those controls for unknown ladders and preserves `reasoning.summary`. Known non-empty +ladders retain the existing per-target effort resolution. This request normalization does not +change target order, attempt accounting, or the existing provider-400 failover classification. + ## Combo streaming commit boundary An HTTP 200 does not by itself commit a streaming combo child. The combo parent runs the child's @@ -536,6 +561,8 @@ not retried. > Decision record: [ADR-0071](../decisions/ADR-0071-combo-streaming-commit-boundary.md) +Console upload-rejection recovery excludes query-bearing and fragment-bearing destinations even when their host and generation path match the canonical endpoint. + Chat helper admission in `src/server/responses/core.ts` follows the [deferred stored-main contract](../providers/openai-tiers.md): only a needed Direct OpenAI helper claims stored main, after terminal vision, routed vision and search exclusions. @@ -543,6 +570,17 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Spark Lite and routing metadata use the same suffix-normalized model object as serialization, including configured bracket-suffix removal. + +## Optional client transport hints + +`dropCodexSafetyBuffering` defaults to false. Canonical OpenAI forward Responses can remove only +the two safety-buffering response headers, matching response.metadata events and top-level +safety_buffering fields. Pull/eager client output boundaries compose this with policy failure +normalization; refusal/error semantics, retryability, cancellation and captured EOF errors remain +intact. Internal inspection observes original upstream frames. Native codex.response.metadata.headers +WebSocket metadata and compact are excluded. This does not disable upstream safety enforcement. + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. diff --git a/structure/transports/streaming-health.md b/structure/transports/streaming-health.md index 6a5574c6d2..71cb6cd72e 100644 --- a/structure/transports/streaming-health.md +++ b/structure/transports/streaming-health.md @@ -1,5 +1,8 @@ # Streaming Health And WebSocket +The configuration-only [plaintext V2 contract](../subagents.md#plaintext-v2-agent-messages) +is scoped to canonical ChatGPT Responses forwarding; other source-area behavior described here is unchanged. + ## Heartbeat and stall deadline The HTTP/SSE bridge emits an SSE comment-line keep-alive (`: opencodex heartbeat`) during upstream @@ -197,6 +200,13 @@ claims stored main, after terminal vision, routed vision and search exclusions. The management quota DTO keeps Combo editing aligned with scoped inference evidence; see [Combo editor routing quota](../gui-and-management-api.md#combo-editor-routing-quota). +Optional Codex transport-hint suppression is scoped to canonical Responses client output; +its defaults and exclusions are owned by [Responses transport](../transports/responses.md). + +Raw reasoning and provider-authored summary deltas both remain real upstream activity; visibility does not change heartbeat or terminal ownership. See [reasoning presentation](../providers/chat-compat.md). + Claude replay carries [Go conversation affinity](../data-planes/inbound-compat.md#claude-affinity-at-final-go-dispatch) privately to final dispatch; preliminary route selection does not inject Go-only headers. Native Chat applies qualifying effort ceilings independently of model pins; pin selection precedes the cap and only pins or cap rewrites enter wire mapping. The [catalog effort contract](../catalog.md#ultra-reasoning-level) records the V1/compaction exemptions and caller-preservation boundary. + +Combo child requests normalize effort and thinking controls against the selected target while retaining reasoning summaries; strict unknown targets preserve caller controls. The [Responses transport owner](responses.md) documents this boundary, and native Chat removes effort only for an explicit empty declaration or no-reasoning model. diff --git a/tests/adapters/bridge-raw-reasoning-hidden.test.ts b/tests/adapters/bridge-raw-reasoning-hidden.test.ts index 4acc66ce62..162d98b206 100644 --- a/tests/adapters/bridge-raw-reasoning-hidden.test.ts +++ b/tests/adapters/bridge-raw-reasoning-hidden.test.ts @@ -77,18 +77,19 @@ describe("hidden raw reasoning (hideThinkingSummary parity for reasoning_raw_del expect(fc).toMatchObject({ call_id: "call_1", name: "read_file" }); }); - test("streamed visible (flag off): raw reasoning rides the expandable summary channel (#2007)", async () => { + test("streamed visible (flag off): raw reasoning rides the content channel (#2007)", async () => { const frames = await collectSse(bridgeToResponsesSSE(replay([ { type: "reasoning_raw_delta", text: "visible raw" }, { type: "done" }, ]), "routed/model")); - expect(frames.some(f => f.event === "response.reasoning_summary_text.delta")).toBe(true); - expect(frames.some(f => f.event === "response.reasoning_text.delta")).toBe(false); + expect(frames.some(f => f.event === "response.reasoning_text.delta")).toBe(true); + expect(frames.some(f => f.event === "response.reasoning_summary_text.delta")).toBe(false); const completed = frames.find(f => f.event === "response.completed")?.data.response as Record; const output = completed.output as Record[]; expect(output[0]).toMatchObject({ type: "reasoning", - summary: [{ type: "summary_text", text: "visible raw" }], + summary: [], + content: [{ type: "reasoning_text", text: "visible raw" }], }); }); @@ -118,14 +119,15 @@ describe("hidden raw reasoning (hideThinkingSummary parity for reasoning_raw_del expect(decodeReasoningEnvelope(reasoning.encrypted_content as string)?.txt).toBe("quiet"); }); - test("non-streaming visible: raw reasoning lands in the summary channel (#2007)", () => { + test("non-streaming visible: raw reasoning lands on the content channel (#2007)", () => { const json = buildResponseJSON([ { type: "reasoning_raw_delta", text: "loud" }, { type: "done" }, ], "routed/model", {}); const output = (json as { output: Record[] }).output; expect(output.find(o => o.type === "reasoning")).toMatchObject({ - summary: [{ type: "summary_text", text: "loud" }], + summary: [], + content: [{ type: "reasoning_text", text: "loud" }], }); }); diff --git a/tests/adapters/bridge.test.ts b/tests/adapters/bridge.test.ts index e9c2b050b6..f283b77015 100644 --- a/tests/adapters/bridge.test.ts +++ b/tests/adapters/bridge.test.ts @@ -87,27 +87,26 @@ describe("Responses bridge reasoning and usage parity", () => { expect(firstOutputs).toBe(1); }); - test("streaming raw reasoning is routed through the expandable summary channel", async () => { + test("streaming raw reasoning rides the content channel like native gpt-oss", async () => { const frames = await collectSse(bridgeToResponsesSSE(replay([ { type: "reasoning_raw_delta", text: "raw detail" }, { type: "done", usage: { inputTokens: 10, outputTokens: 5, cachedInputTokens: 3, reasoningOutputTokens: 2 } }, ]), "routed/model")); - // Chat-completions providers (DeepSeek-style) deliver thinking as raw - // reasoning_content. Codex renders the expandable reasoning trace from the - // Responses summary channel only, so raw reasoning is routed through the - // summary channel (issue #45) instead of the content channel. - expect(frames.find(f => f.event === "response.reasoning_summary_text.delta")?.data) - .toMatchObject({ summary_index: 0, delta: "raw detail" }); - expect(frames.some(f => f.event === "response.reasoning_text.delta")).toBe(false); + // Raw reasoning_content rides the content channel so Codex applies its own display + // policy: the desktop band shows the "Thinking…" placeholder, and raw text appears + // only when show_raw_agent_reasoning is enabled — never as a fake summary. + expect(frames.find(f => f.event === "response.reasoning_text.delta")?.data) + .toMatchObject({ content_index: 0, delta: "raw detail" }); + expect(frames.some(f => f.event === "response.reasoning_summary_text.delta")).toBe(false); const completed = frames.find(f => f.event === "response.completed")?.data.response as Record; const output = completed.output as Record[]; expect(output[0]).toMatchObject({ type: "reasoning", - summary: [{ type: "summary_text", text: "raw detail" }], + summary: [], + content: [{ type: "reasoning_text", text: "raw detail" }], }); - expect((output[0] as { content?: unknown }).content).toBeUndefined(); expect(completed.usage).toMatchObject({ input_tokens: 10, input_tokens_details: { cached_tokens: 3 }, @@ -502,9 +501,9 @@ describe("Responses bridge reasoning and usage parity", () => { const output = json.output as Record[]; expect(output.map(item => item.type)).toEqual(["reasoning", "message"]); expect(output[0]).toMatchObject({ - summary: [{ type: "summary_text", text: "raw json" }], + summary: [], + content: [{ type: "reasoning_text", text: "raw json" }], }); - expect((output[0] as { content?: unknown }).content).toBeUndefined(); expect(json.usage).toMatchObject({ input_tokens: 6, input_tokens_details: { cached_tokens: 1, cache_write_tokens: 2 }, diff --git a/tests/adapters/google/google-adapter.test.ts b/tests/adapters/google/google-adapter.test.ts index 61e1d15d4b..1dd5a7c97a 100644 --- a/tests/adapters/google/google-adapter.test.ts +++ b/tests/adapters/google/google-adapter.test.ts @@ -1,8 +1,10 @@ import { describe, expect, test } from "bun:test"; import { createGoogleAdapter } from "../../../src/adapters/google"; import { chatCompletionsToResponsesBody } from "../../../src/chat/inbound"; +import { bridgeToResponsesSSE, buildResponseJSON } from "../../../src/bridge"; +import { withTestTranslatorBudget } from "../../helpers/translator-budget"; import { parseRequest } from "../../../src/responses/parser"; -import type { OcxParsedRequest } from "../../../src/types"; +import type { AdapterEvent, OcxParsedRequest } from "../../../src/types"; const provider = { adapter: "google", baseUrl: "https://generativelanguage.googleapis.com", apiKey: "key" }; @@ -583,3 +585,134 @@ describe("google adapter — direct -tiered wire renames", () => { } }); }); + +describe("google adapter — Antigravity thought-text opt-in", () => { + // CCA keeps generating thinking either way (thoughtsTokenCount stays non-zero) but returns + // NO `thought` text unless the request sets generationConfig.thinkingConfig.includeThoughts. + // Probed 2026-09-12: gemini-3.8-flash-high answered with 0 thought parts and 321 thoughts + // tokens, then 358-652 chars of reasoning once the key was present. + const ccaProvider = { + adapter: "google", + googleMode: "cloud-code-assist", + baseUrl: "https://daily-cloudcode-pa.googleapis.com", + apiKey: "key", + project: "proj-123", + } as const; + const optedIn = { ...ccaProvider, showThinkingSummary: true } as const; + + function thoughtParsed(modelId: string, effort?: string, hideThinkingSummary?: boolean): OcxParsedRequest { + return { + modelId, + stream: false, + options: { ...(effort ? { reasoning: effort } : {}), ...(hideThinkingSummary ? { hideThinkingSummary } : {}) }, + context: { messages: [{ role: "user", content: "hi" }], tools: [] }, + } as unknown as OcxParsedRequest; + } + + async function thinkingConfig( + providerConfig: Record, + modelId: string, + effort?: string, + hideThinkingSummary?: boolean, + ): Promise | undefined> { + const { body } = await createGoogleAdapter(providerConfig as never) + .buildRequest(thoughtParsed(modelId, effort, hideThinkingSummary)); + const envelope = JSON.parse(body) as { + request: { generationConfig?: { thinkingConfig?: Record } }; + }; + return envelope.request.generationConfig?.thinkingConfig; + } + + test("asks CCA for thought text on the Gemini wire families", async () => { + // Suffix tier ids deliberately state no level — the suffix IS the effort — so the opt-in + // has to stand on its own for those. + expect(await thinkingConfig(optedIn, "gemini-3.8-flash", "high")).toEqual({ includeThoughts: true }); + expect(await thinkingConfig(optedIn, "gemini-3.8-flash-medium")).toEqual({ includeThoughts: true }); + expect(await thinkingConfig(optedIn, "gemini-3.7-flash", "high")) + .toEqual({ thinkingLevel: "high", includeThoughts: true }); + expect(await thinkingConfig(optedIn, "gemini-3.1-pro", "high")) + .toEqual({ thinkingLevel: "high", includeThoughts: true }); + }); + + test("never sends the flag to models that reject or ignore it", async () => { + // gpt-oss answers 400 INVALID_ARGUMENT with the key present, so it would break the turn. + expect(await thinkingConfig(optedIn, "gpt-oss-120b-medium")).toBeUndefined(); + // Claude-on-CCA accepts the key but returns no thought parts, so it stays off that wire. + expect(await thinkingConfig(optedIn, "claude-sonnet-4-6", "high")).toEqual({ thinkingLevel: "high" }); + }); + + test("a provider without the opt-in keeps the CCA wire unchanged", async () => { + expect(await thinkingConfig(ccaProvider, "gemini-3.8-flash", "high")).toBeUndefined(); + expect(await thinkingConfig(ccaProvider, "gemini-3.7-flash", "high")).toEqual({ thinkingLevel: "high" }); + }); + + test("an explicit client opt-out stops the thought text at the source", async () => { + // Same per-request gate the response path uses: hideThinkingSummary is set for an explicit + // reasoning.summary "none", and paying upstream for text the client refused is waste. + expect(await thinkingConfig(optedIn, "gemini-3.8-flash", "high", true)).toBeUndefined(); + expect(await thinkingConfig(optedIn, "gemini-3.7-flash", "high", true)).toEqual({ thinkingLevel: "high" }); + }); +}); + + +describe("CCA thought summary provenance and replay", () => { + const cca = { adapter: "google", googleMode: "cloud-code-assist", baseUrl: "https://daily-cloudcode-pa.googleapis.com", + apiKey: "fixture-key", project: "fixture-project", showThinkingSummary: true } as const; + const signature = "CiQAx-summary-tool-signature-0123456789abcdef"; + for (const stream of [false, true]) for (const hideThinkingSummary of [false, true]) test(`Gemini signature stream=${stream} hidden=${hideThinkingSummary}`, async () => { + const adapter = withTestTranslatorBudget(createGoogleAdapter(cca)); + const parsed = parsedWith([{ role: "user", content: "lookup" }], [ + { name: "lookup", description: "look up", parameters: { type: "object", properties: {} } }, + ]); + parsed.modelId = "gemini-3.8-flash"; + await adapter.buildRequest(parsed); + const payload = { response: { candidates: [{ content: { parts: [ + { thought: true, text: "Provider summary", thoughtSignature: signature }, + { functionCall: { name: "lookup", args: {} } }, + ] }, finishReason: "STOP" }], usageMetadata: { promptTokenCount: 5, candidatesTokenCount: 2 } } }; + const events: AdapterEvent[] = []; + if (stream) { + for await (const event of adapter.parseStream(new Response(`data: ${JSON.stringify(payload)}\n\n`, + { headers: { "content-type": "text/event-stream" } }))) events.push(event); + } else events.push(...await adapter.parseResponse!(Response.json(payload))); + expect(events[0]).toEqual({ type: "thinking_delta", thinking: "Provider summary" }); + expect(events.some(event => event.type === "thinking_signature")).toBe(false); + const call = events.find(event => event.type === "tool_call_start"); + expect(call?.type === "tool_call_start" && call.providerMetadata?.google?.thoughtSignature).toBe(signature); + expect(events.at(-1)?.type).toBe("done"); + let output: Record; + if (stream) { + async function* replay() { yield* events; } + const text = await new Response(bridgeToResponsesSSE(replay(), parsed.modelId, + undefined, undefined, undefined, undefined, undefined, { hideThinkingSummary })).text(); + const payloads = text.split("\n").filter(line => line.startsWith("data: {")).map(line => JSON.parse(line.slice(6))); + output = payloads.find(frame => frame.type === "response.completed").response; + expect(text.includes("response.reasoning_summary_text.delta")).toBe(!hideThinkingSummary); + } else output = buildResponseJSON(events, parsed.modelId, { hideThinkingSummary }); + expect(JSON.stringify(output).includes("Provider summary")).toBe(!hideThinkingSummary); + if (!Array.isArray(output.output)) throw new Error("missing Responses output"); + const continuation = parseRequest({ model: parsed.modelId, input: [ + ...output.output, { type: "function_call_output", call_id: call && "id" in call ? call.id : "", output: "result" }, + ] }); + const next = JSON.parse((await withTestTranslatorBudget(createGoogleAdapter(cca)).buildRequest(continuation)).body); + const parts = next.request.contents.flatMap((turn: { parts: unknown[] }) => turn.parts); + expect(parts).toContainEqual(expect.objectContaining({ functionCall: expect.objectContaining({ name: "lookup" }), thoughtSignature: signature })); + expect(parts).toContainEqual(expect.objectContaining({ functionResponse: expect.objectContaining({ name: "lookup", response: { result: "result" } }) })); + expect(JSON.stringify(next)).not.toContain("no tool result"); + }); + + test("reused adapter resets Gemini summary provenance for a CCA non-Gemini model", async () => { + const adapter = withTestTranslatorBudget(createGoogleAdapter(cca)); + for (const modelId of ["gemini-3.8-flash", "gpt-oss-120b-medium"]) { + const request = parsedWith([{ role: "user", content: "hi" }]); + request.modelId = modelId; + await adapter.buildRequest(request); + const events = await adapter.parseResponse!(Response.json({ response: { candidates: [{ + content: { parts: [{ thought: true, text: "thinking" }] }, finishReason: "STOP", + }] } })); + expect(events[0]).toEqual(modelId.startsWith("gemini-") + ? { type: "thinking_delta", thinking: "thinking" } + : { type: "reasoning_raw_delta", text: "thinking" }); + } + }); +}); diff --git a/tests/adapters/google/google-wire-compiler.test.ts b/tests/adapters/google/google-wire-compiler.test.ts index 1523d11d85..482aa809e6 100644 --- a/tests/adapters/google/google-wire-compiler.test.ts +++ b/tests/adapters/google/google-wire-compiler.test.ts @@ -132,4 +132,28 @@ describe("Google wire compiler", () => { const repaired = JSON.parse(repairGoogleInvalidRequestBody(body, error)!); expect(repaired.request.generationConfig).toEqual({ maxOutputTokens: 4096 }); }); + + test("keeps the includeThoughts opt-in while still dropping unknown thinking keys", () => { + const withOptIn = compileGoogleWireBody({ + generationConfig: { + thinkingConfig: { includeThoughts: true, thinkingLevel: "max", futureThinkingField: true }, + }, + }); + expect(withOptIn.body.generationConfig).toEqual({ + thinkingConfig: { thinkingLevel: "high", includeThoughts: true }, + }); + + // The flag has to survive on its own too: suffix tier ids deliberately carry no + // thinkingLevel, so an includeThoughts-only config is the whole request. + const optInOnly = compileGoogleWireBody({ + generationConfig: { thinkingConfig: { includeThoughts: true } }, + }); + expect(optInOnly.body.generationConfig).toEqual({ thinkingConfig: { includeThoughts: true } }); + + // Non-boolean / absent values must not invent the key. + const notRequested = compileGoogleWireBody({ + generationConfig: { thinkingConfig: { includeThoughts: "yes", thinkingLevel: "high" } }, + }); + expect(notRequested.body.generationConfig).toEqual({ thinkingConfig: { thinkingLevel: "high" } }); + }); }); diff --git a/tests/adapters/openai/openai-chat-hardening.test.ts b/tests/adapters/openai/openai-chat-hardening.test.ts index b051817469..3a8fe98f03 100644 --- a/tests/adapters/openai/openai-chat-hardening.test.ts +++ b/tests/adapters/openai/openai-chat-hardening.test.ts @@ -140,6 +140,23 @@ describe("AgentRouter openai-chat compatibility", () => { expect(body.messages[0]?.content.map(part => part.text)).toEqual([preamble, "responda somente: OK"]); expect(rawBody.messages[0]?.content).toBe("responda somente: OK"); }); + + test("passthrough chat drops reasoning_effort for an explicitly empty capability ladder", () => { + const rawBody = { + messages: [{ role: "user", content: "hi" }], + reasoning_effort: "xhigh", + }; + const request = buildOpenAIChatPassthroughRequest( + provider({ reasoningEfforts: [] }), + rawBody, + "test-model", + false, + ); + const body = JSON.parse(request.body as string) as Record; + + expect(body).not.toHaveProperty("reasoning_effort"); + expect(rawBody.reasoning_effort).toBe("xhigh"); + }); }); function parsed(): OcxParsedRequest { @@ -1455,3 +1472,25 @@ test("tool-call deltas emit heartbeats so a long buffering phase is not read as expect(visible.at(-1)).toMatchObject({ type: "done" }); }); }); + + +describe("native Chat raw reasoning declarations", () => { + const raw = { messages: [{ role: "user", content: "hello" }], reasoning_effort: "enabled" }; + function wire(overrides: Partial) { + const request = buildOpenAIChatPassthroughRequest({ + adapter: "openai-chat", baseUrl: "https://example.test/v1", ...overrides, + }, raw, "target", false); + return JSON.parse(request.body as string) as Record; + } + test("preserves a nonempty wire-only ladder as unknown", () => { + expect(wire({ reasoningEfforts: ["enabled"] }).reasoning_effort).toBe("enabled"); + expect(wire({ reasoningEfforts: [], modelReasoningEfforts: { target: ["enabled"] } }).reasoning_effort).toBe("enabled"); + }); + test("honors explicit empty model overrides over a provider ladder", () => { + expect(wire({ reasoningEfforts: ["high"], modelReasoningEfforts: { target: [] } }).reasoning_effort).toBeUndefined(); + }); + test("honors noReasoningModels over a nonempty model ladder without mutating input", () => { + expect(wire({ noReasoningModels: ["target"], modelReasoningEfforts: { target: ["high"] } }).reasoning_effort).toBeUndefined(); + expect(raw.reasoning_effort).toBe("enabled"); + }); +}); diff --git a/tests/codex-integration/codex-metadata-integrity.test.ts b/tests/codex-integration/codex-metadata-integrity.test.ts index ee70d4829e..329e90a973 100644 --- a/tests/codex-integration/codex-metadata-integrity.test.ts +++ b/tests/codex-integration/codex-metadata-integrity.test.ts @@ -208,33 +208,115 @@ describe("Codex request transport metadata", () => { expect(new Headers(dropped.headers).get(hintHeader)).toBe("model=gpt-5.6-sol"); }); - test("canonical adapter drops Lite only for the Spark wire model", async () => { + test("canonical adapter disables Spark Lite in HTTP headers and WS metadata without mutating input", async () => { + const { prepareCodexWsRequest } = await import("../../src/server/responses/codex-ws-request"); const adapter = createResponsesPassthroughAdapter({ adapter: "openai-responses", authMode: "forward", baseUrl: "https://chatgpt.com/backend-api/codex", headers: { "X-OpenAI-Internal-Codex-Responses-Lite": "true" }, }); for (const [model, incomingLite, expectedLite] of [ - ["gpt-5.3-codex-spark", "true", null], - ["gpt-5.3-codex-spark", undefined, null], + ["gpt-5.3-codex-spark", "true", "false"], + ["gpt-5.3-codex-spark", "false", "false"], + ["gpt-5.3-codex-spark", undefined, "false"], ["gpt-5.6-sol", "true", "true"], + ["gpt-5.6-sol", "false", "false"], + ["gpt-5.6-sol", undefined, "true"], ] as const) { const parsed = minimalParsed(); parsed.modelId = model; - parsed._rawBody = { model, input: [], stream: true }; + parsed._rawBody = { model, input: [], stream: true, + client_metadata: { [liteKey]: "true", other: "preserved" } }; + const before = JSON.stringify(parsed._rawBody); const incoming = new Headers(); if (incomingLite !== undefined) incoming.set(liteHeader, incomingLite); const request = await adapter.buildRequest(parsed, { headers: incoming, }); expect(new Headers(request.headers).get(liteHeader)).toBe(expectedLite); + const prepared = prepareCodexWsRequest(url, { body: request.body, headers: request.headers })!; + expect(JSON.parse(prepared.frameText).client_metadata).toEqual({ + [liteKey]: expectedLite, other: "preserved", + }); + expect(prepared.httpInit.body).toBe(request.body); + expect(JSON.stringify(parsed._rawBody)).toBe(before); + expect(incoming.get(liteHeader)).toBe(incomingLite ?? null); } const routed = minimalParsed(); routed.modelId = "spark-alias"; routed._rawBody = { model: "gpt-5.3-codex-spark", input: [], stream: true }; const request = await adapter.buildRequest(routed, { headers: new Headers({ [liteHeader]: "true" }) }); - expect(new Headers(request.headers).get(liteHeader)).toBeNull(); + expect(new Headers(request.headers).get(liteHeader)).toBe("false"); + const prepared = prepareCodexWsRequest(url, { body: request.body, headers: request.headers })!; + expect(JSON.parse(prepared.frameText).client_metadata[liteKey]).toBe("false"); + + routed.modelId = "gpt-5.3-codex-spark"; + routed._rawBody = { model: "gpt-5.6-sol", input: [], stream: true }; + const otherWireModel = await adapter.buildRequest(routed, { headers: new Headers({ [liteHeader]: "true" }) }); + expect(new Headers(otherWireModel.headers).get(liteHeader)).toBe("true"); + }); + + test("a Lite-shaped Spark body pins Lite back on, whatever the inherited header said", async () => { + const { prepareCodexWsRequest } = await import("../../src/server/responses/codex-ws-request"); + // The catalog keeps use_responses_lite: true for Spark because it selects tool DELIVERY: + // the client catalog rides `input[].additional_tools`, not top-level `tools`. A forwarded or + // configured `false` must not survive on such a body, or the frame advertises non-Lite while + // the tools exist only in the Lite shape and Spark loses them. + const adapter = createResponsesPassthroughAdapter({ + adapter: "openai-responses", authMode: "forward", baseUrl: "https://chatgpt.com/backend-api/codex", + headers: { "X-OpenAI-Internal-Codex-Responses-Lite": "false" }, + }); + const liteShapedInput = [ + { type: "message", role: "user", content: [{ type: "input_text", text: "hi" }] }, + { type: "additional_tools", tools: [{ type: "function", name: "shell", parameters: {} }] }, + ]; + + for (const incomingLite of ["false", "true", undefined] as const) { + const parsed = minimalParsed(); + parsed.modelId = "gpt-5.3-codex-spark"; + parsed._rawBody = { model: "gpt-5.3-codex-spark", input: liteShapedInput, stream: true, + client_metadata: { [liteKey]: "false", other: "preserved" } }; + const incoming = new Headers(); + if (incomingLite !== undefined) incoming.set(liteHeader, incomingLite); + const request = await adapter.buildRequest(parsed, { headers: incoming }); + expect(new Headers(request.headers).get(liteHeader)).toBe("true"); + const prepared = prepareCodexWsRequest(url, { body: request.body, headers: request.headers })!; + expect(JSON.parse(prepared.frameText).client_metadata).toEqual({ + [liteKey]: "true", other: "preserved", + }); + } + + // An empty group is not a Lite tool surface, so the stream fix still applies. + const toolless = minimalParsed(); + toolless.modelId = "gpt-5.3-codex-spark"; + toolless._rawBody = { model: "gpt-5.3-codex-spark", stream: true, + input: [{ type: "additional_tools", tools: [] }] }; + const downgraded = await adapter.buildRequest(toolless, { headers: new Headers() }); + expect(new Headers(downgraded.headers).get(liteHeader)).toBe("false"); + }); + + test("Spark disables Lite without configured headers and retains malformed-metadata HTTP fallback", async () => { + const { prepareCodexWsRequest } = await import("../../src/server/responses/codex-ws-request"); + const adapter = createResponsesPassthroughAdapter({ + adapter: "openai-responses", authMode: "forward", baseUrl: "https://chatgpt.com/backend-api/codex", + }); + for (const client_metadata of [undefined, {}, null, [], { [liteKey]: true }]) { + const parsed = minimalParsed(); + parsed._rawBody = { model: "gpt-5.3-codex-spark", input: [], stream: true, + ...(client_metadata === undefined ? {} : { client_metadata }) }; + const before = JSON.stringify(parsed._rawBody); + const request = await adapter.buildRequest(parsed, { headers: new Headers() }); + expect(new Headers(request.headers).get(liteHeader)).toBe("false"); + const prepared = prepareCodexWsRequest(url, { body: request.body, headers: request.headers }); + if (client_metadata === undefined || JSON.stringify(client_metadata) === "{}") { + expect(JSON.parse(prepared!.frameText).client_metadata).toEqual({ [liteKey]: "false" }); + } else { + expect(prepared).toBeNull(); + expect(JSON.parse(request.body).client_metadata).toEqual(client_metadata); + } + expect(JSON.stringify(parsed._rawBody)).toBe(before); + } }); test("noncanonical adapters neither forward caller Lite nor synthesize a routing hint", async () => { @@ -243,7 +325,10 @@ describe("Codex request transport metadata", () => { adapter: "openai-responses", authMode, baseUrl: "https://gateway.example/v1", headers: { [hintHeader]: "operator-owned" }, }); - const request = await adapter.buildRequest(minimalParsed(), { + const parsed = minimalParsed(); + parsed.modelId = "gpt-5.3-codex-spark"; + parsed._rawBody = { model: parsed.modelId, input: [] }; + const request = await adapter.buildRequest(parsed, { headers: new Headers({ [liteHeader]: "true", [hintHeader]: "caller-owned" }), }); expect(new Headers(request.headers).has(liteHeader)).toBe(false); @@ -363,3 +448,60 @@ describe("Codex request transport metadata", () => { expect(headers.get(hintHeader)).toBe("model=gpt-5.6-luna"); }); }); + + +describe("Spark Lite follows serialized model and surviving tool shape", () => { + const liteHeader = "x-openai-internal-codex-responses-lite"; + const liteKey = "ws_request_header_x_openai_internal_codex_responses_lite"; + test("bracket normalization and both alias directions use the wire model", async () => { + const adapter = createResponsesPassthroughAdapter({ adapter: "openai-responses", authMode: "forward", + baseUrl: "https://chatgpt.com/backend-api/codex", modelSuffixBracketStrip: true }); + for (const [selector, model, expectedModel, expectedLite] of [ + ["alias", "gpt-5.3-codex-spark[1m]", "gpt-5.3-codex-spark", "false"], + ["gpt-5.3-codex-spark", "gpt-5.6-sol", "gpt-5.6-sol", "true"], + ]) { + const parsed = minimalParsed(); + parsed.modelId = selector; + parsed._rawBody = { model, input: [] }; + const before = JSON.stringify(parsed._rawBody); + const built = await adapter.buildRequest(parsed, { headers: new Headers({ [liteHeader]: "true" }) }); + expect(JSON.parse(built.body).model).toBe(expectedModel); + expect(new Headers(built.headers).get(liteHeader)).toBe(expectedLite); + expect(new Headers(built.headers).get("x-codex-routing-hint")).toContain(`model=${expectedModel}`); + expect(JSON.stringify(parsed._rawBody)).toBe(before); + } + }); + + for (const [name, inputTools, topTools, expectedTools, expectedLite] of [ + ["filtered empty", [{ type: "tool_search" }], undefined, [], "false"], + ["reserved functions", [{ type: "namespace", name: "functions", tools: [{ type: "function", name: "lookup", parameters: { type: "object" } }] }], undefined, + [{ type: "namespace", name: "functions", tools: [{ type: "function", name: "lookup", parameters: { type: "object" } }] }], "true"], + ["top-level only", undefined, [{ type: "function", name: "lookup", parameters: { type: "object" } }], undefined, "false"], + ] as const) test(`post-transform body: ${name}`, async () => { + const { prepareCodexWsRequest } = await import("../../src/server/responses/codex-ws-request"); + const adapter = createResponsesPassthroughAdapter({ adapter: "openai-responses", authMode: "forward", + baseUrl: "https://chatgpt.com/backend-api/codex", headers: { [liteHeader]: expectedLite === "true" ? "false" : "true" } }); + const parsed = minimalParsed(); + parsed.modelId = "gpt-5.3-codex-spark"; + parsed._rawBody = { model: parsed.modelId, input: inputTools ? [{ type: "additional_tools", tools: inputTools }] : [], + ...(topTools ? { tools: topTools } : {}) }; + const built = await adapter.buildRequest(parsed); + const body = JSON.parse(built.body); + expect(body.input.find((item: { type: string }) => item.type === "additional_tools")?.tools).toEqual(expectedTools); + if (topTools) expect(body.tools).toEqual(topTools); + expect(new Headers(built.headers).get(liteHeader)).toBe(expectedLite); + const prepared = prepareCodexWsRequest("https://chatgpt.com/backend-api/codex/responses", { body: built.body, headers: built.headers }); + expect(JSON.parse(prepared!.frameText).client_metadata[liteKey]).toBe(expectedLite); + }); + + test("noncanonical static Lite remains operator-owned", async () => { + for (const authMode of ["key", "forward"] as const) { + const adapter = createResponsesPassthroughAdapter({ adapter: "openai-responses", authMode, + baseUrl: "https://gateway.example/v1", headers: { [liteHeader]: "operator-owned" } }); + const parsed = minimalParsed(); + parsed._rawBody = { model: "gpt-5.3-codex-spark", input: [] }; + const built = await adapter.buildRequest(parsed, { headers: new Headers({ [liteHeader]: "true" }) }); + expect(new Headers(built.headers).get(liteHeader)).toBe("operator-owned"); + } + }); +}); diff --git a/tests/codex-integration/combos.test.ts b/tests/codex-integration/combos.test.ts index 76e4f26ca9..0bc1230e0f 100644 --- a/tests/codex-integration/combos.test.ts +++ b/tests/codex-integration/combos.test.ts @@ -297,22 +297,93 @@ describe("combo request cloning", () => { expect(concrete.input).not.toBe(raw.input); }); - test("combo default respects client-owned ignored reasoning values", () => { + test("combo target capability strips unsupported client reasoning controls", () => { expect(concreteComboRequestBody({ model: "combo/x", reasoning: null }, target, "high", []).reasoning).toBeNull(); expect(concreteComboRequestBody( { model: "combo/x", reasoning: { effort: "" } }, target, "high", [], - ).reasoning).toEqual({ effort: "" }); + ).reasoning).toBeUndefined(); expect(concreteComboRequestBody( { model: "combo/x", reasoning: { effort: "banana" } }, target, "high", [], - ).reasoning).toEqual({ effort: "banana" }); + ).reasoning).toBeUndefined(); expect(concreteComboRequestBody( { model: "combo/x", reasoning: { effort: null } }, target, "high", [], - ).reasoning).toEqual({ effort: null }); + ).reasoning).toBeUndefined(); expect(concreteComboRequestBody( { model: "combo/x", reasoning: { summary: "concise" } }, target, "high", ["high"], ).reasoning).toEqual({ summary: "concise", effort: "high" }); }); + test("adaptive normalization strips unsupported controls for an unknown target while preserving summary", () => { + const raw = { + model: "combo/x", + input: "hi", + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }; + const concrete = concreteComboRequestBody(raw, target, null, undefined, "adaptive"); + + expect(concrete).toEqual({ + model: "a/m1", + input: "hi", + reasoning: { summary: "concise" }, + }); + expect(raw).toEqual({ + model: "combo/x", + input: "hi", + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }); + }); + + test("strict normalization preserves reasoning controls for an unknown target", () => { + const raw = { + model: "combo/x", + input: "hi", + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }; + const concrete = concreteComboRequestBody(raw, target, null, undefined, "strict"); + + expect(concrete).toEqual({ + model: "a/m1", + input: "hi", + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }); + }); + + test("explicit empty ladder strips unsupported controls while preserving reasoning summary", () => { + const concrete = concreteComboRequestBody({ + model: "combo/x", + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }, target, "high", []); + + expect(concrete).toEqual({ + model: "a/m1", + reasoning: { summary: "concise" }, + }); + }); + + test("adaptive normalization preserves xhigh for a known reasoning ladder", () => { + const concrete = concreteComboRequestBody({ + model: "combo/x", + reasoning: { effort: "xhigh", summary: "concise" }, + }, target, null, ["low", "medium", "high", "xhigh"], "adaptive"); + + expect(concrete.reasoning).toEqual({ effort: "xhigh", summary: "concise" }); + }); + test("omits combo defaults for unset, no-reasoning, and unknown target capabilities", () => { expect(concreteComboRequestBody({ model: "combo/x" }, target, null, ["high"]).reasoning).toBeUndefined(); // An explicitly empty ladder is how a no-reasoning model is expressed. diff --git a/tests/fixtures/test-layout-expected.json b/tests/fixtures/test-layout-expected.json index 63d04e6aa3..a041e9ca4d 100644 --- a/tests/fixtures/test-layout-expected.json +++ b/tests/fixtures/test-layout-expected.json @@ -376,6 +376,7 @@ "context-compat.test.ts": "codex-integration", "context-history-ownership.test.ts": "server", "context-history.test.ts": "server", + "context-window-seed-repair.test.ts": "providers", "continuation-dedup.test.ts": "responses", "core-lab-boundary.test.ts": "lab", "cost-cap-unknown-evidence.test.ts": "usage", @@ -699,7 +700,6 @@ "model-pinned-effort.test.ts": "codex-integration", "model-presets.test.ts": "providers", "model-rename-migration.test.ts": "providers", - "context-window-seed-repair.test.ts": "providers", "model-selection-guidance.test.ts": "cli", "model-visibility-management-api.test.ts": "codex-integration", "models-feedback-callback.test.ts": "gui", @@ -786,9 +786,9 @@ "openai-chat-hardening.test.ts": "adapters/openai", "openai-chat-invalid-tool-call-diagnostics.test.ts": "adapters/openai", "openai-chat-model-suffix.test.ts": "adapters/openai", - "openai-chat-path-override.test.ts": "adapters/openai", "openai-chat-native-policy.test.ts": "adapters/openai", "openai-chat-parallel-stream.test.ts": "adapters/openai", + "openai-chat-path-override.test.ts": "adapters/openai", "openai-chat-system-order.test.ts": "adapters/openai", "openai-chat-tool-result-images.test.ts": "adapters/openai", "openai-chat-url.test.ts": "adapters/openai", @@ -824,6 +824,8 @@ "pi-path-contract.test.ts": "clients", "pinned-http.test.ts": "lib", "pinned-https-get.test.ts": "images", + "plaintext-v2-agent-messages-server.test.ts": "server", + "plaintext-v2-agent-messages.test.ts": "responses", "plan-video.test.ts": "videos", "plan.test.ts": "images", "policy-execution.test.ts": "routing", @@ -923,6 +925,7 @@ "responses-account-label.test.ts": "responses", "responses-compaction-routing.test.ts": "responses", "responses-compaction.test.ts": "responses", + "responses-console-go-upload-retry.test.ts": "responses", "responses-context-overflow.test.ts": "responses", "responses-custom-tool-guidance.test.ts": "responses", "responses-custom-tool-repair.test.ts": "responses", @@ -946,10 +949,10 @@ "responses-pool-401-refresh.test.ts": "responses", "responses-pool-refresh-attribution.test.ts": "responses", "responses-reasoning-summary-passthrough.test.ts": "responses", - "responses-reasoning-summary-rewrite.test.ts": "responses", "responses-routed-web-search-fields.test.ts": "responses", "responses-self-named-namespace-scrub.test.ts": "responses", "responses-shadow-intercept.test.ts": "responses", + "responses-show-thinking-summary.test.ts": "responses", "responses-snapshot-repair-server.test.ts": "responses", "responses-snapshot-repair.test.ts": "responses", "responses-state-write-amplification.test.ts": "responses", @@ -1038,7 +1041,6 @@ "sidecar-settings-web-search-stream.test.ts": "vision", "sidecar-tracker.test.ts": "vision", "skill-ocx.test.ts": "ci-workflows", - "structure-ssot.test.ts": "ci-workflows", "slug-codec.test.ts": "codex-integration", "sponsor-presets.test.ts": "providers", "sse-client-frame-bounds.test.ts": "responses", @@ -1070,6 +1072,7 @@ "storage-worker-teardown-isolate.test.ts": "storage", "stream-aborted-marker.test.ts": "server", "strict-semver.test.ts": "lib", + "structure-ssot.test.ts": "ci-workflows", "subagent-context-staleness.test.ts": "routing", "subagent-defaults.test.ts": "routing", "subagent-fallback-handle-responses.test.ts": "routing", diff --git a/tests/providers/opencode-go-luna-wire.test.ts b/tests/providers/opencode-go-luna-wire.test.ts index c0afb15553..7866ec9415 100644 --- a/tests/providers/opencode-go-luna-wire.test.ts +++ b/tests/providers/opencode-go-luna-wire.test.ts @@ -163,19 +163,17 @@ describe("OpenCode Go stateless reasoning and continuation routes", () => { const initial = { type: "message", role: "user", content: [{ type: "input_text", text: "Run probe" }] }; const first = await drive({ input: [initial] }); expect(first.document.output[0]).toEqual(reasoning[0]); - expect(first.document.output[1]).toEqual(continuation.summary === "auto" ? { - type: "reasoning", id: `rs_${prefix}_content`, status: "completed", summary: [{ type: "summary_text", text: "Visible thinking" }], - } : reasoning[1]); + // The passthrough keeps native content-channel reasoning in both display modes. + expect(first.document.output[1]).toEqual(reasoning[1]); expect(first.document.output[2]).toEqual(reasoning[2]); expect(first.document.output[3]).toMatchObject(call); expect(first.document.output[4]).toEqual(priorMessage); if (streaming) { - const channel = continuation.summary === "auto" ? "reasoning_summary_text" : "reasoning_text"; - expect(first.text).toContain(`"type":"response.${channel}.delta"`); + expect(first.text).toContain('"type":"response.reasoning_text.delta"'); } const result = { type: "function_call_output", call_id: call.call_id, output: "probe succeeded" }; // Echo exactly the client-visible history through handleResponses. An upstream-shape - // cache would prepend it again after the content-to-summary rewrite (F1). + // cache would prepend it again (F1). const nextBody = { input: continuation.fullHistory ? [initial, ...first.document.output, result] : [result], previous_response_id: first.document.id, store: true, @@ -207,9 +205,8 @@ describe("OpenCode Go stateless reasoning and continuation routes", () => { expect(replay.filter(item => item.type === "reasoning")).toHaveLength(3); expect(replay).toContainEqual(expect.objectContaining({ type: "reasoning", encrypted_content: blob })); expect(JSON.stringify(replay)).toContain("Already summarized"); - if (continuation.summary === "auto") expect(replay).toContainEqual(expect.objectContaining({ - type: "reasoning", summary: [{ type: "summary_text", text: "Visible thinking" }], - })); + // Replay sanitation strips reasoning content in both display modes (F1), so the + // visible "Visible thinking" trace does not re-enter the upstream history. expect(JSON.stringify(replay)).not.toContain("no tool result was recorded"); }); } diff --git a/tests/providers/opencode-zen-rate-limit.test.ts b/tests/providers/opencode-zen-rate-limit.test.ts index 8751082c2a..e8e90d9416 100644 --- a/tests/providers/opencode-zen-rate-limit.test.ts +++ b/tests/providers/opencode-zen-rate-limit.test.ts @@ -6,6 +6,7 @@ import { enrichOpenCodeZenFreeTierMessage, enrichOpenCodeZenRateLimitMessage, enrichOpenCodeZenUpstreamMessage, + isTransientConsoleGoUploadRejection, isOpenCodeZenFreeTierLockIn, isOpenCodeZenRateLimitProvider, } from "../../src/providers/opencode-zen-rate-limit"; @@ -230,3 +231,79 @@ describe("opencode-free keyless tier lock-in (#4121)", () => { expect(lockedOut).not.toContain(OPENCODE_ZEN_OBSERVED_RPM_HINT); }); }); + +describe("Console Go transient upload refusal", () => { + // The observed Go-route refusal, byte-for-byte as Console serves it. + const GO_MESSAGE = "Error from provider (Console Go): Upstream request failed: [invalid_request_error] Invalid upload request."; + // The Zen key route names the same gateway without the Go suffix. + const ZEN_MESSAGE = "Error from provider (Console): Upstream request failed: [invalid_request_error] Invalid upload request."; + const envelope = (message: string) => JSON.stringify({ model: "muse-spark-1.3-contributor", error: { param: null, type: "invalid_request_error", message } }); + const GO_ROUTE = { outboundUrl: "https://opencode.ai/zen/go/v1/responses" }; + const ZEN_ROUTE = { outboundUrl: "https://opencode.ai/zen/v1/responses" }; + + test("accepts the canonical refusal on both canonical Console routes", () => { + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), ...GO_ROUTE })).toBe(true); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(ZEN_MESSAGE), ...ZEN_ROUTE })).toBe(true); + // A custom row pointed at the same destination is still Console: the base URL decides. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), outboundUrl: "https://opencode.ai/zen/go/v1/responses" })).toBe(true); + }); + + test("rejects the refusal text from a non-Console route", () => { + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), outboundUrl: "https://api.deepseek.com/v1/responses" })).toBe(false); + // opencode.ai without the /zen segment is not the Console gateway. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), outboundUrl: "https://opencode.ai/v1/responses" })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE) })).toBe(false); + }); + + test("requires the effective canonical HTTPS endpoint", () => { + const userInfoUrl = new URL("https://opencode.ai/zen/go/v1/responses"); + userInfoUrl.username = "fixture-user"; + userInfoUrl.password = "fixture-password"; + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), outboundUrl: userInfoUrl.href })).toBe(false); + for (const outboundUrl of [ + "http://opencode.ai/zen/go/v1/responses", + "https://opencode.ai:8443/zen/go/v1/responses", + "https://opencode.ai/zen/go/v1/responses?tenant=fixture", + "https://opencode.ai/zen/go/v1/responses#fragment", + "https://opencode.ai.evil.test/zen/go/v1/responses", + "https://opencode.ai/zen-other/v1/responses", + "https://opencode.ai/zen/go/v1/models", + ]) expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), outboundUrl })).toBe(false); + }); + + test("rejects noncanonical envelopes, suffixes, and other statuses", () => { + // A bare string is not the structured envelope. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: JSON.stringify({ error: "Invalid upload request." }), ...GO_ROUTE })).toBe(false); + // A suffix means the gateway said something else; do not guess. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE + " Please retry."), ...GO_ROUTE })).toBe(false); + // Different status: only the gateway 400 is the flap. + expect(isTransientConsoleGoUploadRejection({ status: 500, errorBody: envelope(GO_MESSAGE), ...GO_ROUTE })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: undefined, ...GO_ROUTE })).toBe(false); + // Other 400s on the same wire are verdicts on the request, not flaps. + expect(isTransientConsoleGoUploadRejection({ + status: 400, + errorBody: envelope("Error from provider (Console Go): Upstream request failed: [invalid_request_error] reasoning_effort max requires an active Muse Code subscription for model muse-spark-1.3-contributor."), + ...GO_ROUTE, + })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ + status: 400, + errorBody: JSON.stringify({ type: "error", error: { type: "MissingSessionID", message: "Request is missing x-opencode-session" } }), + ...GO_ROUTE, + })).toBe(false); + }); + + test("rejects partial envelopes and padded messages", () => { + const withError = (error: unknown) => JSON.stringify({ model: "muse-spark-1.3-contributor", error }); + // type carries the refusal identity; a partial envelope is a different error. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ param: null, message: GO_MESSAGE }), ...GO_ROUTE })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ param: null, type: "server_error", message: GO_MESSAGE }), ...GO_ROUTE })).toBe(false); + // param must be present and null. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ type: "invalid_request_error", message: GO_MESSAGE }), ...GO_ROUTE })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ param: "input", type: "invalid_request_error", message: GO_MESSAGE }), ...GO_ROUTE })).toBe(false); + // Padding means the gateway wrapped or appended something. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ param: null, type: "invalid_request_error", message: " " + GO_MESSAGE }), ...GO_ROUTE })).toBe(false); + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: withError({ param: null, type: "invalid_request_error", message: GO_MESSAGE + "\n" }), ...GO_ROUTE })).toBe(false); + // The exact canonical envelope still matches. + expect(isTransientConsoleGoUploadRejection({ status: 400, errorBody: envelope(GO_MESSAGE), ...GO_ROUTE })).toBe(true); + }); +}); diff --git a/tests/responses/openai-responses-passthrough.test.ts b/tests/responses/openai-responses-passthrough.test.ts index d5a49fd39e..f6b411ec4e 100644 --- a/tests/responses/openai-responses-passthrough.test.ts +++ b/tests/responses/openai-responses-passthrough.test.ts @@ -335,6 +335,55 @@ test("canonical forward providers normalize trailing slashes and let the pool ov expect(request.headers["chatgpt-account-id"]).toBe("runtime-account"); }); +test("noncanonical Responses preserves provider-owned safety-buffering hints", async () => { + const upstream = [ + 'event: response.created\ndata: {"type":"response.created","response":{"id":"resp_custom"},"safety_buffering":{"provider_owned":true}}\n\n', + 'event: response.metadata\ndata: {"type":"response.metadata","metadata":{"type":"safety_buffering","provider_owned":true}}\n\n', + 'event: response.completed\ndata: {"type":"response.completed","response":{"id":"resp_custom","status":"completed","output":[]}}\n\n', + "data: [DONE]\n\n", + ].join(""); + const savedFetch = globalThis.fetch; + globalThis.fetch = (async () => new Response(upstream, { headers: { + "content-type": "text/event-stream", + "x-codex-safety-buffering-enabled": "provider-owned", + "x-codex-safety-buffering-faster-model": "provider-model", + } })) as typeof fetch; + try { + for (const providerConfig of [ + { + adapter: "openai-responses", + baseUrl: "https://fixture.test/v1", + authMode: "key" as const, + apiKey: "fixture-key", + }, + { + adapter: "openai-responses", + baseUrl: "https://fixture.test/v1", + authMode: "forward" as const, + headers: { authorization: "Bearer provider-static" }, + }, + ]) { + const config = { + port: 0, + defaultProvider: "fixture", + dropCodexSafetyBuffering: true, + providers: { fixture: providerConfig }, + } as OcxConfig; + const response = await handleResponses(new Request("http://localhost/v1/responses", { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ model: "fixture/model", stream: true, input: "ping" }), + }), config, { model: "", provider: "" }); + + expect(response.headers.get("x-codex-safety-buffering-enabled")).toBe("provider-owned"); + expect(response.headers.get("x-codex-safety-buffering-faster-model")).toBe("provider-model"); + expect(await response.text()).toBe(upstream); + } + } finally { + globalThis.fetch = savedFetch; + } +}); + test("noncanonical pool-required providers use only their configured static credentials", () => { const adapter = createResponsesPassthroughAdapter({ adapter: "openai-responses", @@ -4533,3 +4582,31 @@ describe("raw usage passthrough on the forward path (#41980 parity, #37138 adjac } }); }); + + +test("canonical Responses hint suppression is opt-in at the request boundary", async () => { + const savedFetch = globalThis.fetch; + globalThis.fetch = (async () => new Response([ + 'data: {"type":"response.created","response":{"id":"resp_hint"},"safety_buffering":true}\n\n', + 'data: {"type":"response.metadata","metadata":{"type":"safety_buffering"}}\n\n', + 'data: {"type":"response.completed","response":{"id":"resp_hint","status":"completed","output":[]}}\n\n', + ].join(""), { headers: { "content-type": "text/event-stream", + "x-codex-safety-buffering-enabled": "true", "x-codex-safety-buffering-faster-model": "fixture-model", "x-codex-turn-id": "fixture-turn" } })) as typeof fetch; + try { + for (const dropCodexSafetyBuffering of [undefined, false, true]) { + const config = { port: 0, dropCodexSafetyBuffering, providers: { openai: { + ...provider, codexAccountMode: "direct", upstreamWebsocket: false, + } } } as OcxConfig; + const response = await handleResponses(new Request("http://localhost/v1/responses", { + method: "POST", headers: { "content-type": "application/json", authorization: "Bearer fixture-forward-token" }, + body: JSON.stringify({ model: "openai/gpt-5.6-sol", input: "ping", stream: true }), + }), config, { model: "", provider: "" }); + expect(response.status).toBe(200); + expect(response.headers.has("x-codex-safety-buffering-enabled")).toBe(dropCodexSafetyBuffering !== true); + expect(response.headers.get("x-codex-turn-id")).toBe("fixture-turn"); + const text = await response.text(); + expect(text.includes("safety_buffering")).toBe(dropCodexSafetyBuffering !== true); + expect(text).toContain("response.completed"); + } + } finally { globalThis.fetch = savedFetch; } +}); diff --git a/tests/responses/passthrough-abort.test.ts b/tests/responses/passthrough-abort.test.ts index 46100c6902..913da3eadd 100644 --- a/tests/responses/passthrough-abort.test.ts +++ b/tests/responses/passthrough-abort.test.ts @@ -79,7 +79,7 @@ describe("passthrough relayWithAbort (RC2, passthrough path)", () => { expect(sseBranch).toContain("rewriteBlocks: clientBlockRewrite"); // Elsewhere the failed-tail relay converts mid-stream resets into a clean response.failed. expect(sseBranch).toMatch( - /relaySseWithFailedTail\(\s*rewrittenBody,\s*upstream,\s*reason\s*=>\s*\{\s*responseCompletionCancelled\s*=\s*true;\s*clientGone\.abort\(reason\);\s*\},\s*\{\s*upstreamError:\s*logCtx\.upstreamError\s*\},\s*\)/, + /relaySseWithFailedTail\(\s*rewrittenBody,\s*upstream,\s*reason\s*=>\s*\{\s*responseCompletionCancelled\s*=\s*true;\s*clientGone\.abort\(reason\);\s*\},\s*\{\s*upstreamError:\s*logCtx\.upstreamError,\s*terminalBoundary:\s*codexSafetyBufferingOptions\s*\},\s*\)/, ); expect(sseBranch).toContain("new Response(clientBody"); expect(sseBranch).toContain("markNativePassthroughSseResponse"); diff --git a/tests/responses/passthrough-headers.test.ts b/tests/responses/passthrough-headers.test.ts index 8018e77906..2cd5261609 100644 --- a/tests/responses/passthrough-headers.test.ts +++ b/tests/responses/passthrough-headers.test.ts @@ -1,5 +1,6 @@ import { describe, expect, test } from "bun:test"; -import { sanitizePassthroughHeaders } from "../../src/server"; +import { codexSafetyBufferingFilterOptions, sanitizePassthroughHeaders } from "../../src/server"; +import { createSseTerminalOutputBoundary } from "../../src/server/relay"; describe("passthrough header sanitization (RC5 / F4)", () => { test("content-type: text/event-stream survives sanitization", () => { @@ -43,3 +44,75 @@ describe("passthrough header sanitization (RC5 / F4)", () => { expect(sanitized.get("content-type")).toBe("text/event-stream"); }); }); + +describe("codex safety-buffering hint headers", () => { + const upstream = () => new Headers({ + "content-type": "text/event-stream", + "x-codex-safety-buffering-enabled": "true", + "X-Codex-Safety-Buffering-Faster-Model": "gpt-5.6-luna", + "x-codex-primary-used-percent": "12", + "openai-model": "gpt-6-astra", + }); + + test("forwarded verbatim by default and when the option is off", () => { + for (const options of [undefined, {}, { dropCodexSafetyBuffering: false }]) { + const sanitized = sanitizePassthroughHeaders(upstream(), options); + expect(sanitized.get("x-codex-safety-buffering-enabled")).toBe("true"); + expect(sanitized.get("x-codex-safety-buffering-faster-model")).toBe("gpt-5.6-luna"); + } + }); + + test("dropped case-insensitively when opted in, other x-codex headers survive", () => { + const sanitized = sanitizePassthroughHeaders(upstream(), { dropCodexSafetyBuffering: true }); + expect(sanitized.has("x-codex-safety-buffering-enabled")).toBe(false); + expect(sanitized.has("x-codex-safety-buffering-faster-model")).toBe(false); + expect(sanitized.get("x-codex-primary-used-percent")).toBe("12"); + expect(sanitized.get("openai-model")).toBe("gpt-6-astra"); + expect(sanitized.get("content-type")).toBe("text/event-stream"); + }); + + test("codexSafetyBufferingFilterOptions only enables the drop on an explicit true", () => { + expect(codexSafetyBufferingFilterOptions({})).toEqual({ dropCodexSafetyBuffering: false }); + expect(codexSafetyBufferingFilterOptions({ dropCodexSafetyBuffering: false })) + .toEqual({ dropCodexSafetyBuffering: false }); + expect(codexSafetyBufferingFilterOptions({ dropCodexSafetyBuffering: true })) + .toEqual({ dropCodexSafetyBuffering: true }); + }); +}); + +describe("Codex safety-buffering SSE hints at the client output boundary", () => { + const encoder = new TextEncoder(); + const decoder = new TextDecoder(); + const frames = [ + 'event: response.created\ndata: {"type":"response.created","response":{"id":"resp_1"},"safety_buffering":{"retry_model":"gpt-5.6-luna"}}\n\n', + 'event: response.metadata\ndata: {"type":"response.metadata","metadata":{"type":"safety_buffering","retry_model":"gpt-5.6-luna"}}\n\n', + 'event: response.metadata\ndata: {"type":"response.metadata","metadata":{"type":"other","turn":1}}\n\n', + 'event: response.output_text.delta\ndata: {"type":"response.output_text.delta","delta":"hi"}\n\n', + 'event: response.completed\ndata: {"type":"response.completed","response":{"id":"resp_1","status":"completed"}}\n\n', + ]; + const relay = (options?: { dropCodexSafetyBuffering?: boolean }): string => { + const boundary = createSseTerminalOutputBoundary(options); + let out = ""; + for (const frame of frames) out += decoder.decode(boundary.feed(encoder.encode(frame))); + out += decoder.decode(boundary.finish()); + boundary.dispose(); + return out; + }; + + test("relayed verbatim by default and when the option is off", () => { + for (const options of [undefined, {}, { dropCodexSafetyBuffering: false }]) { + expect(relay(options)).toBe(frames.join("")); + } + }); + + test("metadata event dropped and field stripped when opted in, other events untouched", () => { + const out = relay({ dropCodexSafetyBuffering: true }); + expect(out).not.toContain("safety_buffering"); + expect(out).not.toContain("gpt-5.6-luna"); + expect(out).toContain('data: {"type":"response.created","response":{"id":"resp_1"}}'); + expect(out).toContain(frames[2]); + expect(out).toContain(frames[3]); + expect(out).toContain(frames[4]); + expect(out.match(/^event: /gm)).toHaveLength(4); + }); +}); diff --git a/tests/responses/plaintext-v2-agent-messages.test.ts b/tests/responses/plaintext-v2-agent-messages.test.ts new file mode 100644 index 0000000000..525d23defa --- /dev/null +++ b/tests/responses/plaintext-v2-agent-messages.test.ts @@ -0,0 +1,909 @@ +import { describe, expect, test } from "bun:test"; +import { createResponsesPassthroughAdapter as createResponsesPassthroughAdapterProduction } from "../../src/adapters/openai-responses"; +import { + PlaintextV2AgentMessageRestoreOverflowError, + createPlaintextV2AgentMessageCallRestoreRewrite, + PLAINTEXT_V2_COLLABORATION_NAMESPACE, + preparePlaintextV2AgentMessages, + restorePlaintextV2AgentMessageCalls, + restorePlaintextV2AgentMessageCallsInJson, + restorePlaintextV2AgentMessageCallsInJsonResult, + shouldPreparePlaintextV2AgentMessages, +} from "../../src/responses/plaintext-v2-agent-messages"; +import { withTestTranslatorBudget } from "../helpers/translator-budget"; + +const createResponsesPassthroughAdapter = (...args: Parameters) => + withTestTranslatorBudget(createResponsesPassthroughAdapterProduction(...args)); + +function collaborationTool(name: string, encrypted: boolean = true): Record { + return { + type: "function", + name, + parameters: { + type: "object", + properties: { + message: { + type: "string", + encrypted, + const: { encrypted: true }, + }, + encrypted: { type: "boolean" }, + }, + required: ["message"], + }, + }; +} + +describe("plaintext v2 agent message request preparation", () => { + test("strips only the three message markers and aliases collaboration catalogs", () => { + const body = { + model: "gpt-5.6-sol", + tools: [{ + type: "namespace", + name: "collaboration", + tools: [ + collaborationTool("spawn_agent"), + collaborationTool("send_message"), + collaborationTool("followup_task", false), + collaborationTool("wait_agent"), + ], + }], + input: [{ + type: "additional_tools", + tools: [{ + type: "namespace", + name: "collaboration", + tools: [collaborationTool("followup_task")], + }], + }], + }; + const before = structuredClone(body); + + const prepared = preparePlaintextV2AgentMessages(body); + const result = prepared.body as typeof body; + const namespace = result.tools[0] as typeof body.tools[0]; + const spawn = namespace.tools[0] as ReturnType; + const send = namespace.tools[1] as ReturnType; + const followup = namespace.tools[2] as ReturnType; + const wait = namespace.tools[3] as ReturnType; + const additionalNamespace = result.input[0].tools[0] as { + name: string; + tools: Array>; + }; + const additional = additionalNamespace.tools[0]!; + const message = (tool: Record) => ( + ((tool.parameters as Record).properties as Record>).message + ); + + expect(prepared.namespaceAliased).toBe(true); + expect([...prepared.toolNames].sort()).toEqual([ + "followup_task", + "send_message", + "spawn_agent", + "wait_agent", + ]); + expect(namespace.name).toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(additionalNamespace.name).toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(spawn.name).toBe("start_delegated_task"); + expect(send.name).toBe("deliver_delegated_message"); + expect(followup.name).toBe("continue_delegated_task"); + expect(additional.name).toBe("continue_delegated_task"); + expect(wait.name).toBe("wait_agent"); + expect(message(spawn).encrypted).toBeUndefined(); + expect(message(send).encrypted).toBeUndefined(); + expect(message(additional).encrypted).toBeUndefined(); + expect(message(followup).encrypted).toBe(false); + expect(message(wait).encrypted).toBe(true); + expect(message(spawn).const).toEqual({ encrypted: true }); + expect(((spawn.parameters as Record).properties as Record).encrypted) + .toEqual({ type: "boolean" }); + expect(body).toEqual(before); + }); + + test("does not reinterpret a flat same-named function as the Codex v2 catalog", () => { + const body = { tools: [collaborationTool("spawn_agent")] }; + const prepared = preparePlaintextV2AgentMessages(body); + const tool = (prepared.body as typeof body).tools[0] as Record; + const message = ((tool.parameters as Record).properties as Record>).message; + + expect(message.encrypted).toBe(true); + expect(prepared.body).toBe(body); + expect(prepared.namespaceAliased).toBe(false); + }); + + test("does not strip same-named tools from another namespace", () => { + const body = { + tools: [ + { + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }, + { + type: "namespace", + name: "private_mail", + tools: [collaborationTool("send_message")], + }, + ], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + const privateTool = (prepared.body as typeof body).tools[1]!.tools[0] as Record; + const privateMessage = ((privateTool.parameters as Record).properties as Record>).message; + + expect(privateMessage.encrypted).toBe(true); + expect(privateTool.name).toBe("send_message"); + }); + + test("does not strip an independent flat tool beside a collaboration namespace", () => { + const body = { + tools: [ + { + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }, + collaborationTool("send_message"), + ], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + const flatTool = (prepared.body as typeof body).tools[1] as Record; + const flatMessage = ((flatTool.parameters as Record).properties as Record>).message; + expect(flatMessage.encrypted).toBe(true); + expect(flatTool.name).toBe("send_message"); + }); + + test("aliases selectors and replayed calls with the rewritten collaboration catalog", () => { + const replayedCall = { + type: "function_call", + call_id: "call-old", + namespace: "collaboration", + name: "spawn_agent", + arguments: "{}", + }; + const replayedOutput = { + type: "function_call_output", + call_id: "call-old", + output: { namespace: "collaboration" }, + }; + const body = { + tools: [{ + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent"), collaborationTool("send_message")], + }], + tool_choice: { + type: "allowed_tools", + mode: "required", + tools: [ + { type: "function", namespace: "collaboration", name: "send_message" }, + { type: "function", namespace: "private_mail", name: "send_message" }, + ], + }, + input: [replayedCall, replayedOutput], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + const result = prepared.body as typeof body; + + expect(result.tools[0]!.name).toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(result.tool_choice.tools[0]!.namespace).toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(result.tool_choice.tools[0]!.name).toBe("deliver_delegated_message"); + expect(result.tool_choice.tools[1]!.namespace).toBe("private_mail"); + expect(result.input[0]!.namespace).toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(result.input[0]!.name).toBe("start_delegated_task"); + expect(result.input[1]).toEqual(replayedOutput); + }); + + test("aliases a forced collaboration tool choice", () => { + const body = { + tools: [{ + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent"), collaborationTool("followup_task")], + }], + tool_choice: { type: "function", namespace: "collaboration", name: "followup_task" }, + }; + + const prepared = preparePlaintextV2AgentMessages(body); + expect((prepared.body as typeof body).tool_choice.namespace) + .toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect((prepared.body as typeof body).tool_choice.name).toBe("continue_delegated_task"); + }); + + test("aliases both supported qualified-name forms using declared child names", () => { + const body = { + tools: [{ + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent"), collaborationTool("send_message")], + }], + tool_choice: { type: "function", name: "collaboration__spawn_agent" }, + input: [{ + type: "function_call", + call_id: "call-send", + name: "collaboration.send_message", + arguments: "{}", + }], + }; + + const result = preparePlaintextV2AgentMessages(body).body as typeof body; + expect(result.tool_choice.name) + .toBe(`${PLAINTEXT_V2_COLLABORATION_NAMESPACE}__start_delegated_task`); + expect(result.input[0]!.name) + .toBe(`${PLAINTEXT_V2_COLLABORATION_NAMESPACE}.deliver_delegated_message`); + }); + + test("aliases every duplicate collaboration declaration in one request", () => { + const body = { + tools: [ + { + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }, + { + type: "namespace", + name: "collaboration", + tools: [{ type: "function", name: "wait_agent", parameters: { type: "object" } }], + }, + ], + tool_choice: { type: "function", namespace: "collaboration", name: "wait_agent" }, + }; + + const result = preparePlaintextV2AgentMessages(body).body as typeof body; + expect(result.tools.map(tool => tool.name)).toEqual([ + PLAINTEXT_V2_COLLABORATION_NAMESPACE, + PLAINTEXT_V2_COLLABORATION_NAMESPACE, + ]); + expect(result.tool_choice.namespace).toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + }); + + test("does not reinterpret an independent flattened-looking tool name", () => { + const body = { + tools: [ + { + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }, + { type: "function", name: "collaboration__audit", parameters: { type: "object" } }, + ], + tool_choice: { + type: "allowed_tools", + tools: [{ type: "function", name: "collaboration__audit" }], + }, + input: [{ + type: "function_call", + call_id: "call-audit", + name: "collaboration__audit", + arguments: "{}", + }], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + const result = prepared.body as typeof body; + expect(result.tool_choice.tools[0]!.name).toBe("collaboration__audit"); + expect(result.input[0]!.name).toBe("collaboration__audit"); + }); + + test("skips aliasing when a flat declaration collides with a namespace child", () => { + const body = { + tools: [ + { + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }, + { type: "function", name: "collaboration__spawn_agent", parameters: { type: "object" } }, + ], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + expect(prepared.body).toBe(body); + expect(prepared.namespaceAliased).toBe(false); + }); + + test("aliases a recognized collaboration catalog even when the marker is already absent", () => { + const body = { + tools: [{ + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent", false)], + }], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + expect((prepared.body as typeof body).tools[0]!.name) + .toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect((prepared.body as typeof body).tools[0]!.tools[0]!.name) + .toBe("start_delegated_task"); + expect(prepared.namespaceAliased).toBe(true); + }); + + test("leaves the whole request untouched when a collaboration child already uses a private tool alias", () => { + const body = { + tools: [{ + type: "namespace", + name: "collaboration", + tools: [ + collaborationTool("spawn_agent"), + collaborationTool("start_delegated_task"), + ], + }], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + expect(prepared.body).toBe(body); + expect(prepared.namespaceAliased).toBe(false); + }); + + test("leaves the whole request untouched when a top-level tool uses a fixed message alias", () => { + const body = { + tools: [ + { + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }, + { type: "function", name: "start_delegated_task", parameters: { type: "object" } }, + ], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + expect(prepared.body).toBe(body); + expect(prepared.namespaceAliased).toBe(false); + }); + + test("leaves the whole request untouched when another namespace uses a fixed message alias", () => { + const body = { + tools: [{ + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }], + input: [{ + type: "additional_tools", + tools: [{ + type: "namespace", + name: "foreign", + tools: [{ type: "function", name: "deliver_delegated_message", parameters: { type: "object" } }], + }], + }], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + expect(prepared.body).toBe(body); + expect(prepared.namespaceAliased).toBe(false); + }); + + test("scans tool-search declarations and foreign references for fixed message aliases", () => { + const catalog = [{ + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }]; + const cases = [ + { + tools: catalog, + input: [{ + type: "tool_search_output", + tools: [{ type: "function", name: "continue_delegated_task", parameters: { type: "object" } }], + }], + }, + { + tools: catalog, + tool_choice: { type: "function", name: "start_delegated_task" }, + }, + { + tools: catalog, + input: [{ + type: "function_call", + namespace: "foreign", + name: "deliver_delegated_message", + call_id: "foreign-call", + arguments: "{}", + }], + }, + ]; + + for (const body of cases) { + const prepared = preparePlaintextV2AgentMessages(body); + expect(prepared.body).toBe(body); + expect(prepared.namespaceAliased).toBe(false); + } + }); + + test("leaves the whole request untouched when replay history already uses a private tool alias", () => { + const body = { + tools: [{ + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }], + input: [{ + type: "function_call", + call_id: "call-private-name", + namespace: "collaboration", + name: "start_delegated_task", + arguments: "{}", + }], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + expect(prepared.body).toBe(body); + expect(prepared.namespaceAliased).toBe(false); + }); + + test("leaves the whole request untouched when the private alias already exists", () => { + const body = { + tools: [ + { type: "namespace", name: PLAINTEXT_V2_COLLABORATION_NAMESPACE, tools: [] }, + { + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }, + ], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + expect(prepared.body).toBe(body); + expect(prepared.namespaceAliased).toBe(false); + expect(JSON.stringify(prepared.body)).toContain('"encrypted":true'); + }); + + test("leaves the request untouched when replay history already uses the private alias", () => { + const body = { + tools: [{ + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }], + input: [{ + type: "function_call", + call_id: "call-private", + namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + name: "audit", + arguments: "{}", + }], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + expect(prepared.body).toBe(body); + expect(prepared.namespaceAliased).toBe(false); + }); + + test("leaves the request untouched when tool-search history declares the private alias", () => { + const body = { + tools: [{ + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }], + input: [{ + type: "tool_search_output", + tools: [{ + type: "namespace", + name: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + tools: [], + }], + }], + }; + + const prepared = preparePlaintextV2AgentMessages(body); + expect(prepared.body).toBe(body); + expect(prepared.namespaceAliased).toBe(false); + }); + + test("does not treat a collaboration namespace nested under another namespace as Codex v2", () => { + const depth = 20_000; + const collaboration = { + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + } as Record; + let root: Record = collaboration; + for (let index = 0; index < depth; index++) { + root = { type: "namespace", name: `nest-${index}`, tools: [root] }; + } + + const prepared = preparePlaintextV2AgentMessages({ tools: [root] }); + expect(prepared.namespaceAliased).toBe(false); + }); +}); + +describe("plaintext v2 agent message response restoration", () => { + const declaredToolNames = new Set(["spawn_agent", "send_message"]); + + test("restores tool identities and preserves the plaintext proof and user data", () => { + const payload = JSON.stringify({ + type: "response.completed", + response: { + tool_choice: { + type: "function", + namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + name: "start_delegated_task", + }, + tools: [{ + type: "namespace", + name: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + tools: [{ type: "function", name: "start_delegated_task" }], + }], + output: [ + { + type: "function_call", + namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + name: "start_delegated_task", + arguments: JSON.stringify({ + message: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + }), + encrypted_function_args: [], + }, + { + type: "function_call", + name: `${PLAINTEXT_V2_COLLABORATION_NAMESPACE}__deliver_delegated_message`, + arguments: "{}", + encrypted_function_args: [], + }, + { + type: "function_call_output", + output: { namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE }, + }, + ], + }, + }); + + const restored = JSON.parse( + restorePlaintextV2AgentMessageCallsInJson(payload, declaredToolNames), + ) as { + response: { + tool_choice: Record; + tools: Array>; + output: Array>; + }; + }; + const [namespaced, flattened, toolOutput] = restored.response.output; + + expect(namespaced!.namespace).toBe("collaboration"); + expect(namespaced!.name).toBe("spawn_agent"); + expect(namespaced!.encrypted_function_args).toEqual([]); + expect(JSON.parse(namespaced!.arguments as string).message) + .toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(flattened!.name).toBe("send_message"); + expect(flattened!.encrypted_function_args).toEqual([]); + expect(toolOutput!.output).toEqual({ namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE }); + expect(restored.response.tool_choice.namespace).toBe("collaboration"); + expect(restored.response.tool_choice.name).toBe("spawn_agent"); + expect(restored.response.tools[0]!.name).toBe("collaboration"); + expect((restored.response.tools[0]!.tools as Array>)[0]!.name) + .toBe("spawn_agent"); + }); + + test("restores the identity on streamed function-call argument completion", () => { + const payload = JSON.stringify({ + type: "response.function_call_arguments.done", + item_id: "fc-spawn", + namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + name: `${PLAINTEXT_V2_COLLABORATION_NAMESPACE}__start_delegated_task`, + arguments: JSON.stringify({ message: PLAINTEXT_V2_COLLABORATION_NAMESPACE }), + encrypted_function_args: [], + }); + + const restored = JSON.parse( + restorePlaintextV2AgentMessageCallsInJson(payload, declaredToolNames), + ) as Record; + expect(restored.namespace).toBe("collaboration"); + expect(restored.name).toBe("spawn_agent"); + expect(JSON.parse(restored.arguments as string).message) + .toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(restored.encrypted_function_args).toEqual([]); + }); + + test("rejects invalid JSON but preserves valid payloads without aliases", () => { + expect(() => restorePlaintextV2AgentMessageCallsInJson("not json", declaredToolNames)).toThrow(PlaintextV2AgentMessageRestoreOverflowError); + const payload = '{"type":"response.completed"}'; + expect(restorePlaintextV2AgentMessageCallsInJson(payload, declaredToolNames)).toBe(payload); + }); + + test("restores an unqualified private tool alias in streamed JSON", () => { + const payload = JSON.stringify({ + type: "function_call", + name: "start_delegated_task", + arguments: JSON.stringify({ message: "plain assignment" }), + encrypted_function_args: [], + }); + + const restored = JSON.parse( + restorePlaintextV2AgentMessageCallsInJson(payload, declaredToolNames), + ) as Record; + expect(restored.name).toBe("spawn_agent"); + expect(restored.encrypted_function_args).toEqual([]); + }); + + test("does not restore a fixed alias authenticated by another namespace", () => { + const payload = JSON.stringify({ + type: "function_call", + namespace: "foreign", + name: "start_delegated_task", + arguments: "{}", + }); + + expect(restorePlaintextV2AgentMessageCallsInJson(payload, declaredToolNames)).toBe(payload); + }); + + test("restores only message aliases that this request actually generated", () => { + const prepared = preparePlaintextV2AgentMessages({ + tools: [{ + type: "namespace", + name: "collaboration", + tools: [ + collaborationTool("spawn_agent"), + { type: "custom", name: "send_message" }, + ], + }], + }); + const payload = JSON.stringify({ + type: "function_call", + name: "deliver_delegated_message", + arguments: "{}", + }); + + expect([...prepared.aliasedAgentMessageToolNames]).toEqual(["spawn_agent"]); + expect(() => restorePlaintextV2AgentMessageCallsInJson( + payload, + prepared.toolNames, + prepared.aliasedAgentMessageToolNames, + )).toThrow(PlaintextV2AgentMessageRestoreOverflowError); + }); + + test("leaves foreign calls and nested extension metadata untouched", () => { + const extensionCall = { + type: "function_call", + namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + name: "start_delegated_task", + }; + const payload = JSON.stringify({ + type: "response.completed", + response: { + output: [ + { + type: "function_call", + namespace: "foreign", + name: "audit", + arguments: "{}", + }, + { + type: "function_call", + namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + name: "start_delegated_task", + arguments: "{}", + }, + ], + metadata: { + nested: extensionCall, + values: Array.from({ length: 20_000 }, (_, index) => index), + }, + }, + }); + + const restored = JSON.parse( + restorePlaintextV2AgentMessageCallsInJson(payload, declaredToolNames), + ) as { + response: { + output: Array>; + metadata: { nested: Record; values: number[] }; + }; + }; + + expect(restored.response.output[0]!.namespace).toBe("foreign"); + expect(restored.response.output[1]!.namespace).toBe("collaboration"); + expect(restored.response.metadata.nested).toEqual(extensionCall); + expect(restored.response.metadata.values).toHaveLength(20_000); + }); + + test("fails closed when known identity arrays exceed the work limit", () => { + const value = { + output: Array.from({ length: 10_001 }, () => ({ + type: "function_call", + namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + name: "start_delegated_task", + })), + }; + const payload = JSON.stringify(value); + + expect(restorePlaintextV2AgentMessageCallsInJsonResult(payload, declaredToolNames)).toEqual({ + value: payload, + changed: false, + overflowed: true, + }); + expect(() => restorePlaintextV2AgentMessageCallsInJson(payload, declaredToolNames)) + .toThrow(PlaintextV2AgentMessageRestoreOverflowError); + expect(restorePlaintextV2AgentMessageCalls(value, declaredToolNames)).toEqual({ + value, + changed: false, + overflowed: true, + }); + }); +}); + +describe("plaintext v2 agent message route policy", () => { + test("requires an explicit opt-in, Responses inbound, canonical ChatGPT, and a v2 catalog", () => { + const requestBody = { tools: [{ type: "namespace", name: "collaboration", tools: [collaborationTool("spawn_agent")] }] }; + const baseline = { + enabled: true, + inboundWire: "responses", + canonicalChatGpt: true, + requestBody, + }; + expect(shouldPreparePlaintextV2AgentMessages(baseline)).toBe(true); + const additionalOnly = { input: [{ type: "additional_tools", tools: requestBody.tools }] }; + expect(shouldPreparePlaintextV2AgentMessages({ ...baseline, requestBody: additionalOnly })).toBe(false); + expect(preparePlaintextV2AgentMessages(additionalOnly).namespaceAliased).toBe(false); + expect(shouldPreparePlaintextV2AgentMessages({ ...baseline, enabled: false })).toBe(false); + expect(shouldPreparePlaintextV2AgentMessages({ ...baseline, inboundWire: "anthropic" })).toBe(false); + expect(shouldPreparePlaintextV2AgentMessages({ ...baseline, canonicalChatGpt: false })).toBe(false); + expect(shouldPreparePlaintextV2AgentMessages({ + ...baseline, + requestBody: { tools: [collaborationTool("spawn_agent")] }, + })).toBe(false); + }); +}); + +describe("canonical Responses adapter plaintext v2 integration", () => { + const provider = { + adapter: "openai-responses", + baseUrl: "https://chatgpt.com/backend-api/codex", + authMode: "forward" as const, + }; + + function build(enabled: boolean) { + const rawBody = { + model: "gpt-5.6-sol", + store: false, + stream: true, + input: "delegate", + tools: [{ + type: "namespace", + name: "collaboration", + tools: [collaborationTool("spawn_agent")], + }], + }; + const request = createResponsesPassthroughAdapter(provider).buildRequest({ + modelId: "gpt-5.6-sol", + context: { messages: [] }, + stream: true, + options: {}, + _rawBody: rawBody, + ...(enabled ? { _plaintextV2AgentMessages: true } : {}), + }, { headers: new Headers({ authorization: "Bearer test" }) }); + return { request, rawBody }; + } + + test("changes only the serialized upstream body when enabled", () => { + const { request, rawBody } = build(true); + const sent = JSON.parse(request.body) as typeof rawBody; + const namespace = sent.tools[0]!; + const spawn = namespace.tools[0] as Record; + const message = ((spawn.parameters as Record).properties as Record>).message; + + expect(namespace.name).toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(spawn.name).toBe("start_delegated_task"); + expect(message.encrypted).toBeUndefined(); + expect([...(request.plaintextV2AgentMessageToolNames ?? [])]).toEqual(["spawn_agent"]); + expect([...(request.plaintextV2AgentMessageAliasedToolNames ?? [])]).toEqual(["spawn_agent"]); + expect(rawBody.tools[0]!.name).toBe("collaboration"); + expect(JSON.stringify(rawBody)).toContain('"encrypted":true'); + }); + + test("keeps the upstream collaboration schema unchanged when disabled", () => { + const { request, rawBody } = build(false); + const sent = JSON.parse(request.body); + + expect(sent).toEqual(rawBody); + expect(request.plaintextV2AgentMessageToolNames).toBeUndefined(); + expect(request.plaintextV2AgentMessageAliasedToolNames).toBeUndefined(); + }); +}); + +describe("plaintext V2 refusal boundaries", () => { + const names = new Set(["spawn_agent", "send_message"]); + test("rejects malformed and non-object JSON even without literal aliases", () => { + for (const payload of ["{broken", "[]", "null", '"opaque"']) { + expect(restorePlaintextV2AgentMessageCallsInJsonResult(payload, names).overflowed).toBe(true); + } + }); + test("rejects unknown private identities instead of leaking an unrestorable alias", () => { + expect(() => restorePlaintextV2AgentMessageCallsInJson(JSON.stringify({ + output: [{ type: "function_call", namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, name: "undeclared", arguments: "{}" }], + }), names)).toThrow(PlaintextV2AgentMessageRestoreOverflowError); + }); + test("qualified aliases in nested foreign catalogs prevent any request rewrite", () => { + for (const separator of ["__", "."]) { + const body = { tools: [ + { type: "namespace", name: "collaboration", tools: [collaborationTool("spawn_agent")] }, + { type: "namespace", name: "foreign", tools: [{ type: "function", name: `other${separator}start_delegated_task` }] }, + ] }; + expect(preparePlaintextV2AgentMessages(body).body).toBe(body); + } + }); + test("preserves qualified names explicitly authenticated by a foreign namespace", () => { + const payload = JSON.stringify({ type: "function_call", namespace: "foreign", name: `${PLAINTEXT_V2_COLLABORATION_NAMESPACE}__start_delegated_task`, arguments: "{}" }); + expect(restorePlaintextV2AgentMessageCallsInJson(payload, names)).toBe(payload); + }); +}); + +test("plaintext restoration refuses malformed identity arrays", () => { + for (const value of [{ output: { name: "start_delegated_task" } }, { tools: "collaboration-optimize" }]) { + expect(restorePlaintextV2AgentMessageCallsInJsonResult(JSON.stringify(value), new Set(["spawn_agent"])).overflowed).toBe(true); + } +}); + + +test("restores namespace selectors and allowed namespace choices", () => { + const names = new Set(["spawn_agent"]); + for (const choice of [ + { type: "namespace", name: PLAINTEXT_V2_COLLABORATION_NAMESPACE }, + { type: "allowed_tools", tools: [{ type: "namespace", name: PLAINTEXT_V2_COLLABORATION_NAMESPACE }] }, + ]) { + const restored = restorePlaintextV2AgentMessageCallsInJson(JSON.stringify({ tool_choice: choice }), names); + expect(restored).not.toContain(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(restored).toContain('"collaboration"'); + } +}); + +test("sparse argument events inherit only a compatible existing binding", () => { + const rewrite = createPlaintextV2AgentMessageCallRestoreRewrite(new Set(["spawn_agent", "send_message"])); + rewrite(JSON.stringify({ type: "response.output_item.added", output_index: 0, item: { + type: "function_call", id: "fc1", namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, name: "start_delegated_task", + } })); + expect(JSON.parse(rewrite(JSON.stringify({ type: "response.function_call_arguments.done", item_id: "fc1", name: "start_delegated_task", arguments: "{}" }))).name).toBe("spawn_agent"); + expect(() => rewrite(JSON.stringify({ type: "response.function_call_arguments.done", item_id: "fc1", name: "deliver_delegated_message", arguments: "{}" }))).toThrow(PlaintextV2AgentMessageRestoreOverflowError); +}); + +test("request replay preserves explicit foreign namespace identity", () => { + const replay = { type: "function_call", namespace: "foreign", name: "collaboration__spawn_agent", arguments: "{}" }; + const body = { tools: [{ type: "namespace", name: "collaboration", tools: [collaborationTool("spawn_agent")] }], input: [replay] }; + const prepared = preparePlaintextV2AgentMessages(body); + expect(prepared.namespaceAliased).toBe(true); + expect((prepared.body as typeof body).input[0]).toBe(replay); +}); + + +test("namespace refinement follows every bound coordinate", () => { + const rewrite = createPlaintextV2AgentMessageCallRestoreRewrite(new Set(["spawn_agent"])); + rewrite(JSON.stringify({ type: "response.output_item.added", output_index: 0, item: { + type: "function_call", id: "fc1", call_id: "c1", name: "start_delegated_task", + } })); + rewrite(JSON.stringify({ type: "response.function_call_arguments.done", item_id: "fc1", namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, name: "start_delegated_task", arguments: "{}" })); + expect(() => rewrite(JSON.stringify({ type: "response.completed", response: { output: [{ + type: "function_call", call_id: "c1", namespace: "foreign", name: "spawn_agent", arguments: "{}", + }] } }))).toThrow(PlaintextV2AgentMessageRestoreOverflowError); +}); + + +test("every generated alias spelling restores the exact native dispatch pair", () => { + const names = new Set(["spawn_agent"]); + for (const name of ["start_delegated_task", `${PLAINTEXT_V2_COLLABORATION_NAMESPACE}__start_delegated_task`, `${PLAINTEXT_V2_COLLABORATION_NAMESPACE}.start_delegated_task`]) { + for (const [namespace, marker] of [undefined, null].flatMap(namespace => [undefined, [], ["message"]].map(marker => [namespace, marker] as const))) { + const value = { type: "function_call", name, namespace, arguments: "{}", ...(marker === undefined ? {} : { encrypted_function_args: marker }) }; + const restored = JSON.parse(restorePlaintextV2AgentMessageCallsInJson(JSON.stringify(value), names)); + expect(restored).toMatchObject({ namespace: "collaboration", name: "spawn_agent" }); + expect(restored.encrypted_function_args).toEqual(marker); + } + } + const snapshot = JSON.parse(restorePlaintextV2AgentMessageCallsInJson(JSON.stringify({ + tools: [{ type: "namespace", name: PLAINTEXT_V2_COLLABORATION_NAMESPACE, tools: [{ type: "function", name: "start_delegated_task", parameters: {} }] }], + tool_choice: { type: "function", name: "start_delegated_task" }, + }), names)); + expect(snapshot.tools[0].tools[0]).toEqual({ type: "function", name: "spawn_agent", parameters: {} }); + expect(snapshot.tool_choice).toEqual({ type: "function", namespace: "collaboration", name: "spawn_agent" }); +}); + +test("malformed namespace types cannot bypass private identity restoration", () => { + for (const namespace of [false, 0, {}, []]) { + const payload = JSON.stringify({ type: "function_call", namespace, name: "start_delegated_task", arguments: "{}" }); + expect(restorePlaintextV2AgentMessageCallsInJsonResult(payload, new Set(["spawn_agent"])).overflowed).toBe(true); + } +}); diff --git a/tests/responses/responses-console-go-upload-retry.test.ts b/tests/responses/responses-console-go-upload-retry.test.ts new file mode 100644 index 0000000000..4421848bf2 --- /dev/null +++ b/tests/responses/responses-console-go-upload-retry.test.ts @@ -0,0 +1,263 @@ +import { afterEach, beforeEach, describe, expect, spyOn, test } from "bun:test"; +import { mkdtempSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { markResponseNonReplayable } from "../../src/lib/upstream-retry"; +import { handleResponses } from "../../src/server/responses/core"; +import type { RequestLogContext } from "../../src/server/request-log"; +import type { OcxConfig } from "../../src/types"; +import { removeTreeWithRetry } from "../helpers/remove-tree"; + +const originalFetch = globalThis.fetch; +const originalOpenCodexHome = process.env.OPENCODEX_HOME; + +/** The exact Console Go rejection for a payload it accepts moments later. */ +const UPLOAD_REFUSAL = JSON.stringify({ + model: "muse-spark-1.3-contributor", + error: { + param: null, + type: "invalid_request_error", + message: "Error from provider (Console Go): Upstream request failed: [invalid_request_error] Invalid upload request.", + }, +}); + +/** A deterministic 400 on the same wire: a verdict on the request, never a flap. */ +const EFFORT_REFUSAL = JSON.stringify({ + model: "muse-spark-1.3-contributor", + error: { + param: "reasoning.effort", + type: "invalid_request_error", + message: "Error from provider (Console Go): Upstream request failed: [invalid_request_error] reasoning_effort max requires an active Muse Code subscription for model muse-spark-1.3-contributor.", + }, +}); + +let testDir = ""; + +beforeEach(() => { + testDir = mkdtempSync(join(tmpdir(), "ocx-console-go-upload-retry-")); + process.env.OPENCODEX_HOME = testDir; +}); + +afterEach(() => { + globalThis.fetch = originalFetch; + if (originalOpenCodexHome === undefined) delete process.env.OPENCODEX_HOME; + else process.env.OPENCODEX_HOME = originalOpenCodexHome; + removeTreeWithRetry(testDir); +}); + +function config(): OcxConfig { + return { + defaultProvider: "go", + providers: { + go: { + adapter: "openai-responses", + baseUrl: "https://opencode.ai/zen/go/v1", + authMode: "key", + apiKey: "go-test-key", + }, + other: { + adapter: "openai-responses", + baseUrl: "https://other.example.test/v1", + authMode: "key", + apiKey: "other-test-key", + }, + }, + } as OcxConfig; +} + +function request(stream = false, provider = "go"): Request { + return new Request("http://localhost/v1/responses", { + method: "POST", + headers: { + "content-type": "application/json", + "session_id": "thread-console-go-upload-retry", + }, + body: JSON.stringify({ + model: provider + "/muse-spark-1.3-contributor", + stream, + store: false, + input: [ + { type: "message", role: "user", content: [{ type: "input_text", text: "hi" }] }, + ], + }), + }); +} + +function refusal(status = 400, body = UPLOAD_REFUSAL): Response { + return new Response(body, { status, headers: { "content-type": "application/json" } }); +} + +function success(id: string): Response { + return Response.json({ id, object: "response", status: "completed", model: "muse-spark-1.3-contributor", output: [] }); +} + +describe("Console Go transient upload refusal recovery", () => { + test("replays the refusal once and serves the retry with a byte-identical body", async () => { + const outbound: string[] = []; + globalThis.fetch = (async (_input: RequestInfo | URL, init?: RequestInit) => { + outbound.push(String(init?.body)); + return outbound.length === 1 ? refusal() : success("resp-upload-retry-recovered"); + }) as typeof fetch; + const logCtx: RequestLogContext = { model: "", provider: "" }; + + const response = await handleResponses(request(), config(), logCtx); + + expect(response.status).toBe(200); + expect(outbound).toHaveLength(2); + // The replay must preserve the exact serialized request. + expect(outbound[1]).toBe(outbound[0]); + expect(logCtx.activeAttempt?.recoveryKinds).toEqual(["console-go-upload-retry"]); + }); + + test("does not replay a different 400 from the same wire", async () => { + let sends = 0; + globalThis.fetch = (async () => { + sends += 1; + return refusal(400, EFFORT_REFUSAL); + }) as typeof fetch; + + const response = await handleResponses(request(), config(), { model: "", provider: "" }); + + expect(response.status).toBe(400); + expect(sends).toBe(1); + }); + + test("keeps a repeated refusal visible after the single bounded replay", async () => { + let sends = 0; + globalThis.fetch = (async () => { + sends += 1; + return refusal(); + }) as typeof fetch; + const logCtx: RequestLogContext = { model: "", provider: "" }; + + const response = await handleResponses(request(), config(), logCtx); + + expect(response.status).toBe(400); + expect(sends).toBe(2); + expect(logCtx.activeAttempt?.recoveryKinds).toEqual(["console-go-upload-retry"]); + }); + + test("does not replay the same refusal text from a non-Console provider", async () => { + let sends = 0; + globalThis.fetch = (async () => { + sends += 1; + return refusal(); + }) as typeof fetch; + + const response = await handleResponses(request(false, "other"), config(), { model: "", provider: "" }); + + expect(response.status).toBe(400); + expect(sends).toBe(1); + }); +}); + + +describe("Console destination and translated recovery controls", () => { + test("a query-bearing Console destination never authorizes another POST", async () => { + const cfg = config(); + cfg.providers.go!.baseUrl = "https://opencode.ai/zen/go/v1?tenant=fixture"; + const outbound: string[] = []; + globalThis.fetch = (async input => { outbound.push(String(input)); return refusal(); }) as typeof fetch; + const response = await handleResponses(request(), cfg, { model: "", provider: "" }); + expect(response.status).toBe(400); + expect(outbound).toHaveLength(1); + expect(new URL(outbound[0]!).search).toBe("?tenant=fixture"); + }); + + test("a canonical row name cannot authorize a noncanonical generation path", async () => { + const cfg = config(); + cfg.providers["opencode-go"] = { + ...cfg.providers.go!, + responsesPath: "/unrelated", + chatCompletionsPath: "/unrelated", + }; + const outbound: string[] = []; + globalThis.fetch = (async input => { outbound.push(String(input)); return refusal(); }) as typeof fetch; + const response = await handleResponses(request(false, "opencode-go"), cfg, { model: "", provider: "" }); + expect(response.status).toBe(400); + // Canonical row names normalize their base URL. A configured send path survives + // that normalization and reaches the effective-destination recovery gate. + expect(outbound).toEqual(["https://opencode.ai/zen/go/v1/unrelated"]); + expect(await response.text()).toContain("Invalid upload request."); + }); + + test("normalization to the canonical endpoint keeps its bounded recovery", async () => { + const cfg = config(); + cfg.providers["opencode-go"] = { ...cfg.providers.go!, baseUrl: "https://other.example.test/v1" }; + const outbound: string[] = []; + globalThis.fetch = (async input => { outbound.push(String(input)); return refusal(); }) as typeof fetch; + const response = await handleResponses(request(false, "opencode-go"), cfg, { model: "", provider: "" }); + expect(response.status).toBe(400); + expect(outbound).toEqual([ + "https://opencode.ai/zen/go/v1/responses", + "https://opencode.ai/zen/go/v1/responses", + ]); + expect(await response.text()).toContain("Invalid upload request."); + }); + + for (const adapter of ["openai-responses", "openai-chat"] as const) { + for (const stream of [false, true]) { + test(`${adapter} stream=${stream} replays identical bytes once`, async () => { + const cfg = config(); + cfg.providers.go!.adapter = adapter; + const outbound: string[] = []; + globalThis.fetch = (async (_url: RequestInfo | URL, init?: RequestInit) => { + outbound.push(String(init?.body)); + if (outbound.length === 1) return refusal(); + if (adapter === "openai-responses") { + const completed = { id: "resp_fixture", object: "response", status: "completed", output: [] }; + return stream ? new Response(`event: response.completed\ndata: ${JSON.stringify({ type: "response.completed", response: completed })}\n\n`, { headers: { "content-type": "text/event-stream" } }) : Response.json(completed); + } + if (!stream) return Response.json({ id: "chat_fixture", object: "chat.completion", choices: [{ index: 0, message: { role: "assistant", content: "answer" }, finish_reason: "stop" }] }); + const chunk = { id: "chat_fixture", object: "chat.completion.chunk", choices: [{ index: 0, delta: { content: "answer" }, finish_reason: "stop" }] }; + return new Response(`data: ${JSON.stringify(chunk)}\n\ndata: [DONE]\n\n`, { headers: { "content-type": "text/event-stream" } }); + }) as typeof fetch; + const logCtx: RequestLogContext = { model: "", provider: "" }; + const response = await handleResponses(request(stream), cfg, logCtx); + const body = await response.text(); + expect(response.status).toBe(200); + expect(outbound).toHaveLength(2); + expect(outbound[1]).toBe(outbound[0]); + expect(logCtx.activeAttempt?.recoveryKinds).toEqual(["console-go-upload-retry"]); + expect(body).toContain("completed"); + }); + } + test(`${adapter} abort during backoff sends no replay`, async () => { + const cfg = config(); cfg.providers.go!.adapter = adapter; + const controller = new AbortController(); + let sends = 0; + globalThis.fetch = (async () => { sends++; return refusal(); }) as typeof fetch; + const originalTimeout = globalThis.setTimeout; + const spy = spyOn(globalThis, "setTimeout").mockImplementation(((handler: TimerHandler, ms?: number, ...args: unknown[]) => { + if (ms === 800) queueMicrotask(() => controller.abort()); + return originalTimeout(handler, ms, ...args); + }) as typeof setTimeout); + try { + const response = await handleResponses(request(), cfg, { model: "", provider: "" }, { abortSignal: controller.signal }); + expect(controller.signal.aborted).toBe(true); + expect(response.status).toBe(499); + expect(sends).toBe(1); + } finally { spy.mockRestore(); } + }); + } +}); + + +describe("Console nonreplayable response boundary", () => { + for (const adapter of ["openai-responses", "openai-chat"] as const) { + test(`${adapter} does not replay a marked response`, async () => { + const cfg = config(); cfg.providers.go!.adapter = adapter; + let sends = 0; + globalThis.fetch = (async () => { + sends++; + const response = refusal(); + markResponseNonReplayable(response); + return response; + }) as typeof fetch; + const response = await handleResponses(request(), cfg, { model: "", provider: "" }); + expect(response.status).toBe(400); + expect(sends).toBe(1); + expect(await response.text()).toContain("Invalid upload request."); + }); + } +}); diff --git a/tests/responses/responses-reasoning-summary-passthrough.test.ts b/tests/responses/responses-reasoning-summary-passthrough.test.ts index 5912221354..ab6e099bcc 100644 --- a/tests/responses/responses-reasoning-summary-passthrough.test.ts +++ b/tests/responses/responses-reasoning-summary-passthrough.test.ts @@ -6,10 +6,12 @@ import type { OcxConfig } from "../../src/types"; /** * The passthrough relay for DeepSeek's native /responses endpoint emits - * content-channel reasoning (reasoning_text.delta + content items). The - * summary-channel rewrite must engage only when the client did NOT ask for - * hidden thinking (hideThinkingSummary) - otherwise a client that asked to - * hide reasoning would get it surfaced as visible summary output. + * content-channel reasoning (reasoning_text.delta + content items) in BOTH + * display modes: Codex applies its own raw-reasoning display policy, so a + * requested summary must not rewrite the native passthrough shape either. + * Hidden thinking (hideThinkingSummary) and visible summary get the same + * content-channel passthrough; the hidden variant additionally arrives as an + * envelope-only item upstream when the adapter layer handles suppression. */ function deepseekSeed() { @@ -86,15 +88,16 @@ describe("passthrough reasoning summary rewrite honors hideThinkingSummary", () expect(text).toContain('"content":[{"type":"reasoning_text","text":"think"}]'); }); - test("SSE: requested summary routes raw reasoning through the summary channel", async () => { + test("SSE: requested summary keeps the native content-channel passthrough", async () => { const response = await runHandleResponses( { model: "deepseek-v4-flash", input: "ping", stream: true, reasoning: { effort: "max", summary: "detailed" } }, SSE_UPSTREAM_FRAMES.join(""), "text/event-stream", ); const text = await response.text(); - expect(text).toContain("response.reasoning_summary_text.delta"); - expect(text).toContain('"summary":[{"type":"summary_text","text":"think"}]'); + expect(text).toContain("response.reasoning_text.delta"); + expect(text).not.toContain("response.reasoning_summary_text.delta"); + expect(text).toContain('"content":[{"type":"reasoning_text","text":"think"}]'); }); test("bounded JSON: hidden thinking keeps the content shape", async () => { @@ -108,14 +111,14 @@ describe("passthrough reasoning summary rewrite honors hideThinkingSummary", () expect(text).not.toContain('"summary":[{"type":"summary_text"'); }); - test("bounded JSON: requested summary moves item content into summary", async () => { + test("bounded JSON: requested summary keeps the content shape", async () => { const response = await runHandleResponses( { model: "deepseek-v4-flash", input: "ping", stream: false, reasoning: { effort: "max", summary: "detailed" } }, JSON_UPSTREAM, "application/json", ); const text = await response.text(); - expect(text).toContain('"summary":[{"type":"summary_text","text":"think"}]'); - expect(text).not.toContain('"content":[{"type":"reasoning_text","text":"think"}]'); + expect(text).toContain('"content":[{"type":"reasoning_text","text":"think"}]'); + expect(text).not.toContain('"summary":[{"type":"summary_text"'); }); }); diff --git a/tests/responses/responses-reasoning-summary-rewrite.test.ts b/tests/responses/responses-reasoning-summary-rewrite.test.ts deleted file mode 100644 index 32940a24e7..0000000000 --- a/tests/responses/responses-reasoning-summary-rewrite.test.ts +++ /dev/null @@ -1,286 +0,0 @@ -import { describe, expect, test } from "bun:test"; -import { - createReasoningSummaryChannelPayloadRewrite, - routeUsesContentChannelReasoning, - rewriteReasoningSummaryInJson, - rewriteReasoningSummaryInJsonString, -} from "../../src/server/responses-reasoning-summary-rewrite"; - -const rewrite = createReasoningSummaryChannelPayloadRewrite(); - -function apply(payload: unknown): unknown { - return JSON.parse(rewrite(JSON.stringify(payload))); -} - -describe("responses reasoning summary channel rewrite", () => { - test("routes reasoning_text.delta through the summary channel", () => { - expect(apply({ - type: "response.reasoning_text.delta", - content_index: 0, - delta: "think", - item_id: "rs_1", - output_index: 0, - sequence_number: 4, - })).toEqual({ - type: "response.reasoning_summary_text.delta", - summary_index: 0, - delta: "think", - item_id: "rs_1", - output_index: 0, - sequence_number: 4, - }); - }); - - test("routes reasoning_text.done through the summary channel", () => { - expect(apply({ - type: "response.reasoning_text.done", - content_index: 0, - text: "full thinking", - item_id: "rs_1", - output_index: 0, - })).toEqual({ - type: "response.reasoning_summary_text.done", - summary_index: 0, - text: "full thinking", - item_id: "rs_1", - output_index: 0, - }); - }); - - test("moves reasoning item content into summary on output_item.done", () => { - expect(apply({ - type: "response.output_item.done", - output_index: 0, - item: { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [{ type: "reasoning_text", text: "thinking" }], - summary: [], - }, - })).toEqual({ - type: "response.output_item.done", - output_index: 0, - item: { - type: "reasoning", - id: "rs_1", - status: "completed", - summary: [{ type: "summary_text", text: "thinking" }], - }, - }); - }); - - test("moves reasoning item content into summary inside response.completed", () => { - const payload = { - type: "response.completed", - response: { - id: "resp_1", - status: "completed", - output: [ - { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [{ type: "reasoning_text", text: "thinking" }], - summary: [], - }, - { type: "message", id: "msg_1", status: "completed", content: [{ type: "output_text", text: "OK" }] }, - ], - }, - }; - const result = apply(payload) as { response: { output: Record[] } }; - expect(result.response.output[0]).toEqual({ - type: "reasoning", - id: "rs_1", - status: "completed", - summary: [{ type: "summary_text", text: "thinking" }], - }); - expect(result.response.output[1]).toEqual(payload.response.output[1]); - }); - - test("leaves summary-channel and message events untouched", () => { - const untouched = [ - { type: "response.reasoning_summary_text.delta", summary_index: 0, delta: "s", item_id: "rs_1", output_index: 0 }, - { type: "response.output_text.delta", content_index: 0, delta: "OK", item_id: "msg_1", output_index: 1 }, - { type: "response.output_item.added", output_index: 1, item: { type: "message", id: "msg_1", status: "in_progress", content: [] } }, - ]; - for (const payload of untouched) { - expect(apply(payload)).toEqual(payload); - } - }); - - test("leaves a reasoning item without content text untouched", () => { - expect(apply({ - type: "response.output_item.done", - output_index: 0, - item: { type: "reasoning", id: "rs_1", status: "completed", content: [], summary: [] }, - })).toEqual({ - type: "response.output_item.done", - output_index: 0, - item: { type: "reasoning", id: "rs_1", status: "completed", content: [], summary: [] }, - }); - }); - - test("preserves a summary-channel reasoning item as-is", () => { - expect(apply({ - type: "response.output_item.done", - output_index: 0, - item: { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [], - summary: [{ type: "summary_text", text: "already summarized" }], - }, - })).toEqual({ - type: "response.output_item.done", - output_index: 0, - item: { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [], - summary: [{ type: "summary_text", text: "already summarized" }], - }, - }); - }); - - test("rewrites reasoning items inside a bare completed response document", () => { - const doc = { - id: "resp_1", - object: "response", - status: "completed", - output: [ - { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [{ type: "reasoning_text", text: "thinking" }], - summary: [], - }, - { type: "message", id: "msg_1", status: "completed", content: [{ type: "output_text", text: "OK" }] }, - ], - }; - const result = rewriteReasoningSummaryInJson(doc) as { output: Record[] }; - expect(result.output[0]).toEqual({ - type: "reasoning", - id: "rs_1", - status: "completed", - summary: [{ type: "summary_text", text: "thinking" }], - }); - expect(result.output[1]).toEqual(doc.output[1]); - }); - - test("rewrites reasoning items inside an SSE completed event document", () => { - const doc = { - type: "response.completed", - response: { - id: "resp_1", - status: "completed", - output: [ - { - type: "reasoning", - id: "rs_1", - status: "completed", - content: [{ type: "reasoning_text", text: "thinking" }], - summary: [], - }, - ], - }, - }; - const result = rewriteReasoningSummaryInJson(doc) as { response: { output: Record[] } }; - expect(result.response.output[0]).toEqual({ - type: "reasoning", - id: "rs_1", - status: "completed", - summary: [{ type: "summary_text", text: "thinking" }], - }); - }); - - test("string-level rewrite leaves summary-channel documents untouched", () => { - const doc = JSON.stringify({ - id: "resp_1", - output: [{ type: "reasoning", id: "rs_1", summary: [{ type: "summary_text", text: "already summarized" }] }], - }); - expect(rewriteReasoningSummaryInJsonString(doc)).toBe(doc); - }); - - test("malformed payloads pass through unchanged", () => { - expect(rewrite("not json")).toBe("not json"); - expect(rewrite("[1,2]")).toBe("[1,2]"); - }); - - // `encrypted_content` is opaque, state-bearing provider data, so preserve the complete item - // shape defensively when the client replays it. This rewrite's round-trip was verified against - // DeepSeek, which is stateless and issues no blob; providers that do issue one joined later - // through `preserveReasoningContentModels`. - describe("items carrying encrypted_content", () => { - const blobItem = { - type: "reasoning", - id: "rs_1", - status: "completed", - encrypted_content: "gAAAAAB-upstream-issued-blob", - content: [{ type: "reasoning_text", text: "thinking" }], - summary: [], - }; - - test("are returned byte-for-byte on output_item.done", () => { - const payload = { type: "response.output_item.done", output_index: 0, item: blobItem }; - expect(apply(payload)).toEqual(payload); - }); - - test("are returned byte-for-byte inside response.completed output", () => { - const payload = { - type: "response.completed", - response: { id: "resp_1", output: [blobItem] }, - }; - expect(apply(payload)).toEqual(payload); - }); - - test("are returned byte-for-byte through the non-streaming document rewrite", () => { - const doc = { id: "resp_1", object: "response", output: [blobItem] }; - expect(rewriteReasoningSummaryInJson(doc)).toBe(doc); - const json = JSON.stringify(doc); - expect(rewriteReasoningSummaryInJsonString(json)).toBe(json); - }); - - // Only the stored item is protected: the live trace Codex renders comes from the delta events, - // which carry no blob and are still routed to the summary channel. - test("do not disable the delta rewrite that renders the live trace", () => { - expect(apply({ - type: "response.reasoning_text.delta", - delta: "think", - item_id: "rs_1", - output_index: 0, - })).toMatchObject({ type: "response.reasoning_summary_text.delta", delta: "think" }); - }); - }); -}); - -describe("routeUsesContentChannelReasoning", () => { - test("statelessResponses providers use the content channel", () => { - expect(routeUsesContentChannelReasoning({ statelessResponses: true }, "deepseek-v4-flash")).toBe(true); - }); - - test("preserveReasoningContentModels lists qualify", () => { - expect(routeUsesContentChannelReasoning( - { preserveReasoningContentModels: ["deepseek-v4-flash"] }, - "deepseek-v4-flash", - )).toBe(true); - }); - - test("model matching is case-insensitive on both sides", () => { - expect(routeUsesContentChannelReasoning( - { preserveReasoningContentModels: ["DeepSeek-V4-Flash"] }, - "deepseek-v4-flash", - )).toBe(true); - expect(routeUsesContentChannelReasoning( - { preserveReasoningContentModels: ["deepseek-v4-flash"] }, - "DeepSeek-V4-Flash", - )).toBe(true); - }); - - test("other providers do not", () => { - expect(routeUsesContentChannelReasoning({}, "gpt-5.5")).toBe(false); - }); -}); diff --git a/tests/responses/responses-show-thinking-summary.test.ts b/tests/responses/responses-show-thinking-summary.test.ts new file mode 100644 index 0000000000..a749518b63 --- /dev/null +++ b/tests/responses/responses-show-thinking-summary.test.ts @@ -0,0 +1,173 @@ +import { afterEach, describe, expect, test } from "bun:test"; +import { mkdtempSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { providerConfigSeed } from "../../src/providers/derive"; +import { getProviderRegistryEntry } from "../../src/providers/registry"; +import { handleResponses } from "../../src/server/responses/core"; +import type { OcxConfig, OcxProviderConfig } from "../../src/types"; + +// Provider-opted visible thinking (showThinkingSummary): a provider that serves +// genuine user-facing reasoning surfaces it on the summary channel even when the +// client omits reasoning.summary (the Codex default, which otherwise hides all +// thinking in replay-only envelopes). An explicit client summary of "none" still +// wins and keeps thinking hidden. + +function shownSeed() { + const seed = providerConfigSeed(getProviderRegistryEntry("deepseek")!); + return { ...seed, apiKey: "sk-test", showThinkingSummary: true } as OcxProviderConfig; +} + +function sseFrame(payload: unknown): string { + return "data: " + JSON.stringify(payload) + "\n\n"; +} + +const SSE_UPSTREAM = [ + sseFrame({ type: "response.created", response: { id: "resp_1", status: "in_progress", output: [] } }), + sseFrame({ type: "response.output_item.added", output_index: 0, item: { type: "reasoning", id: "rs_1", status: "in_progress", content: [], summary: [] } }), + sseFrame({ type: "response.reasoning_text.delta", content_index: 0, delta: "think", item_id: "rs_1", output_index: 0 }), + sseFrame({ type: "response.reasoning_text.done", content_index: 0, text: "think", item_id: "rs_1", output_index: 0 }), + sseFrame({ type: "response.output_item.done", output_index: 0, item: { type: "reasoning", id: "rs_1", status: "completed", content: [{ type: "reasoning_text", text: "think" }], summary: [] } }), + sseFrame({ type: "response.completed", response: { id: "resp_1", status: "completed", output: [{ type: "reasoning", id: "rs_1", status: "completed", content: [{ type: "reasoning_text", text: "think" }], summary: [] }] } }), +].join(""); + +async function runHandleResponses(body: Record, seed: OcxProviderConfig) { + const encoder = new TextEncoder(); + globalThis.fetch = (async () => new Response( + new ReadableStream({ + start(controller) { + controller.enqueue(encoder.encode(SSE_UPSTREAM)); + controller.close(); + }, + }), + { status: 200, headers: { "content-type": "text/event-stream" } }, + )) as typeof fetch; + const config = { providers: { deepseek: seed } } as unknown as OcxConfig; + return handleResponses( + new Request("http://localhost/v1/responses", { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify(body), + }), + config, + { model: "", provider: "" }, + { abortSignal: AbortSignal.timeout(5_000) }, + ); +} + +describe("showThinkingSummary provider option", () => { + const originalFetch = globalThis.fetch; + afterEach(() => { globalThis.fetch = originalFetch; }); + + test("provider opt-in never relabels raw content as a summary", async () => { + const response = await runHandleResponses( + { model: "deepseek-v4-flash", input: "ping", stream: true }, + shownSeed(), + ); + const text = await response.text(); + expect(text).not.toContain("response.reasoning_summary_text.delta"); + expect(text).toContain('"content":[{"type":"reasoning_text","text":"think"}]'); + }); + + test("explicit client summary none keeps thinking hidden", async () => { + const response = await runHandleResponses( + { model: "deepseek-v4-flash", input: "ping", stream: true, reasoning: { summary: "none" } }, + shownSeed(), + ); + const text = await response.text(); + expect(text).not.toContain("response.reasoning_summary_text.delta"); + expect(text).toContain("response.reasoning_text.delta"); + }); + + test("without the provider option, omitted summary stays hidden", async () => { + const seed = { ...providerConfigSeed(getProviderRegistryEntry("deepseek")!), apiKey: "sk-test" } as OcxProviderConfig; + const response = await runHandleResponses( + { model: "deepseek-v4-flash", input: "ping", stream: true }, + seed, + ); + const text = await response.text(); + expect(text).not.toContain("response.reasoning_summary_text.delta"); + expect(text).toContain("response.reasoning_text.delta"); + }); + + test("google-antigravity preset opts in", () => { + expect(providerConfigSeed(getProviderRegistryEntry("google-antigravity")!).showThinkingSummary).toBe(true); + expect(providerConfigSeed(getProviderRegistryEntry("deepseek")!).showThinkingSummary).toBeUndefined(); + }); + + for (const stream of [false, true]) for (const [summary, providerFlag, visible] of [ + [undefined, undefined, true], ["none", true, false], [undefined, false, false], ["auto", false, false], ["auto", true, true], + ] as const) test(`CCA summary=${summary} provider=${providerFlag} stream=${stream}`, async () => { + const home = mkdtempSync(join(tmpdir(), "ocx-show-thinking-")); + const prevHome = process.env.OPENCODEX_HOME; + process.env.OPENCODEX_HOME = home; + writeFileSync(join(home, "auth.json"), JSON.stringify({ + "google-antigravity": { + activeAccountId: "active", + accounts: [{ + id: "active", + credential: { + access: "access-token", + refresh: "refresh-token", + expires: Date.now() + 3_600_000, + projectId: "project-id", + }, + }], + }, + })); + const seen: string[] = []; + const requests: Array<{ request: { generationConfig?: { thinkingConfig?: { includeThoughts?: boolean } } } }> = []; + globalThis.fetch = (async (input: RequestInfo | URL, init?: RequestInit) => { + seen.push(String(input)); + requests.push(JSON.parse(String(init?.body))); + const requestedThoughts = requests.at(-1)?.request.generationConfig?.thinkingConfig?.includeThoughts === true; + const payload = { + response: { + candidates: [{ + content: { parts: [...(requestedThoughts ? [{ thought: true, text: "cca-think" }] : []), { text: "OK" }] }, + finishReason: "STOP", + }], + usageMetadata: { promptTokenCount: 10, candidatesTokenCount: 5, totalTokenCount: 15, thoughtsTokenCount: 3 }, + }, + }; + return stream + ? new Response(sseFrame(payload), { headers: { "content-type": "text/event-stream" } }) + : Response.json(payload); + }) as typeof fetch; + try { + const seed = { + ...providerConfigSeed(getProviderRegistryEntry("google-antigravity")!), + liveModels: false, + models: ["gemini-3.8-flash"], + } as OcxProviderConfig; + // Simulate a saved provider row written before the registry learned the flag: + // the request path must backfill it from the registry entry (routedProviderConfig), + // enrichProviderFromRegistry never runs there. + delete seed.showThinkingSummary; + if (providerFlag !== undefined) seed.showThinkingSummary = providerFlag; + const config = { providers: { "google-antigravity": seed } } as unknown as OcxConfig; + const response = await handleResponses( + new Request("http://localhost/v1/responses", { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ model: "google-antigravity/gemini-3.8-flash", input: "ping", stream, reasoning: { effort: "low", ...(summary ? { summary } : {}) } }), + }), + config, + { model: "", provider: "" }, + { abortSignal: AbortSignal.timeout(10_000) }, + ); + const text = await response.text(); + expect(seen).toHaveLength(1); + expect(seen[0]).toContain(stream ? "v1internal:streamGenerateContent" : "v1internal:generateContent"); + expect(text.includes('"summary":[{"type":"summary_text","text":"cca-think"}]')).toBe(visible); + expect(requests[0].request.generationConfig?.thinkingConfig?.includeThoughts === true) + .toBe(visible && providerFlag !== false); + if (stream) expect(text).toContain("response.completed"); + expect(text).toContain("OK"); + } finally { + if (prevHome === undefined) delete process.env.OPENCODEX_HOME; + else process.env.OPENCODEX_HOME = prevHome; + rmSync(home, { recursive: true, force: true }); + } + }); +}); diff --git a/tests/responses/sse-failed-tail.test.ts b/tests/responses/sse-failed-tail.test.ts index f21301ffb7..0a38424f42 100644 --- a/tests/responses/sse-failed-tail.test.ts +++ b/tests/responses/sse-failed-tail.test.ts @@ -422,3 +422,43 @@ describe("relaySseWithFailedTail", () => { expect(out).not.toContain("event: response.failed"); }); }); + + +describe("optional Codex hint filtering preserves relay semantics", () => { + for (const eager of [false, true]) { + const relay = (chunks: string[], upstreamError?: string) => eager + ? relaySseEagerBounded(sourceStream(chunks), new AbortController(), parityHooks, + { upstreamError, terminalBoundary: { dropCodexSafetyBuffering: true } }) + : relaySseWithFailedTail(sourceStream(chunks), new AbortController(), undefined, + { upstreamError, terminalBoundary: { dropCodexSafetyBuffering: true } }); + test(`policy failure plus hint is composed, eager=${eager}`, async () => { + const frame = `event: error\r\ndata: ${JSON.stringify({ type: "error", safety_buffering: { enabled: true }, + error: { code: "cyber_policy", message: "blocked by upstream policy", type: "invalid_request_error" } })}\r\n\r\n`; + const text = await drain(relay([frame.slice(0, 19), frame.slice(19)])); + expect(text).not.toContain("safety_buffering"); + expect(text).toContain('"type":"response.failed"'); + expect(text).toContain('"code":"cyber_policy"'); + expect(text).toContain('"retryable":false'); + expect(text.match(/data: \[DONE\]/g)).toHaveLength(1); + }); + test(`metadata removal retains other frames and one terminal, eager=${eager}`, async () => { + const metadata = 'data: {"type":"response.metadata","metadata":{"type":"safety_buffering"}}\n\n'; + const other = 'data: {"type":"codex.response.metadata","headers":{"x-codex-safety-buffering-enabled":"true"}}\n\n'; + const malformed = 'data: {malformed}\n\n'; + const terminal = 'data: {"type":"response.completed","response":{"status":"completed"},"safety_buffering":true}\n\ndata: [DONE]\n\n'; + const text = await drain(relay([metadata.slice(0, 7), metadata.slice(7), other, malformed, terminal])); + expect(text).not.toContain('"type":"safety_buffering"'); + expect(text).not.toContain('"safety_buffering":true'); + expect(text).toContain(other); + expect(text).toContain(malformed); + expect(text).toContain('"type":"response.completed"'); + expect(text.match(/data: \[DONE\]/g)).toHaveLength(1); + }); + test(`hint-only EOF preserves captured error fallback, eager=${eager}`, async () => { + const text = await drain(relay(['data: {"type":"response.metadata","metadata":{"type":"safety_buffering"}}\n\n'], "provider unavailable")); + expect(text).toContain("provider unavailable"); + expect(text).not.toContain("adapter_eof"); + expect(text).not.toContain("safety_buffering"); + }); + } +}); diff --git a/tests/responses/ws-upstream-reuse.test.ts b/tests/responses/ws-upstream-reuse.test.ts index b957fdb317..48b720d2ea 100644 --- a/tests/responses/ws-upstream-reuse.test.ts +++ b/tests/responses/ws-upstream-reuse.test.ts @@ -3,6 +3,8 @@ import { codexWsUpstreamFetch } from "../../src/server/responses/ws-upstream"; import { runOptionalShutdownHooks } from "../../src/lib/optional-shutdown-hooks"; import { CodexWsPool, codexWsPool } from "../../src/server/responses/codex-ws-pool"; import { prepareCodexWsRequest } from "../../src/server/responses/codex-ws-request"; +import { createResponsesPassthroughAdapter } from "../../src/adapters/openai-responses"; +import { withTestTranslatorBudget } from "../helpers/translator-budget"; const URL = "https://chatgpt.com/backend-api/codex/responses"; const realWebSocket = globalThis.WebSocket; @@ -307,3 +309,31 @@ test("a Lite mode change retires the old handshake", async () => { expect(Socket.all).toHaveLength(2); expect(Socket.all[0]!.readyState).toBe(3); }); + +test("adapter Spark Lite override retires a legacy socket and reuses the disabled identity", async () => { + const liteHeader = "x-openai-internal-codex-responses-lite"; + const liteKey = "ws_request_header_x_openai_internal_codex_responses_lite"; + const options = init(); + const rawBody = { ...JSON.parse(options.body as string), model: "gpt-5.3-codex-spark", + client_metadata: { thread_id: "fixture-thread", turn_id: "fixture-turn", [liteKey]: "true" } }; + const before = JSON.stringify(rawBody); + const adapter = withTestTranslatorBudget(createResponsesPassthroughAdapter({ + adapter: "openai-responses", authMode: "forward", baseUrl: "https://chatgpt.com/backend-api/codex", + })); + const built = await adapter.buildRequest({ modelId: "spark-alias", context: { messages: [] }, + stream: true, options: {}, _rawBody: rawBody, + }, { headers: new Headers(options.headers) }); + const current = { ...options, body: built.body, headers: built.headers }; + // Keep the exact same Spark model/scope/headers; only the old delete-only Lite policy differs. + const legacyHeaders = new Headers(current.headers); + legacyHeaders.delete(liteHeader); + await drain({ ...current, headers: legacyHeaders }); + await drain(current); + await drain(current); + expect(Socket.all).toHaveLength(2); + expect(Socket.all.map(socket => socket.readyState)).toEqual([3, 1]); + expect(Socket.all.map(socket => socket.frames.map(frame => + (frame.client_metadata as Record)[liteKey]))).toEqual([["true"], ["false", "false"]]); + expect(Socket.all.flatMap(socket => socket.frames).every(frame => frame.model === rawBody.model)).toBe(true); + expect(JSON.stringify(rawBody)).toBe(before); +}); diff --git a/tests/responses/ws-upstream.test.ts b/tests/responses/ws-upstream.test.ts index 80c6291780..1260305d5d 100644 --- a/tests/responses/ws-upstream.test.ts +++ b/tests/responses/ws-upstream.test.ts @@ -1,3 +1,7 @@ +import { + PLAINTEXT_V2_AGENT_MESSAGE_RESTORE_OVERFLOW_MESSAGE, + PLAINTEXT_V2_COLLABORATION_NAMESPACE, +} from "../../src/responses/plaintext-v2-agent-messages"; import { afterEach, beforeEach, describe, expect, jest, test } from "bun:test"; import { providerFetch } from "../../src/server/responses/fetch-helpers"; import { handleResponses } from "../../src/server/responses"; @@ -346,6 +350,206 @@ describe("handleResponses Codex WS relay selection", () => { }); } + function plaintextV2CollaborationRequest(): Request { + return new Request("http://localhost/v1/responses", { + method: "POST", + headers: { "content-type": "application/json", authorization: "Bearer test" }, + body: JSON.stringify({ + model: "gpt-5.5", + stream: true, + tools: [{ type: "namespace", name: "collaboration", tools: [{ + type: "function", name: "spawn_agent", parameters: { + type: "object", properties: { message: { type: "string", encrypted: true } }, + }, + }, { type: "function", name: "send_message", parameters: { type: "object" } }] }], + input: [ + { + type: "additional_tools", + tools: [{ + type: "namespace", + name: "collaboration", + tools: [ + { + type: "function", + name: "spawn_agent", + parameters: { + type: "object", + properties: { message: { type: "string", encrypted: true } }, + }, + }, + { type: "function", name: "send_message", parameters: { type: "object" } }, + ], + }], + }, + { type: "message", role: "user", content: [{ type: "input_text", text: "delegate" }] }, + ], + }), + }); + } + + test.each(["collaboration-optimize", null])("plaintext v2 WS restoration handles namespace=%s", async namespace => { + installFake(ws => { + ws.emit("open", {}); + ws.emit("message", { + data: JSON.stringify({ + type: "response.created", + response: { + id: "r-plaintext-v2-ws", + object: "response", + status: "in_progress", + output: [], + }, + }), + }); + ws.emit("message", { + data: JSON.stringify({ + type: "response.output_item.added", + output_index: 0, + item: { + type: "function_call", + id: "fc_spawn", + call_id: "call-spawn", + namespace, + name: "start_delegated_task", + arguments: "", + encrypted_function_args: [], + status: "in_progress", + }, + }), + }); + ws.emit("message", { + data: JSON.stringify({ + type: "response.function_call_arguments.done", + item_id: "fc_spawn", + output_index: 0, + namespace, + name: "collaboration-optimize__start_delegated_task", + arguments: JSON.stringify({ message: "plain WS assignment" }), + encrypted_function_args: [], + }), + }); + ws.emit("message", { + data: JSON.stringify({ + type: "response.output_item.done", + output_index: 0, + item: { + type: "function_call", + id: "fc_spawn", + call_id: "call-spawn", + namespace, + name: "start_delegated_task", + arguments: JSON.stringify({ message: "plain WS assignment" }), + encrypted_function_args: [], + status: "completed", + }, + }), + }); + ws.emit("message", { + data: JSON.stringify({ + type: "response.completed", + response: { + id: "r-plaintext-v2-ws", + status: "completed", + output: [{ + type: "function_call", + id: "fc_spawn", + call_id: "call-spawn", + namespace, + name: "start_delegated_task", + arguments: JSON.stringify({ message: "plain WS assignment" }), + encrypted_function_args: [], + status: "completed", + }], + }, + }), + }); + }); + const config = { ...forwardConfig(), plaintextV2AgentMessages: true } as OcxConfig; + const request = plaintextV2CollaborationRequest(); + + const response = await handleResponses(request, config, { model: "", provider: "" }, { + codexWsRuntimeIdentity: BOUNDED_WS_RUNTIME, + }); + expect(FakeWebSocket.instances).toHaveLength(1); + const frame = JSON.parse(FakeWebSocket.instances[0]!.sent[0]!) as { + type: string; + stream?: unknown; + input: Array>; + }; + const additionalTools = frame.input.find(item => item.type === "additional_tools") as { + tools: Array<{ + name: string; + tools: Array<{ + name: string; + parameters: { properties: { message: Record } }; + }>; + }>; + }; + expect(frame.type).toBe("response.create"); + expect(frame.stream).toBeUndefined(); + expect(additionalTools.tools[0]!.name).toBe("collaboration-optimize"); + expect(additionalTools.tools[0]!.tools[0]!.name).toBe("start_delegated_task"); + expect(additionalTools.tools[0]!.tools[0]!.parameters.properties.message.encrypted).toBeUndefined(); + + expect(isEagerRelaySseResponse(response)).toBe(true); + const clientText = await response.text(); + expect(clientText).toContain("response.function_call_arguments.done"); + const argumentDoneLine = clientText.split("\n") + .find(line => line.includes('"response.function_call_arguments.done"'))!; + const argumentDone = JSON.parse(argumentDoneLine.replace(/^data: /, "")) as Record; + expect(argumentDone.namespace).toBe("collaboration"); + expect(argumentDone.name).toBe("spawn_agent"); + expect(argumentDone.encrypted_function_args).toEqual([]); + const completedLine = clientText.split("\n") + .find(line => line.includes('"response.completed"'))!; + const completed = JSON.parse(completedLine.replace(/^data: /, "")) as { + response: { output: Array> }; + }; + expect(completed.response.output[0]!.namespace).toBe("collaboration"); + expect(completed.response.output[0]!.name).toBe("spawn_agent"); + expect(completed.response.output[0]!.encrypted_function_args).toEqual([]); + }); + + test("plaintext v2 restoration overflow fails closed on the WS upstream path", async () => { + installFake(ws => { + ws.emit("open", {}); + ws.emit("message", { + data: JSON.stringify({ + type: "response.completed", + response: { + id: "r-plaintext-v2-ws-overflow", + status: "completed", + output: Array.from({ length: 10_000 }, (_, index) => ({ + type: "function_call", + call_id: `call-${index}`, + namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + name: "start_delegated_task", + arguments: "{}", + })), + }, + }), + }); + }); + const config = { ...forwardConfig(), plaintextV2AgentMessages: true } as OcxConfig; + + const response = await handleResponses( + plaintextV2CollaborationRequest(), + config, + { model: "", provider: "" }, + { codexWsRuntimeIdentity: BOUNDED_WS_RUNTIME }, + ); + + expect(FakeWebSocket.instances).toHaveLength(1); + expect(isEagerRelaySseResponse(response)).toBe(true); + const clientText = await response.text(); + expect(clientText).toContain("event: response.failed"); + expect(clientText).toContain(PLAINTEXT_V2_AGENT_MESSAGE_RESTORE_OVERFLOW_MESSAGE); + expect(clientText).toContain("data: [DONE]"); + expect(clientText).not.toContain(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(clientText).not.toContain("start_delegated_task"); + expect(FakeWebSocket.instances[0]!.closed).toBe(true); + }); + test("a successful WS upgrade bypasses the configured legacy tee path", async () => { installFake(ws => { ws.emit("open", {}); diff --git a/tests/server/config.test.ts b/tests/server/config.test.ts index b096b5857c..8dbc482403 100644 --- a/tests/server/config.test.ts +++ b/tests/server/config.test.ts @@ -777,6 +777,19 @@ describe("opencodex config defaults", () => { }); }); + test("codex safety-buffering header drop is an explicit top-level opt-in", () => { + const defaults = getDefaultConfig(); + expect(defaults.dropCodexSafetyBuffering).toBe(false); + expect(validateConfigCandidate({ ...defaults, dropCodexSafetyBuffering: true })).toMatchObject({ + ok: true, + config: { dropCodexSafetyBuffering: true }, + }); + expect(validateConfigCandidate({ ...defaults, dropCodexSafetyBuffering: "yes" })).toMatchObject({ + ok: false, + error: expect.stringContaining("dropCodexSafetyBuffering"), + }); + }); + test("usage and MCP config overrides change the effective bound while defaults remain compatible", () => { const defaults = getDefaultConfig(); expect(defaults.managementUsageMaxReadBytes).toBe(64 * 1024 * 1024); @@ -1216,6 +1229,46 @@ describe("opencodex config defaults", () => { } }); + test("plaintextV2AgentMessages is explicit and degrades invalid hand edits", () => { + const base = { + port: 12345, + providers: { + custom: { + adapter: "openai-responses", + baseUrl: "https://example.test/v1", + }, + }, + defaultProvider: "custom", + }; + expect(getDefaultConfig().plaintextV2AgentMessages).toBeUndefined(); + + writeConfig({ ...base, plaintextV2AgentMessages: true }); + expect(loadConfig()).toMatchObject({ ...base, plaintextV2AgentMessages: true }); + expect(validateConfigCandidate({ ...base, plaintextV2AgentMessages: true })).toMatchObject({ + ok: true, + config: { plaintextV2AgentMessages: true }, + }); + + for (const invalid of [null, "true", 1, {}]) { + writeConfig({ ...base, plaintextV2AgentMessages: invalid }); + const diagnostics = readConfigDiagnostics(); + expect(diagnostics).toMatchObject({ + source: "file", + error: null, + config: base, + }); + expect(diagnostics.config.plaintextV2AgentMessages).toBeUndefined(); + expect(diagnostics.warnings).toContain( + "plaintextV2AgentMessages ignored: expected a boolean", + ); + expect(validateConfigCandidate({ ...base, plaintextV2AgentMessages: invalid })).toMatchObject({ + ok: false, + error: expect.stringContaining("plaintextV2AgentMessages"), + }); + expect(backupNames()).toEqual([]); + } + }); + test("native subagent-default sync is opt-in and ignores malformed opt-ins without falling back", () => { const base = { port: 12345, diff --git a/tests/server/plaintext-v2-agent-messages-server.test.ts b/tests/server/plaintext-v2-agent-messages-server.test.ts new file mode 100644 index 0000000000..cfbba40249 --- /dev/null +++ b/tests/server/plaintext-v2-agent-messages-server.test.ts @@ -0,0 +1,692 @@ +import { warnPlaintextV2AgentMessagesStartup } from "../../src/server"; +import { afterEach, beforeEach, describe, expect, test } from "bun:test"; +import { mkdtempSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import { join } from "node:path"; +import { saveCodexAccountCredential } from "../../src/codex/account-store"; +import { clearAccountQuota, updateAccountQuota } from "../../src/codex/auth-api"; +import { clearCodexUpstreamHealth, clearThreadAccountMap } from "../../src/codex/routing"; +import { CODEX_FORWARD_BASE_URL } from "../../src/providers/openai-tiers"; +import { PROVIDER_REGISTRY } from "../../src/providers/registry"; +import { + PLAINTEXT_V2_AGENT_MESSAGE_RESTORE_OVERFLOW_MESSAGE, + PLAINTEXT_V2_COLLABORATION_NAMESPACE, +} from "../../src/responses/plaintext-v2-agent-messages"; +import { clearResponseStateForTests, expandPreviousResponseInput } from "../../src/responses/state"; +import { handleResponses } from "../../src/server/responses"; +import type { OcxConfig } from "../../src/types"; + +const originalFetch = globalThis.fetch; +beforeEach(() => { clearResponseStateForTests(); }); +afterEach(() => { + globalThis.fetch = originalFetch; + clearResponseStateForTests(); +}); + +function config( + enabled: boolean, + snapshotRepair = false, + streamMode?: "auto" | "legacy-tee" | "eager-relay", +): OcxConfig { + return { + defaultProvider: "native", + providers: { + native: { + adapter: "openai-responses", + baseUrl: CODEX_FORWARD_BASE_URL, + authMode: "forward", + ...(snapshotRepair ? { responsesSnapshotRepair: true } : {}), + }, + }, + plaintextV2AgentMessages: enabled, + ...(streamMode ? { streamMode } : {}), + } as OcxConfig; +} + +function collaborationRequest(options: { + input?: unknown[]; + model?: string; + previousResponseId?: string; + toolChoice?: unknown; +} = {}): Request { + return new Request("http://localhost/v1/responses", { + method: "POST", + headers: { "content-type": "application/json" }, + body: JSON.stringify({ + model: options.model ?? "native/gpt-5.6-sol", + store: false, + stream: true, + input: options.input ?? [{ + type: "message", + role: "user", + content: [{ type: "input_text", text: "delegate" }], + }], + ...(options.previousResponseId ? { previous_response_id: options.previousResponseId } : {}), + ...(options.toolChoice ? { tool_choice: options.toolChoice } : {}), + tools: [{ + type: "namespace", + name: "collaboration", + tools: [ + { + type: "function", + name: "spawn_agent", + parameters: { + type: "object", + properties: { message: { type: "string", encrypted: true } }, + required: ["message"], + }, + }, + { type: "function", name: "send_message", parameters: { type: "object" } }, + ], + }], + }), + }); +} + +async function withPoolHome(run: () => Promise): Promise { + const home = mkdtempSync(join(tmpdir(), "ocx-plaintext-v2-pool-")); + const previousOpencodexHome = process.env.OPENCODEX_HOME; + const previousCodexHome = process.env.CODEX_HOME; + process.env.OPENCODEX_HOME = home; + process.env.CODEX_HOME = home; + clearCodexUpstreamHealth(); + clearThreadAccountMap(); + clearAccountQuota(); + try { + return await run(); + } finally { + clearCodexUpstreamHealth(); + clearThreadAccountMap(); + clearAccountQuota(); + rmSync(home, { recursive: true, force: true }); + if (previousOpencodexHome === undefined) delete process.env.OPENCODEX_HOME; + else process.env.OPENCODEX_HOME = previousOpencodexHome; + if (previousCodexHome === undefined) delete process.env.CODEX_HOME; + else process.env.CODEX_HOME = previousCodexHome; + } +} + +function completedResponsePayload(id = "resp-plaintext-v2") { + return { + id, + status: "completed", + output: [{ + type: "function_call", + call_id: "call-spawn", + namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + name: "start_delegated_task", + arguments: JSON.stringify({ message: "plain assignment" }), + encrypted_function_args: [], + }], + }; +} + +function overLimitResponsePayload(id = "resp-plaintext-v2-overflow") { + return { + id, + status: "completed", + output: Array.from({ length: 10_000 }, (_, index) => ({ + type: "function_call", + call_id: `call-${index}`, + namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + name: "start_delegated_task", + arguments: "{}", + })), + }; +} + +describe("plaintext v2 agent messages at the Responses server boundary", () => { + test.each(["json", "legacy-tee", "eager-relay"] as const)("null namespace restores before %s delivery and continuation storage", async mode => { + const id = `resp-null-namespace-${mode}`; + const item = { ...completedResponsePayload(id).output[0]!, namespace: null }; + const payload = { id, status: "completed", output: [item] }; + globalThis.fetch = (async () => mode === "json" ? Response.json(payload) : new Response( + `event: response.output_item.added\ndata: ${JSON.stringify({ type: "response.output_item.added", output_index: 0, item })}\n\n` + + `event: response.completed\ndata: ${JSON.stringify({ type: "response.completed", response: payload })}\n\ndata: [DONE]\n\n`, + { headers: { "content-type": "text/event-stream" } }, + )) as typeof fetch; + const response = await handleResponses(collaborationRequest(), config(true, false, mode === "json" ? undefined : mode), { model: "", provider: "" }); + const text = await response.text(); + expect(text).toContain('"namespace":"collaboration"'); + expect(text).toContain('"name":"spawn_agent"'); + expect(text).not.toContain('"name":"start_delegated_task"'); + const replay = expandPreviousResponseInput({ previous_response_id: id, input: [] }) as { input: Array> }; + expect(replay.input.find(value => value.type === "function_call")).toMatchObject({ namespace: "collaboration", name: "spawn_agent" }); + expect(JSON.stringify(replay)).not.toContain("start_delegated_task"); + }); + + test("rewrites the canonical request and restores every SSE response snapshot", async () => { + const sentBodies: string[] = []; + globalThis.fetch = (async (_input: RequestInfo | URL, init?: RequestInit) => { + sentBodies.push(typeof init?.body === "string" ? init.body : ""); + const response = completedResponsePayload(); + return new Response( + `event: response.output_item.added\ndata: ${JSON.stringify({ + type: "response.output_item.added", + output_index: 0, + item: response.output[0], + })}\n\nevent: response.function_call_arguments.done\ndata: ${JSON.stringify({ + type: "response.function_call_arguments.done", + item_id: "fc-spawn", + namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + name: `${PLAINTEXT_V2_COLLABORATION_NAMESPACE}__start_delegated_task`, + arguments: JSON.stringify({ message: "plain assignment" }), + encrypted_function_args: [], + })}\n\nevent: response.completed\ndata: ${JSON.stringify({ + type: "response.completed", + response, + })}\n\ndata: [DONE]\n\n`, + { status: 200, headers: { "content-type": "text/event-stream" } }, + ); + }) as typeof fetch; + + const response = await handleResponses( + collaborationRequest(), + config(true), + { model: "", provider: "" }, + ); + const clientBody = await response.text(); + const sentBody = JSON.parse(sentBodies[0]!) as { + tools: Array<{ + name: string; + tools: Array<{ + name: string; + parameters: { properties: { message: Record } }; + }>; + }>; + }; + + expect(sentBodies).toHaveLength(1); + expect(sentBody.tools[0]!.name).toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(sentBody.tools[0]!.tools[0]!.name).toBe("start_delegated_task"); + expect(sentBody.tools[0]!.tools[0]!.parameters.properties.message.encrypted).toBeUndefined(); + expect(clientBody).not.toContain(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(clientBody).toContain('"namespace":"collaboration"'); + expect(clientBody).toContain('"name":"spawn_agent"'); + expect(clientBody).toContain('"encrypted_function_args":[]'); + }); + + test("restores the namespace in bounded JSON responses", async () => { + globalThis.fetch = (async () => new Response(JSON.stringify(completedResponsePayload()), { + status: 200, + headers: { "content-type": "application/json" }, + })) as typeof fetch; + + const response = await handleResponses( + collaborationRequest(), + config(true), + { model: "", provider: "" }, + ); + const clientBody = await response.text(); + + expect(clientBody).not.toContain(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(clientBody).toContain('"namespace":"collaboration"'); + expect(clientBody).toContain('"encrypted_function_args":[]'); + }); + + test("rejects an unclassified successful response while restoration is required", async () => { + globalThis.fetch = (async () => new Response( + JSON.stringify(completedResponsePayload()), + { status: 200 }, + )) as typeof fetch; + + const response = await handleResponses( + collaborationRequest(), + config(true), + { model: "", provider: "" }, + ); + const clientBody = await response.text(); + + expect(response.status).toBe(502); + expect(clientBody).toContain("unsupported content type"); + expect(clientBody).not.toContain(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(clientBody).not.toContain("start_delegated_task"); + }); + + test("restores aliases after SSE snapshot repair copies request tools and tool choice", async () => { + globalThis.fetch = (async () => { + const response = completedResponsePayload("resp-snapshot-sse"); + return new Response( + `event: response.completed\ndata: ${JSON.stringify({ + type: "response.completed", + response, + })}\n\ndata: [DONE]\n\n`, + { status: 200, headers: { "content-type": "text/event-stream" } }, + ); + }) as typeof fetch; + + const response = await handleResponses( + collaborationRequest({ + toolChoice: { type: "function", namespace: "collaboration", name: "spawn_agent" }, + }), + config(true, true), + { model: "", provider: "" }, + ); + const completedLine = (await response.text()).split("\n") + .find(line => line.includes('"response.completed"'))!; + const completed = JSON.parse(completedLine.replace(/^data: /, "")) as { + response: { + tool_choice: { namespace: string }; + tools: Array<{ name: string }>; + output: Array<{ namespace: string; encrypted_function_args: unknown[] }>; + }; + }; + + expect(completed.response.tool_choice.namespace).toBe("collaboration"); + expect(completed.response.tools[0]!.name).toBe("collaboration"); + expect(completed.response.output[0]!.namespace).toBe("collaboration"); + expect(completed.response.output[0]!.encrypted_function_args).toEqual([]); + }); + + test("restores aliases after bounded JSON snapshot repair", async () => { + globalThis.fetch = (async () => new Response( + JSON.stringify(completedResponsePayload("resp-snapshot-json")), + { status: 200, headers: { "content-type": "application/json" } }, + )) as typeof fetch; + + const response = await handleResponses( + collaborationRequest({ + toolChoice: { type: "function", namespace: "collaboration", name: "spawn_agent" }, + }), + config(true, true), + { model: "", provider: "" }, + ); + const completed = await response.json() as { + tool_choice: { namespace: string }; + tools: Array<{ name: string }>; + output: Array<{ namespace: string; encrypted_function_args: unknown[] }>; + }; + + expect(completed.tool_choice.namespace).toBe("collaboration"); + expect(completed.tools[0]!.name).toBe("collaboration"); + expect(completed.output[0]!.namespace).toBe("collaboration"); + expect(completed.output[0]!.encrypted_function_args).toEqual([]); + }); + + test("keeps the marker and reserved namespace when the option is disabled", async () => { + const sentBodies: string[] = []; + globalThis.fetch = (async (_input: RequestInfo | URL, init?: RequestInit) => { + sentBodies.push(typeof init?.body === "string" ? init.body : ""); + return new Response("data: [DONE]\n\n", { + status: 200, + headers: { "content-type": "text/event-stream" }, + }); + }) as typeof fetch; + + await handleResponses(collaborationRequest(), config(false), { model: "", provider: "" }); + const sentBody = JSON.parse(sentBodies[0]!) as { + tools: Array<{ name: string; tools: Array<{ parameters: { properties: { message: Record } } }> }>; + }; + + expect(sentBodies).toHaveLength(1); + expect(sentBody.tools[0]!.name).toBe("collaboration"); + expect(sentBody.tools[0]!.tools[0]!.parameters.properties.message.encrypted).toBe(true); + }); + + test("keeps the whole request unchanged when tool-search history conflicts with the alias", async () => { + const sentBodies: string[] = []; + globalThis.fetch = (async (_input: RequestInfo | URL, init?: RequestInit) => { + sentBodies.push(typeof init?.body === "string" ? init.body : ""); + return new Response("data: [DONE]\n\n", { + status: 200, + headers: { "content-type": "text/event-stream" }, + }); + }) as typeof fetch; + + await handleResponses(collaborationRequest({ + input: [{ + type: "tool_search_output", + tools: [{ + type: "namespace", + name: PLAINTEXT_V2_COLLABORATION_NAMESPACE, + tools: [], + }], + }], + }), config(true), { model: "", provider: "" }); + + const sent = JSON.parse(sentBodies[0]!) as { + tools: Array<{ + name: string; + tools: Array<{ parameters: { properties: { message: Record } } }>; + }>; + }; + expect(sent.tools[0]!.name).toBe("collaboration"); + expect(sent.tools[0]!.tools[0]!.parameters.properties.message.encrypted).toBe(true); + }); + + test("rebuilds the plaintext alias after a canonical pool quota retry", async () => { + await withPoolHome(async () => { + const poolConfig = { + defaultProvider: "openai", + activeCodexAccountId: "pool-a", + autoSwitchThreshold: 0, + providers: { + openai: { + adapter: "openai-responses", + baseUrl: CODEX_FORWARD_BASE_URL, + authMode: "forward", + codexAccountMode: "pool", + }, + }, + codexAccounts: ["pool-a", "pool-b"].map(id => ({ + id, + email: `${id}@example.test`, + isMain: false, + chatgptAccountId: `${id}_chatgpt`, + })), + plaintextV2AgentMessages: true, + } as OcxConfig; + for (const [index, id] of ["pool-a", "pool-b"].entries()) { + saveCodexAccountCredential(id, { + accessToken: `${id}-access-token`, + refreshToken: `${id}-refresh-token`, + expiresAt: Date.now() + 300_000, + chatgptAccountId: `${id}_chatgpt`, + }); + updateAccountQuota(id, 10 + index * 10); + } + + const sentBodies: string[] = []; + const entitlementSnapshot = { + modelsByAccount: new Map([ + ["pool-a", new Set(["gpt-5.6-sol"])], + ["pool-b", new Set(["gpt-5.6-sol"])], + ]), + confirmedAccountIds: new Set(["pool-a", "pool-b"]), + credentialIdentities: new Map(), + }; + globalThis.fetch = (async (_input: RequestInfo | URL, init?: RequestInit) => { + sentBodies.push(typeof init?.body === "string" ? init.body : ""); + if (sentBodies.length === 1) { + return Response.json({ error: { message: "rate limited" } }, { + status: 429, + headers: { "retry-after": "1" }, + }); + } + return Response.json(completedResponsePayload("resp-pool-retry")); + }) as typeof fetch; + + const response = await handleResponses( + collaborationRequest({ model: "gpt-5.6-sol" }), + poolConfig, + { model: "", provider: "" }, + { resolveCodexModelEntitlements: async () => entitlementSnapshot }, + ); + const clientBody = await response.text(); + + expect(sentBodies).toHaveLength(2); + for (const body of sentBodies) { + const sent = JSON.parse(body) as { tools: Array<{ name: string }> }; + expect(sent.tools[0]!.name).toBe(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + } + expect(clientBody).toContain('"namespace":"collaboration"'); + expect(clientBody).not.toContain(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + }); + }); + + test("fails closed for over-limit streamed responses in both relay modes", async () => { + for (const streamMode of ["legacy-tee", "eager-relay"] as const) { + globalThis.fetch = (async () => new Response( + `event: response.completed\ndata: ${JSON.stringify({ + type: "response.completed", + response: overLimitResponsePayload(`resp-${streamMode}`), + })}\n\ndata: [DONE]\n\n`, + { status: 200, headers: { "content-type": "text/event-stream" } }, + )) as typeof fetch; + + const response = await handleResponses( + collaborationRequest(), + config(true, false, streamMode), + { model: "", provider: "" }, + ); + const clientBody = await response.text(); + + expect(response.status).toBe(200); + expect(clientBody).toContain('"type":"response.failed"'); + expect(clientBody).toContain(PLAINTEXT_V2_AGENT_MESSAGE_RESTORE_OVERFLOW_MESSAGE); + expect(clientBody).toContain("data: [DONE]"); + expect(clientBody).not.toContain(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(clientBody).not.toContain("start_delegated_task"); + } + }); + + test("rejects over-limit bounded JSON before HTTP or WebSocket reframing", async () => { + const fixtureId = "plaintext-v2-bounded-json-fixture"; + const fixtureModel = "fixture-model"; + const mutableRegistry = PROVIDER_REGISTRY as unknown as Array>; + mutableRegistry.push({ + id: fixtureId, + label: "Plaintext V2 bounded JSON fixture", + baseUrl: CODEX_FORWARD_BASE_URL, + adapter: "openai-responses", + authKind: "forward", + models: [fixtureModel], + defaultModel: fixtureModel, + modelResponsesUpstreamStreaming: { [fixtureModel]: false }, + }); + const fixtureConfig = { + defaultProvider: fixtureId, + providers: { + [fixtureId]: { + adapter: "openai-responses", + baseUrl: CODEX_FORWARD_BASE_URL, + authMode: "forward", + }, + }, + plaintextV2AgentMessages: true, + } as OcxConfig; + + try { + for (const inboundTransport of [undefined, "websocket"] as const) { + globalThis.fetch = (async () => Response.json(overLimitResponsePayload())) as typeof fetch; + const response = await handleResponses( + collaborationRequest({ model: `${fixtureId}/${fixtureModel}` }), + fixtureConfig, + { model: "", provider: "" }, + inboundTransport + ? { inboundWire: "responses", inboundTransport } + : undefined, + ); + const clientBody = await response.text(); + + expect(response.status).toBe(502); + expect(response.headers.get("content-type")).toContain("application/json"); + expect(clientBody).toContain(PLAINTEXT_V2_AGENT_MESSAGE_RESTORE_OVERFLOW_MESSAGE); + expect(clientBody).not.toContain(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(clientBody).not.toContain("start_delegated_task"); + expect(clientBody).not.toContain("data: [DONE]"); + } + } finally { + const index = mutableRegistry.findIndex(entry => entry.id === fixtureId); + if (index >= 0) mutableRegistry.splice(index, 1); + } + }); + + test("rejects an over-limit JSON response and does not retain it for continuation", async () => { + const sentBodies: string[] = []; + let requestIndex = 0; + globalThis.fetch = (async (_input: RequestInfo | URL, init?: RequestInit) => { + sentBodies.push(typeof init?.body === "string" ? init.body : ""); + requestIndex += 1; + const payload = requestIndex === 1 + ? overLimitResponsePayload() + : { id: "resp-after-overflow", status: "completed", output: [] }; + return Response.json(payload); + }) as typeof fetch; + + const first = await handleResponses( + collaborationRequest(), + config(true), + { model: "", provider: "" }, + ); + const firstBody = await first.text(); + expect(first.status).toBe(502); + expect(firstBody).toContain(PLAINTEXT_V2_AGENT_MESSAGE_RESTORE_OVERFLOW_MESSAGE); + expect(firstBody).not.toContain(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(firstBody).not.toContain("start_delegated_task"); + + const second = await handleResponses( + collaborationRequest({ + previousResponseId: "resp-plaintext-v2-overflow", + input: [{ + type: "message", + role: "user", + content: [{ type: "input_text", text: "continue" }], + }], + }), + config(false), + { model: "", provider: "" }, + ); + const secondBody = await second.text(); + expect({ status: second.status, body: secondBody, sends: sentBodies.length }).toEqual({ + status: 400, + body: expect.stringContaining("continuation state is unavailable or expired"), + sends: 1, + }); + }); + + test.each([undefined, "websocket"] as const)("stores the client namespace across an option change on %s", async inboundTransport => { + const sentBodies: string[] = []; + let requestIndex = 0; + globalThis.fetch = (async (_input: RequestInfo | URL, init?: RequestInit) => { + sentBodies.push(typeof init?.body === "string" ? init.body : ""); + requestIndex += 1; + const payload = requestIndex === 1 + ? completedResponsePayload("resp-toggle-plaintext-v2") + : { id: "resp-after-toggle", status: "completed", output: [] }; + return new Response(JSON.stringify(payload), { + status: 200, + headers: { "content-type": "application/json" }, + }); + }) as typeof fetch; + + const first = await handleResponses( + collaborationRequest(), + config(true), + { model: "", provider: "" }, + { inboundWire: "responses", inboundTransport }, + ); + await first.text(); + const second = await handleResponses( + collaborationRequest({ + previousResponseId: "resp-toggle-plaintext-v2", + input: [{ type: "function_call_output", call_id: "call-spawn", output: "done" }], + }), + config(false), + { model: "", provider: "" }, + { inboundWire: "responses", inboundTransport }, + ); + await second.text(); + + const replay = JSON.parse(sentBodies[1]!) as { input: Array> }; + const replayedCall = replay.input.find(item => item.type === "function_call"); + expect(replayedCall?.namespace).toBe("collaboration"); + expect(JSON.stringify(replay)).not.toContain(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + }); +}); + + +test("plaintext startup warning requires explicit opt-in and names retention", () => { + const warnings: string[] = []; + const original = console.warn; + console.warn = (...args: unknown[]) => { warnings.push(args.join(" ")); }; + try { + warnPlaintextV2AgentMessagesStartup({}); + warnPlaintextV2AgentMessagesStartup({ plaintextV2AgentMessages: false }); + expect(warnings).toEqual([]); + warnPlaintextV2AgentMessagesStartup({ plaintextV2AgentMessages: true }); + expect(warnings.join(" ")).toContain("Codex history"); + expect(warnings.join(" ")).toContain("HTTPS"); + expect(warnings.join(" ")).toContain("local response/debug state"); + } finally { console.warn = original; } +}); + +for (const streamMode of ["legacy-tee", "eager-relay"] as const) { + for (const refusal of ["malformed", "unknown-alias", "conflicting-binding"] as const) { + test(`${streamMode} ${refusal} is refused without caching or retry`, async () => { + let sends = 0; + const sent: string[] = []; + globalThis.fetch = (async (_url: RequestInfo | URL, init?: RequestInit) => { + sends += 1; + sent.push(String(init?.body)); + if (sends > 1) return new Response(JSON.stringify({ id: "after-refusal", status: "completed", output: [] }), { headers: { "content-type": "application/json" } }); + const completed = completedResponsePayload("refused-plaintext"); + const first = { type: "response.output_item.added", output_index: 0, item: completed.output[0] }; + const invalid = refusal === "malformed" ? "{malformed" + : JSON.stringify({ type: "response.output_item.done", output_index: 0, item: { + ...completed.output[0], name: refusal === "unknown-alias" ? "unknown_private" : "deliver_delegated_message", + } }); + return new Response(`data: ${JSON.stringify(first)}\n\ndata: ${invalid}\n\ndata: ${JSON.stringify({ type: "response.completed", response: completed })}\n\ndata: [DONE]\n\n`, { headers: { "content-type": "text/event-stream" } }); + }) as typeof fetch; + const response = await handleResponses(collaborationRequest(), config(true, false, streamMode), { model: "", provider: "" }); + const text = await response.text(); + expect(text).toContain("response.failed"); + expect(text).not.toContain(PLAINTEXT_V2_COLLABORATION_NAMESPACE); + expect(text).not.toContain("start_delegated_task"); + expect(sends).toBe(1); + const next = await handleResponses(collaborationRequest({ previousResponseId: "refused-plaintext", input: [{ type: "message", role: "user", content: [{ type: "input_text", text: "next" }] }] }), config(false), { model: "", provider: "" }); + expect(next.status).toBe(400); + expect(await next.text()).toContain("previous_response_not_found"); + expect(sent).toHaveLength(1); + expect(sends).toBe(1); + }); + } +} + +test("malformed bounded JSON is a single-attempt 502", async () => { + let sends = 0; + globalThis.fetch = (async () => { sends += 1; return new Response("{malformed", { headers: { "content-type": "application/json" } }); }) as typeof fetch; + const response = await handleResponses(collaborationRequest(), config(true), { model: "", provider: "" }); + expect(response.status).toBe(502); + expect(await response.text()).not.toContain("{malformed"); + expect(sends).toBe(1); +}); + +test("cross-coordinate namespace conflict cannot publish continuation", async () => { + let sends = 0; + globalThis.fetch = (async () => { + sends += 1; + const events = [ + { type: "response.output_item.added", output_index: 0, item: { type: "function_call", id: "fc1", call_id: "c1", name: "start_delegated_task", arguments: "" } }, + { type: "response.function_call_arguments.done", item_id: "fc1", namespace: PLAINTEXT_V2_COLLABORATION_NAMESPACE, name: "start_delegated_task", arguments: "{}" }, + { type: "response.completed", response: { id: "refused-coordinates", status: "completed", output: [{ type: "function_call", call_id: "c1", namespace: "foreign", name: "spawn_agent", arguments: "{}" }] } }, + ]; + return new Response(events.map(event => `data: ${JSON.stringify(event)}\n\n`).join("") + "data: [DONE]\n\n", { headers: { "content-type": "text/event-stream" } }); + }) as typeof fetch; + const response = await handleResponses(collaborationRequest(), config(true, false, "eager-relay"), { model: "", provider: "" }); + expect(await response.text()).toContain("response.failed"); + const next = await handleResponses(collaborationRequest({ previousResponseId: "refused-coordinates" }), config(false), { model: "", provider: "" }); + expect(next.status).toBe(400); + expect(sends).toBe(1); +}); + + +test("concurrent native requests do not share plaintext alias metadata", async () => { + const pending: Array<{ enabled: boolean; resolve: (value: Response) => void }> = []; + let bothReady!: () => void; + const ready = new Promise(resolve => { bothReady = resolve; }); + globalThis.fetch = ((_url, init) => { + const body = JSON.parse(String(init?.body)); + const enabled = body.tools.some((tool: { name: string }) => tool.name === PLAINTEXT_V2_COLLABORATION_NAMESPACE); + const response = new Promise(resolve => { pending.push({ enabled, resolve }); }); + if (pending.length === 2) bothReady(); + return response; + }) as typeof fetch; + const first = handleResponses(collaborationRequest(), config(true), { model: "", provider: "" }); + const second = handleResponses(collaborationRequest(), config(false), { model: "", provider: "" }); + await ready; + for (const request of [...pending].reverse()) { + const payload = completedResponsePayload(request.enabled ? "concurrent-enabled" : "concurrent-disabled"); + if (!request.enabled) payload.output[0]!.namespace = "foreign"; + request.resolve(Response.json(payload)); + } + const [enabledResponse, disabledResponse] = await Promise.all([first, second]); + const enabledText = await enabledResponse.text(); + const disabledText = await disabledResponse.text(); + expect(enabledText).toContain('"namespace":"collaboration"'); + expect(enabledText).toContain('"name":"spawn_agent"'); + expect(enabledText).not.toContain("start_delegated_task"); + expect(disabledText).toContain('"namespace":"foreign"'); + expect(disabledText).toContain('"name":"start_delegated_task"'); + expect(pending).toHaveLength(2); +}); diff --git a/tests/server/server-combo-failover-e2e.test.ts b/tests/server/server-combo-failover-e2e.test.ts index 6f649c52d3..436f62a0ff 100644 --- a/tests/server/server-combo-failover-e2e.test.ts +++ b/tests/server/server-combo-failover-e2e.test.ts @@ -1182,6 +1182,88 @@ describe("server combo failover 030 activation matrix", () => { expectMappedReceipt(hydrated[0]!); }); + test("adaptive combo normalizes unknown and empty target capability before the upstream wire", async () => { + const bodies: Array> = []; + const upstream = serve(async request => { + bodies.push(await request.json() as Record); + return chatSuccess("normalized", "m1"); + }); + const request = { + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", + thinking_budget: 8192, + thinking: { type: "enabled" }, + }; + + const unknownResponse = await post( + comboConfig( + { a: provider("openai-chat", baseUrl(upstream), "key-a") }, + undefined, + { reasoningEffortMode: "adaptive" }, + ), + request, + ); + expect(unknownResponse.status).toBe(200); + + const emptyResponse = await post( + comboConfig( + { a: provider("openai-chat", baseUrl(upstream), "key-a", { reasoningEfforts: [] }) }, + undefined, + { reasoningEffortMode: "adaptive" }, + ), + request, + ); + expect(emptyResponse.status).toBe(200); + + const knownResponse = await post( + comboConfig( + { a: provider("openai-chat", baseUrl(upstream), "key-a", { + reasoningEfforts: ["low", "medium", "high", "xhigh"], + }) }, + undefined, + { reasoningEffortMode: "adaptive" }, + ), + request, + ); + expect(knownResponse.status).toBe(200); + + expect(bodies).toHaveLength(3); + for (const body of bodies.slice(0, 2)) { + expect(body).not.toHaveProperty("reasoning_effort"); + expect(body).not.toHaveProperty("thinking_budget"); + expect(body).not.toHaveProperty("thinking"); + } + expect(bodies[2]!.reasoning_effort).toBe("xhigh"); + }); + + test("adaptive Responses combo preserves summary after removing unsupported controls", async () => { + const bodies: Array> = []; + const upstream = serve(async request => { + bodies.push(await request.json() as Record); + return Response.json(responsesSuccess("normalized", "m1")); + }); + for (const reasoningEfforts of [undefined, [], ["high", "xhigh"]]) { + const response = await post(comboConfig({ + a: provider("openai-responses", baseUrl(upstream), "key-a", { + ...(reasoningEfforts === undefined ? {} : { reasoningEfforts }), + modelSupportsReasoningSummaries: { m1: true }, + }), + }, undefined, { reasoningEffortMode: "adaptive" }), { + reasoning: { effort: "xhigh", summary: "concise" }, + reasoning_effort: "xhigh", thinking_budget: 8192, thinking: { type: "enabled" }, + }); + expect(response.status).toBe(200); + } + expect(bodies).toHaveLength(3); + for (const body of bodies.slice(0, 2)) { + expect(body.reasoning).toEqual({ summary: "concise" }); + expect(body).not.toHaveProperty("reasoning_effort"); + expect(body).not.toHaveProperty("thinking_budget"); + expect(body).not.toHaveProperty("thinking"); + } + expect(bodies[2]!.reasoning).toMatchObject({ effort: "xhigh", summary: "concise" }); + }); + test("all-target exhaustion promotes the final attempt reasoning wire to the logical row", async () => { const a = serve(() => Response.json({ error: { message: "first overloaded" } }, { status: 503 })); const b = serve(() => Response.json({ error: { message: "last overloaded" } }, { status: 503 })); @@ -3937,3 +4019,35 @@ describe("combo compact failover", () => { expect(await response.text()).toContain("empty summary"); }); }); + + +describe("thinking-summary defaults follow the serving combo route", () => { + for (const firstVisible of [true, false]) for (const summary of [undefined, "none", "auto"]) { + test(`fallback from ${firstVisible} with summary=${summary}`, async () => { + const observed: Array<[string, boolean | undefined]> = []; + customRunTurn = async (parsed, _incoming, emit) => { + observed.push([parsed.modelId, parsed.options.hideThinkingSummary]); + if (parsed.modelId === "m1") { + emit({ type: "error", message: "provider unavailable", status: 503, retryable: true }); + return; + } + emit({ type: "thinking_delta", thinking: "Actual provider summary" }); + emit({ type: "text_delta", text: "Final fallback answer" }); + emit({ type: "done" }); + }; + const config = comboConfig({ + a: provider("test-run-turn", "https://a.test/v1", "key-a", { showThinkingSummary: firstVisible }), + b: provider("test-run-turn", "https://b.test/v1", "key-b", { showThinkingSummary: !firstVisible }), + }); + const response = await post(config, { ...(summary ? { reasoning: { summary } } : {}) }); + const output = JSON.stringify(await response.json()); + expect(response.status).toBe(200); + expect(observed).toEqual([ + ["m1", summary === "none" || (!summary && !firstVisible)], + ["m2", summary === "none" || (!summary && firstVisible)], + ]); + expect(output.includes("Actual provider summary")).toBe(summary === "auto" || (!summary && !firstVisible)); + expect(output).toContain("Final fallback answer"); + }); + } +}); diff --git a/tests/server/server-xai-chat-reasoning-streaming.test.ts b/tests/server/server-xai-chat-reasoning-streaming.test.ts index 23da211fd0..0b1db50d0d 100644 --- a/tests/server/server-xai-chat-reasoning-streaming.test.ts +++ b/tests/server/server-xai-chat-reasoning-streaming.test.ts @@ -151,7 +151,7 @@ describe("xAI OAuth Chat reasoning streaming", () => { let received = ""; await Promise.race([ (async () => { - while (!received.includes("response.reasoning_summary_text.delta")) { + while (!received.includes("response.reasoning_text.delta")) { const chunk = await reader!.read(); if (chunk.done) throw new Error("stream ended before the first xAI reasoning delta"); received += decoder.decode(chunk.value, { stream: true }); @@ -187,7 +187,7 @@ describe("xAI OAuth Chat reasoning streaming", () => { if (chunk.done) break; received += decoder.decode(chunk.value, { stream: true }); } - const reasoningIndex = received.indexOf("response.reasoning_summary_text.delta"); + const reasoningIndex = received.indexOf("response.reasoning_text.delta"); const contentIndex = received.indexOf("response.output_text.delta"); const completedIndex = received.indexOf("response.completed"); expect(reasoningIndex).toBeGreaterThanOrEqual(0);