Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs-site/src/content/docs/guides/codex-app-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,7 +105,7 @@ metadata instead of an older-template approximation.
| OpenAI (API key) | Exactly eight namespaced rows: `gpt-5.5`, `gpt-5.6`, Sol/Terra/Luna, and the three `*-pro` virtual ids (1,050,000 context; 922,000 max input for all eight) |
| OpenRouter | `openrouter/openai/gpt-5.6-sol`, `openrouter/openai/gpt-5.6-terra`, `openrouter/openai/gpt-5.6-luna` (1,050,000) |
| Cursor | Static fallback includes `cursor/gpt-5.6-sol`, `cursor/gpt-5.6-terra`, and `cursor/gpt-5.6-luna` (1,000,000), plus `cursor/grok-4.5` and `cursor/grok-4.5-fast` (500,000); live account discovery decides which remain visible. |
| xAI | Live discovery is authoritative; the fallback catalog defaults to `xai/grok-4.5` with a 500,000-token window and `low` / `medium` / `high` reasoning controls. |
| xAI | Live discovery is authoritative. The fallback catalog includes `xai/grok-4.6` and defaults to `xai/grok-4.5`; both have 500,000-token windows. Grok 4.6 exposes `low` / `medium` / `high` / `xhigh` (upstream default: `high`), while Grok 4.5 stops at `high`. |

The pinned GPT-5.6 entries preserve the exact upstream ladder. Sol and Terra expose `low` through
`ultra`; Luna stops at `max`. Sol defaults to `low`, while Terra and Luna default to `medium`.
Expand Down
2 changes: 1 addition & 1 deletion docs-site/src/content/docs/ko/guides/codex-app-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ GPT-5.6에만 사용합니다. 오래된 템플릿으로 근사하지 않고 모
| OpenAI(API key) | 정확히 여덟 개의 네임스페이스 행: `gpt-5.5`, `gpt-5.6`, Sol/Terra/Luna, 그리고 세 개의 `*-pro` 가상 id (모두 컨텍스트 1,050,000; 최대 입력 922,000) |
| OpenRouter | `openrouter/openai/gpt-5.6-sol`, `openrouter/openai/gpt-5.6-terra`, `openrouter/openai/gpt-5.6-luna` (1,050,000) |
| Cursor | 정적 폴백에는 `cursor/gpt-5.6-sol`, `cursor/gpt-5.6-terra`, `cursor/gpt-5.6-luna` (1,000,000)와 `cursor/grok-4.5`, `cursor/grok-4.5-fast` (500,000)가 들어갑니다. 실시간 계정 탐색이 어떤 항목을 계속 보일지 정합니다. |
| xAI | 실시간 탐색이 기준입니다. 폴백 카탈로그의 기본값은 `xai/grok-4.5`이고, 컨텍스트 500,000과 `low` / `medium` / `high` 추론 제어를 제공합니다. |
| xAI | 실시간 탐색이 기준입니다. 폴백 카탈로그에는 `xai/grok-4.6`이 포함되며 기본값은 `xai/grok-4.5`입니다. 두 모델 모두 컨텍스트 창은 500,000입니다. Grok 4.6은 `low` / `medium` / `high` / `xhigh`(업스트림 기본값: `high`)를 제공하고, Grok 4.5는 `high`까지만 제공합니다. |

고정된 GPT-5.6 항목은 업스트림 ladder를 그대로 보존합니다. Sol과 Terra는 `low`부터 `ultra`까지 노출하고,
Luna는 `max`에서 멈춥니다. Sol의 기본값은 `low`이고, Terra와 Luna의 기본값은 `medium`입니다. `ultra`는
Expand Down
2 changes: 1 addition & 1 deletion docs-site/src/content/docs/ru/guides/codex-app-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,7 +90,7 @@ per-model identity и метаданные вместо приближения
| OpenAI (API key) | Ровно восемь namespaced-строк: `gpt-5.5`, `gpt-5.6`, Sol/Terra/Luna и три виртуальных id `*-pro` (контекст 1,050,000; максимум входа 922,000 у всех восьми) |
| OpenRouter | `openrouter/openai/gpt-5.6-sol`, `openrouter/openai/gpt-5.6-terra`, `openrouter/openai/gpt-5.6-luna` (1,050,000) |
| Cursor | Статический fallback включает `cursor/gpt-5.6-sol`, `cursor/gpt-5.6-terra` и `cursor/gpt-5.6-luna` (1,000,000), а также `cursor/grok-4.5` и `cursor/grok-4.5-fast` (500,000); какие из них останутся видимыми, решает live-discovery аккаунта. |
| xAI | Live-discovery авторитетно; fallback-каталог по умолчанию содержит `xai/grok-4.5` с окном 500,000 токенов и reasoning-control `low` / `medium` / `high`. |
| xAI | Live-discovery авторитетно. Fallback-каталог включает `xai/grok-4.6`, а моделью по умолчанию остаётся `xai/grok-4.5`; у обеих окно 500,000 токенов. Grok 4.6 поддерживает `low` / `medium` / `high` / `xhigh` (upstream-default: `high`), а Grok 4.5 — только до `high`. |

Закреплённые записи GPT-5.6 сохраняют точную upstream-лестницу. Sol и Terra дают диапазон от
`low` до `ultra`; у Luna верхняя ступень — `max`. По умолчанию у Sol стоит `low`, а у Terra и
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -70,7 +70,7 @@ visibility = "list"
| OpenAI(API key) | 恰好八个命名空间行:`gpt-5.5`、`gpt-5.6`、Sol/Terra/Luna,以及三个 `*-pro` 虚拟 id(八个条目均为 1,050,000 context / 922,000 max input) |
| OpenRouter | `openrouter/openai/gpt-5.6-sol`、`openrouter/openai/gpt-5.6-terra`、`openrouter/openai/gpt-5.6-luna`(1,050,000) |
| Cursor | 静态回退包含 `cursor/gpt-5.6-sol`、`cursor/gpt-5.6-terra`、`cursor/gpt-5.6-luna`(1,000,000),以及 `cursor/grok-4.5` 和 `cursor/grok-4.5-fast`(500,000);实时账户发现会决定最终哪些条目仍然可见。 |
| xAI | 实时发现具有权威性;回退目录默认使用 `xai/grok-4.5`,上下文窗口为 500,000,并提供 `low` / `medium` / `high` reasoning 控制。 |
| xAI | 实时发现具有权威性。回退目录包含 `xai/grok-4.6`,默认模型仍为 `xai/grok-4.5`;两者的上下文窗口均为 500,000。Grok 4.6 提供 `low` / `medium` / `high` / `xhigh`(上游默认值为 `high`),Grok 4.5 最高为 `high`。 |

固定的 GPT-5.6 条目保留了精确的上游阶梯。Sol 和 Terra 暴露从 `low` 到 `ultra` 的档位;Luna 只到 `max`。Sol 默认是 `low`,Terra 和 Luna 默认是 `medium`。`ultra` 是面向客户端的最大 reasoning 加主动委派选项,在后端会以 `max` 传入。选择器里的一个条目只表示目录已经准备好:关联的账户或 API key 仍然必须有权使用该模型。

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -92,7 +92,7 @@ GPT-5.6,以便提供每個模型真實的身份和後設資料,而不是套
| OpenAI(API key) | 恰好八個帶名稱空間的列:`gpt-5.5`、`gpt-5.6`、Sol/Terra/Luna 與三個 `*-pro` 虛擬 id(全部八個都是 1,050,000 context;922,000 max input) |
| OpenRouter | `openrouter/openai/gpt-5.6-sol`、`openrouter/openai/gpt-5.6-terra`、`openrouter/openai/gpt-5.6-luna`(1,050,000) |
| Cursor | 靜態回退目錄包含 `cursor/gpt-5.6-sol`、`cursor/gpt-5.6-terra`、`cursor/gpt-5.6-luna`(1,000,000),以及 `cursor/grok-4.5`、`cursor/grok-4.5-fast`(500,000);帳號的即時發現結果決定最終顯示哪些模型。 |
| xAI | 以即時發現結果為準;回退目錄預設使用 `xai/grok-4.5`,視窗為 500,000 token,並提供 `low` / `medium` / `high` reasoning 控制。 |
| xAI | 以即時發現結果為準。回退目錄包含 `xai/grok-4.6`,預設模型仍為 `xai/grok-4.5`;兩者的 context window 均為 500,000。Grok 4.6 提供 `low` / `medium` / `high` / `xhigh`(上游預設值為 `high`),Grok 4.5 最高為 `high`。 |

固定的 GPT-5.6 條目會保留精確的上游 reasoning 階梯。Sol 和 Terra 從 `low` 到 `ultra`,Luna
最高到 `max`。Sol 預設使用 `low`,Terra 和 Luna 預設使用 `medium`。`ultra` 是用戶端側的
Expand Down
11 changes: 7 additions & 4 deletions src/providers/registry.ts
Original file line number Diff line number Diff line change
Expand Up @@ -950,8 +950,8 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [
// devlog/model_update/260709_model_refresh/001_xai_lineup.md.
// grok-4.20-multi-agent-0309 is intentionally absent: the OAuth chat-completions
// transport returns 400 ("Multi Agent requests are not allowed on chat completions").
// 260813: grok-4.6 added per the new docs.x.ai/developers/grok-4-6 page; specs mirrored
// from grok-4.5 until the official capability/pricing tables settle.
// 260813: grok-4.6 added per docs.x.ai/developers/grok-4-6. Context/vision still match
// grok-4.5; the reasoning ladder does not — 4.6 adds the documented xhigh rung.
models: ["grok-4.6", "grok-4.5", "grok-4.3", "grok-4.20-0309-reasoning", "grok-4.20-0309-non-reasoning", "grok-build-0.1", "grok-composer-2.5-fast"],
defaultModel: "grok-4.5",
// Vision lineup per docs.x.ai model-capabilities/images/understanding: the grok-4.x chat
Expand All @@ -973,8 +973,11 @@ export const PROVIDER_REGISTRY: readonly ProviderRegistryEntry[] = [
// (docs.x.ai prompt-caching/multi-turn, verified 2026-07-13 — devlog/_plan/260713_grok_caching).
// Models that never emit reasoning simply have no thinking parts to replay (no-op).
preserveReasoningContentModels: ["grok-4.6", "grok-4.5", "grok-4.3", "grok-4.20-0309-reasoning"],
// grok-4.5 reasoning is always-on with low/medium/high control (no off tier upstream).
modelReasoningEfforts: { "grok-4.6": ["low", "medium", "high"], "grok-4.5": ["low", "medium", "high"] },
// grok-4.5 reasoning is always-on with low/medium/high (no off tier, no xhigh).
// grok-4.6 adds xhigh per docs.x.ai/developers/model-capabilities/text/reasoning;
// xAI documents high as the upstream default.
modelReasoningEfforts: { "grok-4.6": ["low", "medium", "high", "xhigh"], "grok-4.5": ["low", "medium", "high"] },
modelDefaultReasoningEfforts: { "grok-4.6": "high" },
modelContextWindows: {
"grok-4.6": 500_000,
"grok-4.5": 500_000,
Expand Down
1 change: 1 addition & 0 deletions tests/catalog-vision-sidecar-modalities.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -156,6 +156,7 @@ describe("vision-capable provider models feed combo modalities", () => {
test("xAI grok chat models declare image input in the registry", () => {
const xai = PROVIDER_REGISTRY.find(entry => entry.id === "xai");
for (const model of [
"grok-4.6",
"grok-4.5",
"grok-4.3",
"grok-4.20-0309-reasoning",
Expand Down
2 changes: 2 additions & 0 deletions tests/effort-policy.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -219,6 +219,8 @@ describe("supportedLadderFor (real routeModel routes)", () => {
} as Partial<OcxConfig>);
const route = routeModel(config, "xai/grok-4.5");
expect(supportedLadderFor(route)).toEqual(["low", "medium", "high"]);
const grok46 = routeModel(config, "xai/grok-4.6");
expect(supportedLadderFor(grok46)).toEqual(["low", "medium", "high", "xhigh"]);
const noReasoning = routeModel(config, "xai/grok-composer-2.5-fast");
expect(supportedLadderFor(noReasoning)).toEqual([]);
});
Expand Down
37 changes: 37 additions & 0 deletions tests/oauth-provider-reconcile.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@ import { join } from "node:path";
import { loadConfig } from "../src/config";
import { OAUTH_PROVIDERS, reconcileOAuthProviders, upsertOAuthProvider } from "../src/oauth";
import { getCredential, saveCredential } from "../src/oauth/store";
import { routeModel } from "../src/router";
import type { OcxConfig } from "../src/types";

const originalHome = process.env.OPENCODEX_HOME;
Expand Down Expand Up @@ -189,4 +190,40 @@ describe("OAuth provider reconciliation", () => {
reconcileOAuthProviders(config);
expect(config.providers.kimi.requiresReasoningPlaceholderModels).toEqual([]);
});

test("refreshes Grok 4.6 levels while runtime fills the default without overwriting user intent", () => {
const home = mkdtempSync(join(tmpdir(), "ocx-grok-46-reconcile-"));
homes.push(home);
process.env.OPENCODEX_HOME = home;
const staleXai = structuredClone(OAUTH_PROVIDERS.xai.providerConfig);
staleXai.modelReasoningEfforts = {
"grok-4.6": ["low", "medium", "high"],
"grok-4.5": ["low", "medium", "high"],
};
delete staleXai.modelDefaultReasoningEfforts;
const config = {
port: 10100,
defaultProvider: "xai",
providers: {
xai: {
...staleXai,
note: "user-owned-note",
},
},
} satisfies OcxConfig;

expect(reconcileOAuthProviders(config)).toBe(true);
expect(config.providers.xai.modelReasoningEfforts?.["grok-4.6"])
.toEqual(["low", "medium", "high", "xhigh"]);
expect(config.providers.xai.modelDefaultReasoningEfforts).toBeUndefined();
expect(routeModel(config, "xai/grok-4.6").provider.modelDefaultReasoningEfforts?.["grok-4.6"])
.toBe("high");
expect(config.providers.xai.note).toBe("user-owned-note");
expect(reconcileOAuthProviders(config)).toBe(false);

config.providers.xai.modelDefaultReasoningEfforts = { "grok-4.6": "medium" };
expect(reconcileOAuthProviders(config)).toBe(false);
expect(routeModel(config, "xai/grok-4.6").provider.modelDefaultReasoningEfforts?.["grok-4.6"])
.toBe("medium");
});
});
21 changes: 21 additions & 0 deletions tests/provider-registry-parity.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -670,9 +670,14 @@ describe("provider registry parity", () => {
}
expect(OAUTH_PROVIDERS.xai.providerConfig.defaultModel).toBe("grok-4.5");
expect(OAUTH_PROVIDERS.xai.providerConfig.liveModels).toBe(true);
expect(OAUTH_PROVIDERS.xai.providerConfig.models).toContain("grok-4.6");
expect(OAUTH_PROVIDERS.xai.providerConfig.models).toContain("grok-4.5");
expect(OAUTH_PROVIDERS.xai.providerConfig.modelContextWindows?.["grok-4.6"]).toBe(500_000);
expect(OAUTH_PROVIDERS.xai.providerConfig.modelContextWindows?.["grok-4.5"]).toBe(500_000);
expect(OAUTH_PROVIDERS.xai.providerConfig.modelReasoningEfforts?.["grok-4.6"]).toEqual(["low", "medium", "high", "xhigh"]);
expect(OAUTH_PROVIDERS.xai.providerConfig.modelReasoningEfforts?.["grok-4.5"]).toEqual(["low", "medium", "high"]);
expect(OAUTH_PROVIDERS.xai.providerConfig.modelDefaultReasoningEfforts).toEqual({ "grok-4.6": "high" });
expect(OAUTH_PROVIDERS.xai.providerConfig.modelReasoningEffortMap).toBeUndefined();
expect(OAUTH_PROVIDERS.xai.providerConfig.noVisionModels).toContain("grok-build-0.1");
const antigravityRegistry = PROVIDER_REGISTRY.find(entry => entry.id === "google-antigravity");
expect(antigravityRegistry?.liveModels).toBe(true);
Expand Down Expand Up @@ -847,6 +852,22 @@ describe("provider registry parity", () => {
.toEqual(["low", "medium", "high", "max", "ultra"]);
});

test("grok-4.6 advertises the documented xhigh rung from the xai registry seed", () => {
const xai = PROVIDER_REGISTRY.find(entry => entry.id === "xai");
const seed = providerConfigSeed(xai!);
const model = applyProviderConfigHints("xai", seed, { id: "grok-4.6", provider: "xai" });
expect(model.contextWindow).toBe(500_000);
expect(model.reasoningEfforts).toEqual(["low", "medium", "high", "xhigh"]);

const entries = buildCatalogEntries(nativeTemplate() as never, [], [model]);
const entry = entries.find(e => e.slug === "xai/grok-4.6");
expect(entry).toBeTruthy();
expect(entry?.context_window).toBe(500_000);
expect((entry?.supported_reasoning_levels as { effort: string }[]).map(l => l.effort))
.toEqual(["low", "medium", "high", "xhigh", "max", "ultra"]);
expect(entry?.default_reasoning_level).toBe("high");
});

// The id-list assertion above only proves the preset exists. Pin the contract a user actually
// depends on: which endpoint the key is sent to, which adapter parses the stream, and that the
// vendor-namespaced seed models survive into a real catalog entry.
Expand Down
57 changes: 57 additions & 0 deletions tests/reasoning-effort.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -99,6 +99,63 @@ describe("provider-specific reasoning effort mapping", () => {
});
});

test("xAI grok-4.6 forwards xhigh while grok-4.5 still clamps it to high", () => {
const config: OcxConfig = {
port: 10100,
defaultProvider: "xai",
providers: {
xai: {
adapter: "openai-chat",
baseUrl: "https://api.x.ai/v1",
apiKey: "key",
},
},
};
const grok46 = routeModel(config, "xai/grok-4.6");
const grok45 = routeModel(config, "xai/grok-4.5");

expect(configuredReasoningEfforts(grok46.provider, grok46.modelId)).toEqual(["low", "medium", "high", "xhigh"]);
expect(configuredReasoningEfforts(grok45.provider, grok45.modelId)).toEqual(["low", "medium", "high"]);

const grok46Xhigh = buildChatRequest(grok46.provider, grok46.modelId, { reasoning: "xhigh" });
const grok46Max = buildChatRequest(grok46.provider, grok46.modelId, { reasoning: "max" });
const grok45Xhigh = buildChatRequest(grok45.provider, grok45.modelId, { reasoning: "xhigh" });

expect(JSON.parse(grok46Xhigh.body).reasoning_effort).toBe("xhigh");
expect(JSON.parse(grok46Max.body).reasoning_effort).toBe("xhigh");
expect(JSON.parse(grok45Xhigh.body).reasoning_effort).toBe("high");
expect(mapReasoningEffort(grok46.provider, grok46.modelId, "xhigh")).toBe("xhigh");
expect(mapReasoningEffort(grok45.provider, grok45.modelId, "xhigh")).toBe("high");
});

test("xAI grok-4.6 preserves an explicit narrower ladder and provider-wide downgrade map", () => {
const config: OcxConfig = {
port: 10100,
defaultProvider: "xai",
providers: {
xai: {
adapter: "openai-chat",
baseUrl: "https://api.x.ai/v1",
authMode: "key",
apiKey: "key",
modelReasoningEfforts: {
"grok-4.6": ["low", "medium", "high"],
"grok-4.5": ["low", "medium", "high"],
},
reasoningEffortMap: { xhigh: "high", max: "high" },
},
},
};
const grok46 = routeModel(config, "xai/grok-4.6");
const grok45 = routeModel(config, "xai/grok-4.5");

expect(configuredReasoningEfforts(grok46.provider, grok46.modelId)).toEqual(["low", "medium", "high"]);
expect(mapReasoningEffort(grok46.provider, grok46.modelId, "xhigh")).toBe("high");
expect(mapReasoningEffort(grok46.provider, grok46.modelId, "max")).toBe("high");
expect(configuredReasoningEfforts(grok45.provider, grok45.modelId)).toEqual(["low", "medium", "high"]);
expect(mapReasoningEffort(grok45.provider, grok45.modelId, "xhigh")).toBe("high");
});

test("Neuralwatt GLM-5.2 sends direct max and preserves reasoning history", () => {
const provider: OcxProviderConfig = {
adapter: "openai-chat",
Expand Down
Loading