fix(acp): pick the reasoning-effort ladder from the selected model - #286
fix(acp): pick the reasoning-effort ladder from the selected model#286Astro-Han wants to merge 3 commits into
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 36e6a05e1f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
36e6a05 to
4242863
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4242863562
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
4242863 to
97dd43d
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 97dd43d578
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
97dd43d to
ca3b71f
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ca3b71f362
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
ca3b71f to
65bb180
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 65bb1807a2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
65bb180 to
8604664
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 860466473f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
8604664 to
eae92ca
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: eae92cad85
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
eae92ca to
3d6d68b
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3d6d68b317
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
3d6d68b to
4e66a56
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4e66a561f3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
2ffb88f to
8a400be
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8a400be590
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
8a400be to
47e3335
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 47e3335117
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
47e3335 to
9cba5f6
Compare
|
@Leeeon233 The PR is ready for review, thanks! |
9cba5f6 to
0902fc2
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0902fc2ee8
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
0902fc2 to
3088326
Compare
The capability cache treated the probed model's thought-level list as global. An adapter can now publish each model's reasoning-effort ladder under the Lody-owned "_meta.lody.modelReasoningEfforts" namespace on the session response (the Grok adapter does), and the capability normalization merges that map with the existing legacy "model[effort]" derivation into modelReasoningEfforts. No cache-version bump: the startup capability refresh re-probes every configured agent unconditionally and every session rewrites its entry, so existing caches pick up the map without a global invalidation that would hide registry/custom selectors until the refresh completes. "summarizeAgentRunConfigCapabilities" and "resolveAgentRunConfigSelection" already consume the map, so MCP callers and dispatch validation get per-model efforts without further change. Ref: LodyAI#149 Model: glm-5.3
3088326 to
5aaddd3
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5aaddd3053
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The composer's reasoning-effort picker rendered whatever thought-level list the capability probe captured, no matter which model was selected: selecting Grok 4.6 while the probe ran under a model with a high/low/max ladder hid "medium" and offered the invalid "max" (LodyAI#149). A new normalizeReasoningEffortSelectors rebuilds the thought-level selector from the per-model ladder the capability source publishes (cached runtime map or static builtin table), reusing probe-time labels and descriptions for values the ladder still offers and synthesizing labels for the rest. It runs in both the selector builders and the selection resolution, so a model moved by authoritative validation re-normalizes against the RESOLVED model, and it skips Codex, which keeps its own hardcoded extended-effort table for now. The map is a required field of the selection input, so the session, draft, and Role composer call sites cannot drop it silently; the Role editor passes its pinned model so its picker and compatibility check agree on the same ladder. The static builtin Grok table gains per-model ladders so the picker is correct before any probe, and resolveConfigOptions now surfaces the cached per-model map alongside configOptions. Model: glm-5.3
…tract Model: glm-5.3
5aaddd3 to
510fe86
Compare
Related issue
Closes #149
Problem / pressure
The reasoning-effort picker showed one flat list — the ladder measured for whatever model was current at probe time — for every model. Selecting Grok 4.6 while the probe had run under a model with a
high/low/maxladder hidmedium/xhighand offeredmax, which xAI rejects with400 Invalid reasoning effort. Grok 4.5 and 4.6 differ the same way (xhighonly on 4.6). The Grok adapter kept one flat ladder and never rebuilt it on model change, and nothing carried a model-independent view to the UI.Summary
model_changednotification intoconfig_option_update, and publishes the per-model view as_meta.lody.modelReasoningEffortson the session response. The root gitlink is bumped after that PR merges, per the runtime-upgrade convention from chore: upgrade built-in AI agent runtimes #201.model[effort]derivation — now scoped to builtin Codex, because Claude uses the same brackets for context windows (opus[1m]).ACP_CAPABILITY_CACHE_VERSIONgoes to 7, because an entry probed before that guard stored a bogus ladder for every agent that spells other variants with the same brackets (a Claude probe left{ opus: ['1m'] }), and the picker would now rebuild that model's ladder from it.summarizeAgentRunConfigCapabilitiesandresolveAgentRunConfigSelectionalready consume the map, so MCP options and dispatch validation pick it up unchanged.normalizeReasoningEffortSelectorsrebuilds the thought-level selector from the selected model's ladder, in the selector builders and again in the selection resolution so it follows an authoritatively resolved model. Applying a whole saved configuration (a recent run config or a Role) that also moves the model passes its effort through unvalidated, because the selectors still describe the outgoing model; the selection resolution re-validates it against the resolved one. It is a fallback, never an override — Codex's hand-maintained extended tiers and Claude's adapter-side rebuild keep their own paths, and a model the map does not cover keeps the probe-time list, since the adapter owns the wire and rejects an unsupported effort with a visible warning.Before / after
high/low/maxxhigh/high/medium/lowlody_session_create_optionsreports one flatreasoningEffortValuesmodels[].reasoningEffortValuesTest plan
packages/acp-extension-grok:npm test— 39 passing (model-switch rebuild, effort validation against the selected model, published meta).apps/cli:npx vitest run tests/acp-capability-normalization.test.ts src/mcp/lody-mcp-server.test.ts src/commands/session.test.ts— passing.packages/components:npx vitest run tests/acp-selector-options.test.ts tests/acp-session-config-selection.test.ts tests/session-config-selection-oscillation.test.tsx tests/agent-role-form.test.ts tests/recent-run-configs.test.ts— passing (the two localStorage failures inrecent-run-configsreproduce on a clean checkout).packages/shared: full run, 995 passing.pnpm typecheck,pnpm lint,pnpm lint:i18n,check:code-collab-imports,check:platform-boundaries,check:public-boundary— pass. Fulltest:ciwas not run; the known environment-related component failures reproduce on a clean checkout without this change.Context handoff
Instructions for reviewing agents
normalizeReasoningEffortSelectorsinpackages/components/src/components/shared/acp-selector-options.ts, and the zod-scoped_meta.lodyread plus Codex-only legacy merge inapps/cli/src/agent/acp-capability-normalization.ts._meta.lody.modelReasoningEffortscontract is documented in AGENTS.md rather than declared in acp-extension-core — say the word if you prefer it in core.Authoring context
_metamust not leak into Lody business code; the adapter's startupsessionConfigtranslation is untouched.mediumrather than the ladder's first (most expensive) entry; an uncovered model keeps the probe-time list and relies on the adapter's rejection warning.configOptionValues, which the resolver does not read, and none publishes the map today; no end-to-end Grok run on a real machine.