Skip to content

feat(sc-codex): gpt-6 model aliases and reasoning effort option (0.14.0) - #116

Draft
ehrvs wants to merge 3 commits into
randlee:developfrom
ehrvs:feat/sc-codex-models-effort
Draft

ehrvs wants to merge 3 commits into
randlee:developfrom
ehrvs:feat/sc-codex-models-effort

Conversation

@ehrvs

@ehrvs ehrvs commented Sep 28, 2026

Copy link
Copy Markdown

Summary

Lets callers pick the Codex model and reasoning effort through sc-codex. For example, "use sc-codex with sol low" runs codex exec --yolo --model gpt-6-sol -c model_reasoning_effort="low" <prompt>.

Package: sc-codex · Type: feat (breaking) · Version: 0.13.0 → 0.14.0

Models

  • The model list is defined once in ai_cli/task_tool.py. The Pydantic model, the argparse help and the JSON schema all come from it, and a test checks that they match.
  • Aliases: sol → gpt-6-sol, astra → gpt-6-astra, luna → gpt-6-luna, terra → gpt-5.6-terra. codex or no model means the default.
  • Full names also accepted: gpt-6-{sol,astra,luna}, gpt-5.6-{sol,terra,luna}, gpt-5.5.
  • Default model: gpt-6-astra.
  • BREAKING: the old aliases are removed (mini, max, codex-mini, codex-max, gpt-5, gpt-5.2, gpt-5.2-codex, gtp-5). An unknown model now fails with an error that lists the valid names.
  • ChatGPT-account fallback: kept but simplified. If Codex rejects the model for a ChatGPT account, the run retries once on gpt-5.5, with effort lowered to xhigh, and the retry is logged.

Reasoning effort

  • New --effort option on sc_codex_task.py and ai_cli run, plus a reasoning_effort field in the JSON payload.
  • Levels: low|medium|high|xhigh|max|ultra, checked per model before launch. The luna models stop at max and gpt-5.5 stops at xhigh. minimal is rejected because no current model supports it.
  • Unset: no -c setting is passed, so ~/.codex/config.toml applies.
  • Command line wins: --model and --effort override the payload, including a payload with an invalid value. The overrides are saved in the payload, so background runs and logs show what actually ran.

Docs / housekeeping

  • Docs: the command doc, SKILL.md, the package README and the ai_cli README are updated. They include how to turn phrases like "sol low" into flags, and a note that max and ultra are efforts, not models.
  • README fixes: the Quick Start used flags that don't exist, and the README still said 0.7.x.
  • Background default: the docs said two different things. They now match the code: on for sc_codex_task.py, off for ai_cli run.
  • CHANGELOG: new 0.14.0 entry, with a note that 0.8.0–0.13.0 were released without entries.
  • Agent registry: the sc-codex agent is added to .claude/agents/registry.yaml.
  • Version: bumped with scripts/set-package-version.py, which also updated the registries.
  • Fixtures: mini becomes luna, and test_parallel_minis.yaml becomes test_parallel_luna.yaml.

Review

  • Skill review of sc-codex against the repo's guidelines, done before implementing.
  • Fable review: no blockers, approved.
  • Codex gpt-6-astra (high effort): one should-fix, the override order.
  • All should-fix items and nits from both reviews are addressed in 2534477.

Testing

  • pytest tests/: 1444 passed, 2 failed. Both also fail on a clean origin/develop in this environment (go not installed; sc-compose render_template can't be imported).
  • tests/test_ai_cli_task_runner.py: 61 tests covering aliases and the default, unknown-model errors, -c present or absent, unsupported effort per model, command-line overrides of the payload (including invalid payloads), the background payload carrying effort, the fallback, and schema/code parity.
  • scripts/validate-all.py: 9/9 pass. scripts/audit-versions.py --verbose: 0 failures.
  • One live smoke run with --model luna --effort low returned output; the log shows gpt-6-luna / low.

Out of scope (follow-ups)

  • Rewrite the sc-codex agent stub and make its responses a fenced-JSON {success, data, error} object
  • shell=True in hook execution (task_runner.py)
  • scripts/sc_shared.py, which appears unused
  • codex-agent skill naming convention
  • sc-launchpad / sc-launch-term still use their own model aliases (codex → gpt-5.6-terra)

Notes for reviewer

  • sol/luna now mean gpt-6. The uncommitted edit in ~/.claude/scripts used gpt-5.6. Use gpt-5.6-sol / gpt-5.6-luna explicitly for the older ones.
  • Installed copy is stale. ~/.claude/scripts has an uncommitted 0.12.0 hand-edit. Reinstall sc-codex after merging to replace it.

🤖 Generated with Claude Code

ehrvs and others added 3 commits September 27, 2026 20:45
- New model catalog (sol/astra/luna/terra/codex aliases, gpt-6 and
  gpt-5.6 slugs, gpt-5.5); default model is gpt-6-astra.
- --effort {low,medium,high,xhigh,max,ultra} on sc_codex_task.py and
  ai_cli run, plus optional reasoning_effort payload/schema field,
  validated per model and passed as -c model_reasoning_effort="...".
  Omitted when unset so ~/.codex/config.toml applies.
- CLI flags override the payload; effort survives background runs and
  is logged as a top-level field.
- ChatGPT-account fallback now retries once on gpt-5.5 with effort
  clamped to xhigh.

BREAKING CHANGE: legacy aliases gpt-5.2-codex, codex-max/max,
codex-mini/mini, gpt-5, gpt-5.2 and gtp-5 are removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bump version via set-package-version.py, add CHANGELOG entry, and
register the sc-codex agent in .claude/agents/registry.yaml.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Apply --model/--effort to the raw payload before validation in
  sc_codex_task.py and ai_cli run, persisting the resolved model so
  background payloads and logs reflect what ran.
- Add cmd_run tests (effort override, unsupported effort, override of
  invalid payload values) and schema nullability parity checks.
- Allow null for optional fields in task_tool.schema.json.
- Docs: effort-vs-model disambiguation, "alongside --json" wording,
  minimal effort note, --runner codex in ai_cli examples, argv-accurate
  -c quoting; CHANGELOG note replaces range heading.
- Rename fixture test_parallel_minis.yaml -> test_parallel_luna.yaml.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant