Skip to content

fix(cli): discover service models and apply reasoning controls - #37

Merged
sizzlecar merged 1 commit into
mainfrom
fix/api-model-reasoning
Sep 17, 2026
Merged

sizzlecar merged 1 commit into
mainfrom
fix/api-model-reasoning

Conversation

@sizzlecar

@sizzlecar sizzlecar commented Sep 17, 2026 •

Copy link
Copy Markdown
Owner

/model currently lists configured profiles even when the connected service exposes different model IDs, and the generic OpenAI adapter does not send a selected reasoning control. Refresh the current service's model API, preserve the resolved endpoint and credentials when selecting its model IDs, and expose typed reasoning controls through the CLI and TUI.

  • /model discovers live service models; /model profiles explicitly selects configured profiles. Empty responses and discovery errors are visible. Existing optional context-capacity discovery and generation-only endpoint fallback remain supported.
  • --reasoning and /reasoning distinguish an omitted default, explicit none, declared effort levels, and binary thinking on/off. Unknown capabilities remain unknown; binary-only models do not acquire invented effort levels. Unsupported adapters reject non-default controls.
  • Selections survive Host reconstruction and session switching in the running TUI. Restart behavior continues to use CLI/configuration; this change does not write the user's base configuration.

Validation on Orchestral 274d975: formatting, workspace check, workspace tests (1,015 passed; 18 existing opt-in tests ignored), all-features Clippy with warnings denied, and the WASM web compile check passed. Focused CLI tests (161 passed; 1 ignored) and OpenAI adapter tests (37 passed) also passed. A local HTTP-fixture PTY test exercised discovery, authenticated model switching, binary reasoning controls on actual adapter requests, and session resume.

Real-model validation used pinned local Qwen3.5-9B Q4_K_M on Metal, separately with F16 and INT8 KV:

  • TUI controls — Ferrum f50080ac, Orchestral 274d975: native PTY sessions started from a configured placeholder, selected the actual API ID with /model, and displayed only service-declared default/on/off reasoning controls. Off/on/default sent false/true/omitted request fields. All six short requests completed with stop, nonempty answers, zero failures and no remaining requests. One initial CLI launch rejected an unsupported attention-policy flag before loading the model; that failure was retained, and successful runs used the supported ferrum.toml setting.
  • Tools and session continuity — Ferrum 51de5a51, Orchestral 274d975: the frozen Rust interop harness passed both formats with a 16,384-token context, 512-token output limit and 8 GiB total runtime budget; health confirmed Metal portable execution and the requested KV storage. Each format made four successful model requests across three separate CLI processes: a genuine file_read exchange and answer, history recall after the source file was deleted, and exact retrieval from 128 records. The first response ended with tool_calls; all three final answers ended with stop. Input/output usage was 2632/26, 2731/35, 2800/36 and 6860/24 in both formats. The initial prompt token IDs, sanitized request body and effective sampling matched exactly across formats. Both servers ended healthy with zero failures, active requests or queued requests; binary hashes remained unchanged and all owned processes were stopped.

Evidence is retained outside the repository under /private/tmp/ferrum-int8-local-validation-20260917/: orchestral-qwen35-9b-274d975-f50080ac-tui/report.json for the six TUI requests, and orchestral-qwen35-9b-274d975-51de5a51-paired/report.json for the tool/session pair, alongside product request bundles, isolated journals, stdout/stderr and health snapshots. These are finite interoperability samples on one model/backend, not performance, context-limit, broad quality or all-model qualification. Only the initial paired input is required to match; generated histories may diverge.

@sizzlecar
sizzlecar marked this pull request as ready for review September 17, 2026 13:01
@sizzlecar
sizzlecar merged commit bf154aa into main Sep 17, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant