Skip to content

feat: Model type with extraBody passthrough, portable top_logprobs, transport hardening - #12

Merged
expilu merged 1 commit into
mainfrom
feaure/model-config
Sep 28, 2026
Merged

expilu merged 1 commit into
mainfrom
feaure/model-config

Conversation

@expilu

@expilu expilu commented Sep 28, 2026

Copy link
Copy Markdown
Owner

Summary

Moves everything about how the model is reached and served out of Question into a new reusable Model type, adds an extraBody passthrough for engine-/model-specific request settings, and hardens the transport layer. Breaking to the Question shape, which is expected pre-1.0 (documented in the changeset).

Model type & extraBody

  • New src/types/model.ts: Model { apiBaseUrl, apiKey, model, extraBody? }. Define once, pass to every choice().
  • Question becomes { model: Model, mode?, maxRetries?, timeoutMs?, state, instructions, criteria }.
  • extraBody fields are forwarded verbatim into the chat completions request body — the escape hatch for settings the library does not model:
    • chat_template_kwargs thinking toggles (llama.cpp / vLLM / SGLang, per-model keys)
    • reasoning_effort: 'none' (OpenAI, OpenRouter, Ollama /v1), think: false (Ollama native)
  • Merge semantics (applyExtraBody): a strict reserved set (model, messages, stream, logprobs, top_logprobs, max_tokens, temperature) protects System 1's single-token trick; chat_template_kwargs merges one level deep so the built-in enable_thinking: false default survives and user keys win per key; __proto__/constructor keys are ignored.
  • No provider field: engines ignore unknown body fields, and capability gaps (no logprobs on Ollama /v1, Anthropic) fail loudly instead.

System 1 portability

  • top_logprobs default 50 → 20: the highest portable window (OpenAI and OpenRouter cap at 20, vLLM's server default --max-logprobs is 20). llama.cpp accepts ≤ 50, so 20 is safe everywhere logprobs exist.

Transport hardening

  • Refuses redirects (redirect: 'error') — the Bearer token never follows a 3xx.
  • Response bodies are read under a 10 MB safety cap (readBodyCapped): declared Content-Length over the cap is refused before reading; streamed bodies are capped mid-read. Over-cap failures are not retried.
  • maxRetries is sanitized: NaN (previously an infinite retry loop) falls back to the default, negatives clamp to 0, fractions floor.
  • Malformed entries in top_logprobs (null/missing token) are skipped instead of crashing.
  • Error messages strip query strings from URLs (query-borne credentials can't leak into logs); the endpoint path is joined through the URL API (query strings on the base URL survive); invalid base URLs throw a clear Invalid model.apiBaseUrl error.

CI / release supply chain

  • All GitHub Actions pinned by commit SHA (tags kept as comments; Dependabot-compatible).
  • scripts/release.mjs spawns every subprocess as an argv array via execFileSync — no shell interpolation.

Docs

  • README: usage now builds const model = {...} and passes it to choice(); new "Model and model-specific settings" section; new Compatibility matrix (llama.cpp, vLLM, SGLang, OpenRouter, OpenCode Zen, OpenAI, Gemini, Ollama, Anthropic) with per-engine extraBody cookbook.
  • Changeset (minor, 0.3.0) documenting the breaking shape change, the top_logprobs default change, and the hardening behavior changes.

Testing

  • 98 tests passing; coverage stays at the CI-enforced 100% (lines/functions/branches/statements), including the new apply-extra-body and read-body-capped utils.
  • New wire-level tests: thinking-default override via extraBody, exotic engine keys forwarded, reserved keys ignored, over-cap bodies refused without retrying, non-finite maxRetries handled, URL join/query handling, prototype-smuggling defense.

- Added a new Model type to encapsulate model-specific API details including endpoint, credentials, and extra request settings.
- Updated Question type to include the Model type, allowing for more structured model configuration.
- Enhanced ChatCompletionRequest to support engine-specific fields via an index signature.
- Implemented applyExtraBody utility to merge extra fields into chat completion requests while preserving reserved keys.
- Introduced readBodyCapped function to safely handle response bodies, preventing excessive memory usage from oversized responses.
- Updated generateText function to incorporate new model handling and response body safety checks.
- Added comprehensive tests for new functionalities and ensured existing tests reflect the updated structure.
@expilu
expilu merged commit 5e807d8 into main Sep 28, 2026
5 of 6 checks passed
@github-actions github-actions Bot mentioned this pull request Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant