Add explicit local runtime profiles - #23
Conversation
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 50aead2b3e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| from .client import Client | ||
|
|
||
| kwargs = dict(provider_kwargs) | ||
| kwargs.setdefault("strict", True) |
There was a problem hiding this comment.
Enforce strict mode for the selected Ollama client
When local_client() selects Ollama and the service fails after the readiness probe, setting strict=True does not fulfil the advertised local-only contract: OllamaProvider.summarize() catches the HTTP/CLI failures and returns the input prompt as a deterministic fallback, while stream() similarly falls back through summarize(). This can silently present fabricated output as a successful local-model response; the Ollama execution paths must honour strict mode before this helper can promise that no deterministic shim is used.
Useful? React with 👍 / 👎.
| for provider in candidates: | ||
| candidate_model = _model_for_provider(provider, model, models) | ||
| status = _probe_local_provider( |
There was a problem hiding this comment.
Reject an unspecified model before reporting readiness
When callers accept the public model=None default and omit models, candidate_model remains None, which the shared probes treat as available regardless of the returned model catalogue. The resulting Ollama client sends requests without a model, while the oMLX constructor silently changes the unverified selection to the literal model "omlx"; therefore local_client() can report successful readiness and then fail on its first request. Require a provider-specific model or deliberately select and return one of the probed models.
Useful? React with 👍 / 👎.
| ## Pending work | ||
|
|
||
| 1. Open the local-runtime profile pull request. | ||
| 2. Run lint, mypy, Python-version test matrix, build checks as appropriate, and fix all branch failures. |
There was a problem hiding this comment.
Restore the complete pending project state
The rewritten Pending work section contains only tasks for this feature branch and removes existing unresolved project work such as the missing Claude raw passthrough and Gemini/Grok chat surfaces, so STATUS.md is no longer a complete current snapshot. Restore those still-applicable items and an explicit recent-changes/next-steps account rather than replacing the repository-wide state with this branch's checklist.
AGENTS.md reference: AGENTS.md:L86-L88
Useful? React with 👍 / 👎.
| ## Architecture | ||
|
|
||
| The primary application surface is `modelito.Client`, backed by registered provider adapters. Providers implement a small common interface and may expose richer raw/streaming capabilities when available. Local runtime readiness is separated from provider construction: |
There was a problem hiding this comment.
Retain inline diagrams in STATUS
This rewrite deletes the existing inline SVG architecture and request-flow diagrams even though the replacement still describes a meaningful multi-stage architecture and adds another local-runtime selection flow. Keep useful diagrams in STATUS.md as required instead of reducing the architecture section to prose alone.
AGENTS.md reference: AGENTS.md:L90-L92
Useful? React with 👍 / 👎.
Adds an explicit local-only selection layer for applications that need predictable deployment semantics without changing the established
Client(provider="auto")contract.Runtime policy
portable: Ollama as the cross-platform path.mac-performance: BaseRT → vllm-mlx → oMLX → Ollama on Apple Silicon.auto: resolves to the Mac-oriented profile on Apple Silicon andportableelsewhere.prefer=override so measured machine/workload results can reorder candidates.select_local_runtime()for read-only readiness/selection.local_client()for strict local-only construction with no hosted or deterministic fallback.The default order is deliberately a curated deployment starting point, not a universal performance ranking. BaseRT, vllm-mlx, oMLX, Ollama, and raw MLX-LM have materially different execution, batching, and caching strategies; the target workload should be benchmarked on the target machine.
Added in this PR
modelito-benchmark-local, covering first-request TTFT, warm-prefix TTFT, context-growth latency, first useful streamed phrase, estimated decode rate, cancellation/stream-close behaviour, and optional process RSS.docs/LOCAL-RUNTIMES.md.STATUS.mdstate, pending tasks, recent changes, next steps, and inline architecture/request-flow diagrams.CI remains intentionally scoped to the repository's established checks. A newly introduced global Black gate was removed because current Black would reformat pre-existing files across the repository, which is a separate formatting migration rather than part of this runtime feature.
No release, tag, or version bump is included.