You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Add explicit portable, mac-performance, and auto local runtime profiles without changing general Client(provider="auto") behaviour.
Include Ollama, BaseRT, vllm-mlx, and oMLX in local runtime discovery/selection; keep raw MLX-LM as a benchmark reference rather than an automatic provider.
Add conservative runtime capability metadata for streaming, prefix/prompt caching, cancellation, structured output, tool calls, and model discovery.
Add a conversational local benchmark covering cold TTFT, warm-prefix TTFT, first useful phrase latency, estimated decode rate, context growth, cancellation/stream-close behaviour, and optional process RSS.
Make local-only construction strict so runtime failures cannot silently become deterministic fallback output.
Bind an unspecified requested model to a model actually reported by the selected local runtime.
Reconcile README/API/local-runtime documentation and repository-wide STATUS state.
Hardware-specific performance ranking remains empirical: the benchmark harness is part of this goal; results for Tomas's M1 Max 64 GB must be recorded only when the benchmark is actually run on that machine, not inferred from CI or external benchmarks.
Goal for the August 2026 local-runtime pass:
portable,mac-performance, andautolocal runtime profiles without changing generalClient(provider="auto")behaviour.main.tom-work-admin.Hardware-specific performance ranking remains empirical: the benchmark harness is part of this goal; results for Tomas's M1 Max 64 GB must be recorded only when the benchmark is actually run on that machine, not inferred from CI or external benchmarks.