Skip to content

Complete local runtime strategy and benchmarking pass #24

Description

@krahd

Goal for the August 2026 local-runtime pass:

  • Add explicit portable, mac-performance, and auto local runtime profiles without changing general Client(provider="auto") behaviour.
  • Include Ollama, BaseRT, vllm-mlx, and oMLX in local runtime discovery/selection; keep raw MLX-LM as a benchmark reference rather than an automatic provider.
  • Add conservative runtime capability metadata for streaming, prefix/prompt caching, cancellation, structured output, tool calls, and model discovery.
  • Add a conversational local benchmark covering cold TTFT, warm-prefix TTFT, first useful phrase latency, estimated decode rate, context growth, cancellation/stream-close behaviour, and optional process RSS.
  • Make local-only construction strict so runtime failures cannot silently become deterministic fallback output.
  • Bind an unspecified requested model to a model actually reported by the selected local runtime.
  • Reconcile README/API/local-runtime documentation and repository-wide STATUS state.
  • Obtain a completely green CI run for PR Add explicit local runtime profiles #23.
  • Resolve all current PR review findings.
  • Merge PR Add explicit local runtime profiles #23 to main.
  • Verify post-merge CI and update tom-work-admin.

Hardware-specific performance ranking remains empirical: the benchmark harness is part of this goal; results for Tomas's M1 Max 64 GB must be recorded only when the benchmark is actually run on that machine, not inferred from CI or external benchmarks.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions