Last updated: 2026-08-25 06:39
Modelito is a compact, dependency-light Python library for provider-agnostic LLM access. It supports hosted and local providers, OpenAI-compatible serving, streaming and structured responses, embeddings, readiness probes, token/timeout helpers, Ollama administration, and deterministic/offline-friendly test fallbacks.
- Package metadata version:
1.4.6. - Python: 3.10–3.12.
- Licence: MIT.
- Hosted providers include OpenAI, Anthropic/Claude, Gemini, and Grok.
- Local providers include Ollama, BaseRT, vllm-mlx, oMLX, and generic OpenAI-compatible HTTP servers.
modelito-serveexposes OpenAI-compatible models, chat-completions, and embeddings endpoints.modelito-doctorandcheck_provider_ready()provide read-only provider diagnostics.modelito-benchmark-localmeasures conversational latency for already-running OpenAI-compatible local runtimes.- Raw OpenAI chat-completions passthrough is supported by raw-capable providers for tool-calling and metadata-preserving integrations.
- Ollama helpers cover installation detection, local service control, model lifecycle/readiness, download, preload, and diagnostics while keeping mutating operations explicit.
RecordingProvider/ReplayProviderprovide JSONL cassette recording and deterministic replay.
The explicit local-only surface is separate from the established
Client(provider="auto") contract:
portable: Ollama as the common macOS/Linux/Windows path;mac-performance: on Apple Silicon, BaseRT → vllm-mlx → oMLX → Ollama;auto:mac-performanceon Apple Silicon andportableelsewhere;MODELITO_LOCAL_PROFILEenvironment configuration;select_local_runtime()for read-only selection and diagnostics;local_client()for a strict local-only client with no hosted or deterministic fallback;- provider-specific model, endpoint, and API-key mappings;
- explicit
prefer=ordering so benchmark results can override defaults; local_runtime_capabilities()for conservative runtime-family capability metadata.
The candidate order is a deployment starting point, not a universal performance claim. Modelito does not encode a claim that any one local runtime is fastest for every model or workload.
The portable path. Current Apple-Silicon releases can use MLX and cache/snapshot optimisations, while later 2026 releases also use llama.cpp paths to broaden model and hardware support.
An Apple-Silicon native-Metal runtime exposed through an OpenAI-compatible server. Modelito integrates only through the local HTTP API.
An Apple-Silicon MLX server with OpenAI-compatible serving, caching/batching, structured-output and cancellation capabilities in current upstream releases.
An MLX-native server oriented towards persistent conversational/agent workloads, including continuous batching and tiered prefix/KV caching.
The existing OpenAICompatibleHTTPProvider remains the escape hatch for
llama.cpp, LM Studio, vLLM, MLX-LM's HTTP server, and other compatible
endpoints. Arbitrary endpoints are not auto-selected.
Raw MLX-LM remains a benchmark/reference path rather than another automatic
runtime provider. Its prompt caching is useful for repeated conversational
contexts. The benchmark CLI recognises mlx-lm as a label for comparisons.
modelito-benchmark-local records:
- first-request TTFT;
- first phrase-like streamed latency;
- estimated decode tokens/s;
- warm-prefix TTFT over repeated requests;
- context-growth TTFT;
- client stream-close latency and a post-cancellation probe;
- optional sampled server process RSS.
The benchmark embeds its own caveats. First-request TTFT is only a cold-model measurement when the server/model was actually cold. Token rate uses Modelito's token-count helper, client stream-close time does not prove server-side cancellation acknowledgement, and process RSS can under-report Metal/unified memory on macOS.
This benchmark is intended to compare equivalent workloads on the target
machine. It does not replace runtime-specific instrumentation such as
vllm-mlx's bench-serve command.
The primary application surface is modelito.Client, backed by registered
provider adapters. Providers implement a small common interface and may expose
richer raw/streaming capabilities when available. The local-runtime selector is
policy layered above those providers rather than another provider protocol.
Development installation:
python -m pip install -e '.[dev]'Typical checks:
python scripts/check_no_legacy_dicts.py
ruff check .
black --check .
mypy modelito --ignore-missing-imports
pytest -q --ignore=tests/integration tests
python -m buildLocal runtime policy, benchmark usage, capability caveats, and upstream sources
are documented in docs/LOCAL-RUNTIMES.md.
- Full non-integration suite: 397 tests passed and 1 skipped on 2026-08-25.
- Focused Ollama/client regression suite: 68 tests passed on 2026-08-25.
- The available pytest configuration emits one warning because the installed
pytest does not recognise
asyncio_default_fixture_loop_scope. - Ruff and mypy were not available in the current development environment, so lint and type checks could not be run.
python3 -m compileall -q modelitopassed. Black could not start because its installed version imports a removed private symbol from the installed Click.- The package build was not run for the current Ollama settings change.
- Ollama native
/api/chatand/api/generaterequests now preserve message roles, explicitly select synchronous or streaming responses, place known generation settings in the nativeoptionsobject, and use separate documentation-based allowlists for each endpoint's top-level controls. Undocumented and unknown flat settings are not guessed;settings["options"]remains the explicit escape hatch. - Ollama strict enforcement now lives in the base provider as well as the
package/registry compatibility export.
strict=Truesummary and streaming requests use direct HTTP and never fall back to CLI or deterministic output;strict=Falseretains the resilient fallback chain. - Added explicit
portable,mac-performance, andautolocal-runtime profiles without altering the establishedClient(provider="auto")contract. - Added BaseRT and vllm-mlx OpenAI-compatible provider presets alongside oMLX and Ollama.
- Added provider-specific model/endpoint/key mappings, readiness-based model
resolution, benchmark-overridable
prefer=ordering, and conservative capability metadata. - Added
modelito-benchmark-localfor first-turn, warm-prefix, context-growth, decode, cancellation-close, and approximate RSS measurements. - Added a strict-aware Ollama surface so local-only clients propagate runtime failures instead of returning deterministic fallback text; the package-root export now uses the same class as the provider registry.
- Explicitly configured the historical Ruff lint contract after Ruff 0.16 expanded its unconfigured default rule set. This prevents an unpinned tool update from silently redefining repository lint policy.
- Restored repository-wide pending work and inline current-state diagrams in this status snapshot.
- Do not encode a universal local-runtime performance ranking.
- Keep Ollama as the portable path.
- Keep a Mac-oriented path because Apple-Silicon-specific runtimes expose materially different caching, serving, and execution strategies.
- Preserve existing
Client(provider="auto")behaviour. - A local-only client must fail explicitly when no requested local backend or model is ready.
- Provider-specific model identifiers and endpoints must be expressible independently.
- Local runtime defaults remain benchmark-overridable with
prefer=. - Capability metadata is conservative:
conditionalandunknownare used instead of guessing model- or version-specific support. - Speech/VAD/TTS/ASR orchestration does not belong in Modelito's LLM runtime abstraction.
- No release, tag, or version bump is part of this work.
These repository-wide tasks remain open and are not superseded by the local runtime work:
ClaudeProviderstill has noraw_complete()/raw_stream()surface, so it cannot serve as a Pi tool-calling backend through the Modelito HTTP server. Anthropic tool-call response translation requires a dedicated follow-up.GeminiProviderandGrokProviderstill have nochat()implementation; they remain lower-priority compatibility shims.
The repository now contains the benchmark needed for workload-specific local selection, but a real Apple-Silicon performance ranking must be measured on the actual target machine with comparable model families, quantisations, runtime configuration, and cache state. GitHub-hosted Linux CI cannot provide that evidence.
Until those measurements exist, mac-performance is intentionally a curated
candidate order with an explicit prefer= override rather than an empirical
winner table.
- Keep the current CI green and resolve all PR review findings before merging the local-runtime work.
- Run the conversational benchmark on the target Apple-Silicon machine with equivalent models/configurations and record results separately from runtime marketing benchmarks.
- Keep reviewing provider additions against the portable-common-surface rule.
- Continue monitoring Ollama raw passthrough behaviour and keep docs/tests aligned with OpenAI-compatible payload expectations.
- Address Claude raw passthrough and Gemini/Grok chat surfaces in dedicated, bounded follow-ups rather than expanding the local-runtime PR.
- Maintain a small stable provider protocol surface.
- Keep hosted SDK dependencies optional.
- Expand provider-specific helpers only when they are clearly useful and well-contained.
- Local backend performance and model support change quickly upstream.
- Model identifiers and formats differ between runtimes.
- BaseRT, vllm-mlx, and oMLX are Apple-Silicon-oriented; hosted Linux CI can validate adapters and policy but not their native execution.
- Readiness probes establish availability, not latency, quality, memory pressure, cache effectiveness, or thermal behaviour.
- Capability metadata can become stale and should be reviewed when upstream runtime behaviour materially changes.
- Deterministic fallbacks remain useful for tests but are inappropriate for the strict local-only runtime path.
- API key storage should not move into a built-in encrypted database in the core package.
- Cloud-provider integrations should remain lightweight shims by default.
- The core value of the package is provider-agnostic normalisation, optional local tooling, and dependency-light embeddability.
- CI intentionally excludes integration tests by path/flags to keep default hosted CI fast and safe.
Last updated: 2026-08-25 06:39