Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
52 commits
Select commit Hold shift + click to select a range
87c1906
Document local runtime profiles
krahd Aug 8, 2026
d076408
Add local runtime profile helpers
krahd Aug 8, 2026
08c96ee
Add strict local runtime selection
krahd Aug 8, 2026
be3ef53
Use local runtime selector API in docs
krahd Aug 8, 2026
f8ffa5a
Export local runtime selection helpers
krahd Aug 8, 2026
a1efe25
Test local runtime profiles
krahd Aug 8, 2026
43016e4
Cite current local runtime capabilities
krahd Aug 8, 2026
50aead2
Refresh project status for local runtime profiles
krahd Aug 8, 2026
354f28c
Add BaseRT local provider
krahd Aug 9, 2026
df1cbb0
Register BaseRT provider
krahd Aug 9, 2026
58f5971
Add BaseRT readiness probe
krahd Aug 9, 2026
709832c
Add BaseRT to Mac performance profile
krahd Aug 9, 2026
411b8f3
Test BaseRT Mac runtime selection
krahd Aug 9, 2026
cbb5b4c
Add BaseRT provider tests
krahd Aug 9, 2026
534f06f
Document current Mac runtime options
krahd Aug 9, 2026
35dfae2
Require a loaded model for local selection
krahd Aug 9, 2026
3a38863
Test loaded-model local selection
krahd Aug 9, 2026
22bca55
Make local runtime candidate typing explicit
krahd Aug 9, 2026
4ed14f3
Export BaseRT provider
krahd Aug 9, 2026
933f86b
feat: add vllm-mlx provider
krahd Aug 9, 2026
5a80f6c
feat: register vllm-mlx provider
krahd Aug 9, 2026
1f04aaa
feat: probe vllm-mlx readiness
krahd Aug 9, 2026
e7759f1
feat: add runtime capabilities and vllm-mlx selection
krahd Aug 9, 2026
b1c1a66
feat: expose local runtime capabilities and vllm-mlx
krahd Aug 9, 2026
e8fb383
fix: keep package exports explicit
krahd Aug 9, 2026
5340c9d
feat: align doctor with local runtime profiles
krahd Aug 9, 2026
c317b9c
test: cover vllm-mlx and runtime capabilities
krahd Aug 9, 2026
ce156c0
test: add vllm-mlx provider coverage
krahd Aug 9, 2026
4c08012
test: cover doctor local runtime family
krahd Aug 9, 2026
076955b
feat: add conversational local runtime benchmark
krahd Aug 9, 2026
91b7dbb
test: cover local benchmark helpers
krahd Aug 9, 2026
5ca5f04
feat: expose local runtime benchmark CLI
krahd Aug 9, 2026
ec21acf
docs: document complete local runtime policy
krahd Aug 9, 2026
feffbed
docs: update modelito project status
krahd Aug 9, 2026
4701307
fix: satisfy local runtime lint
krahd Aug 9, 2026
cf46f12
fix: sort local runtime imports
krahd Aug 9, 2026
f1df9b1
fix: sort benchmark imports
krahd Aug 9, 2026
617838f
docs: tighten local runtime claims and examples
krahd Aug 9, 2026
fd6494d
docs: document local runtime profiles and benchmark
krahd Aug 9, 2026
6734797
docs: add local runtime public API
krahd Aug 9, 2026
ffa5e30
ci: enforce black formatting
krahd Aug 9, 2026
4c8576e
ci: show black formatting diffs
krahd Aug 9, 2026
65eabd8
Format provider protocol types
krahd Aug 9, 2026
f22885a
Make Ollama strict mode honour local-only contract
krahd Aug 9, 2026
416768d
Route Ollama through strict-aware provider
krahd Aug 9, 2026
0d6b030
Test strict Ollama local failure semantics
krahd Aug 9, 2026
1a62e99
Format provider module
krahd Aug 9, 2026
209bbe0
Stabilize Ruff lint contract
krahd Aug 9, 2026
62e8630
Expose strict-aware Ollama provider
krahd Aug 9, 2026
8455e16
Test strict Ollama package export
krahd Aug 9, 2026
2e007e9
Restore complete project status and diagrams
krahd Aug 9, 2026
3470c69
Keep local-runtime PR on existing lint contract
krahd Aug 9, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
83 changes: 67 additions & 16 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,11 +3,12 @@ modelito

Modelito is a compact, dependency-light Python library that provides provider-
agnostic abstractions and connectors for large language models (LLMs). It
offers lightweight shims for OpenAI, Claude, Gemini, oMLX, local Ollama deployments,
and local OpenAI-compatible servers (llama.cpp, vLLM, LM Studio), plus
utilities for token counting, timeout estimation, and small helpers to manage
Ollama servers when needed. The library is designed for easy integration into
applications and CI pipelines.
offers lightweight shims for OpenAI, Claude, Gemini, local Ollama deployments,
BaseRT, vllm-mlx, oMLX, and generic OpenAI-compatible servers (llama.cpp,
vLLM, LM Studio, and similar runtimes), plus utilities for token counting,
timeout estimation, provider readiness, and small helpers to manage Ollama
servers when needed. The library is designed for easy integration into
applications and CI pipelines.

Quick start
-----------
Expand Down Expand Up @@ -78,6 +79,7 @@ pip install dist/*.whl
See the `docs/` folder for more details:
- [ARCHITECTURE.md](docs/ARCHITECTURE.md) — Core design, Provider Protocol, and SDK hierarchy
- [USAGE.md](docs/USAGE.md) — Usage guide and examples
- [LOCAL-RUNTIMES.md](docs/LOCAL-RUNTIMES.md) — Local runtime profiles, capabilities, and benchmarking
- [local-openai-compatible.md](docs/local-openai-compatible.md) — Using local OpenAI-compatible servers
- [INSTALL.md](docs/INSTALL.md), [API.md](docs/API.md) — Installation and API reference
- [RELEASE.md](docs/RELEASE.md) — Release checklist and publication steps
Expand Down Expand Up @@ -121,6 +123,10 @@ Provided shims and utilities:
- `ClaudeProvider` — will use the official Anthropic SDK when installed,
falling back to deterministic behavior otherwise.
- `GeminiProvider`, `GrokProvider` — lightweight shims.
- `BaseRTProvider` — thin BaseRT preset built on
`OpenAICompatibleHTTPProvider`.
- `VLLMMLXProvider` — thin vllm-mlx preset built on
`OpenAICompatibleHTTPProvider`.
- `OMLXProvider` — thin oMLX preset built on `OpenAICompatibleHTTPProvider`.
- `OllamaProvider` — HTTP-aware provider that can call a local Ollama HTTP API
through stdlib helpers and can fall back to the local Ollama CLI or
Expand All @@ -132,11 +138,10 @@ The client layer recognises the same provider stack through `ChatProvider`,
`MessageInput`, and structured response helpers such as `Client.chat()` and
`Client.chat_json()`.

`OpenAICompatibleHTTPProvider`, `OMLXProvider`, `OpenAIProvider`, and
`OllamaProvider` also
expose `raw_complete()` and `raw_stream()` for OpenAI-compatible passthrough.
`Client.chat_parsed()` remains the structured JSON convenience path for Python
applications.
`OpenAICompatibleHTTPProvider`, `BaseRTProvider`, `VLLMMLXProvider`,
`OMLXProvider`, `OpenAIProvider`, and `OllamaProvider` expose `raw_complete()`
and `raw_stream()` for OpenAI-compatible passthrough. `Client.chat_parsed()`
remains the structured JSON convenience path for Python applications.

For quick diagnostics, use the provider readiness API or CLI:

Expand All @@ -154,6 +159,51 @@ dict conversion.
python -m modelito doctor --provider omlx --model omlx
```

Local runtime profiles
----------------------

For applications that require local execution rather than Modelito's general
provider auto-selection, use `local_client()` or `select_local_runtime()`:

```py
from modelito import local_client

client = local_client(
profile="mac-performance",
models={
"basert": "my-base-model",
"vllm-mlx": "my-vllm-model",
"omlx": "my-omlx-model",
"ollama": "my-ollama-tag",
},
)
```

The profiles are:

- `portable`: Ollama as the common cross-platform path;
- `mac-performance`: on Apple Silicon, try BaseRT, vllm-mlx, oMLX, then
Ollama;
- `auto`: use `mac-performance` on Apple Silicon and `portable` elsewhere.

That order is a deployment starting point, not a universal speed ranking.
`prefer=` can override it after benchmarking the target workload. Local
selection is strict: it does not silently fall back to a hosted provider or to
the deterministic offline shim.

`local_runtime_capabilities()` exposes conservative metadata for streaming,
prefix caching, cancellation, structured output, tool calls, and model
discovery. Model- or configuration-dependent features are marked
`conditional`; uncertain claims are marked `unknown`.

For latency-sensitive local work, `modelito-benchmark-local` measures first
request TTFT, warm-prefix TTFT, first useful streamed phrase, estimated decode
rate, context-growth latency, cancellation/stream-close behavior, and optional
process RSS against an already-running OpenAI-compatible server. Raw MLX-LM is
supported as a benchmark reference without being added to automatic runtime
selection. See [docs/LOCAL-RUNTIMES.md](docs/LOCAL-RUNTIMES.md) for the caveats
and current upstream-source audit.

Server mode for non-Python clients:

```sh
Expand Down Expand Up @@ -207,9 +257,9 @@ Example `~/.pi/agent/models.json` provider entry:
```

Tool-calling workflows require raw passthrough support. Modelito currently
implements that on `OpenAICompatibleHTTPProvider`, `OMLXProvider`,
`OpenAIProvider`, and `OllamaProvider` via Ollama's
`/v1/chat/completions` endpoint.
implements that on `OpenAICompatibleHTTPProvider`, `BaseRTProvider`,
`VLLMMLXProvider`, `OMLXProvider`, `OpenAIProvider`, and `OllamaProvider` via
their OpenAI-compatible chat-completions paths.

The package also exposes a small Ollama administration layer for local model
operations, including install backend detection, remote catalog metadata,
Expand Down Expand Up @@ -269,8 +319,9 @@ compatible with existing duck-typed providers — it requires only:
- `summarize(messages, settings=None)` -> `str`

All built-in providers shipped with the package (`OpenAIProvider`,
`ClaudeProvider`, `GeminiProvider`, `OMLXProvider`, `OllamaProvider`, `GrokProvider`) satisfy
the `Provider` protocol structurally. The `Provider` Protocol is decorated with
`ClaudeProvider`, `GeminiProvider`, `BaseRTProvider`, `VLLMMLXProvider`,
`OMLXProvider`, `OllamaProvider`, `GrokProvider`) satisfy the `Provider`
protocol structurally. The `Provider` Protocol is decorated with
`@runtime_checkable`, so you can use `isinstance()` checks at runtime when
you need to enforce the contract in application code.

Expand Down Expand Up @@ -399,4 +450,4 @@ These modules are not currently presented as stable top-level package exports
in this README. Prefer the documented `Client`, provider adapters, connector,
and server entrypoints for application integrations.

See the `tests/` directory for comprehensive coverage and usage examples.
See the `tests/` directory for comprehensive coverage and usage examples.
Loading
Loading