From ea57680d0d849c3291edbaf3da3a38cb20863411 Mon Sep 17 00:00:00 2001 From: James Dumay Date: Sat, 12 Sep 2026 18:48:12 +1000 Subject: [PATCH 1/4] docs(skippy): add KV-cache disk tier operator guide Document disk-cache modes, configuration precedence, budgets, observability, maintenance, failure handling, and recovery. --- docs/README.md | 1 + docs/skippy/KV_CACHE_DISK.md | 250 +++++++++++++++++++++++++++++++++++ 2 files changed, 251 insertions(+) create mode 100644 docs/skippy/KV_CACHE_DISK.md diff --git a/docs/README.md b/docs/README.md index 238750fe06..0f02077dd6 100644 --- a/docs/README.md +++ b/docs/README.md @@ -30,6 +30,7 @@ Use this hub to find project guides that are not owned by a single Rust crate. | [skippy/TOPOLOGY_PLANNER.md](skippy/TOPOLOGY_PLANNER.md) | Stage topology planning behavior | | [skippy/CONFIGURATION.md](skippy/CONFIGURATION.md) | Authoritative operator matrix for Skippy config keys and rejection boundaries | | [skippy/PROMPT_CACHE.md](skippy/PROMPT_CACHE.md) | OpenAI prompt-prefix cache behavior, defaults, telemetry, and benchmark flow | +| [skippy/KV_CACHE_DISK.md](skippy/KV_CACHE_DISK.md) | Operator guide for the node-local disk prompt cache (L3): config, modes, status, prune/clear, corruption handling | | [skippy/PIPELINED_VERIFY_WINDOW.md](skippy/PIPELINED_VERIFY_WINDOW.md) | Native MTP, anchored N-gram extension, VerifyWindow protocol, pipeline behavior, and telemetry | | [skippy/SUFFIX_NGRAM_PROPOSER.md](skippy/SUFFIX_NGRAM_PROPOSER.md) | Long exact-suffix proposer design, invariants, telemetry, and benchmark contract | | [skippy/DATA_FLOW.md](skippy/DATA_FLOW.md) | Stage data flow and transport details | diff --git a/docs/skippy/KV_CACHE_DISK.md b/docs/skippy/KV_CACHE_DISK.md new file mode 100644 index 0000000000..e22d2c24db --- /dev/null +++ b/docs/skippy/KV_CACHE_DISK.md @@ -0,0 +1,250 @@ +# KV-Cache Disk Tier — Operator Guide + +The node-local **disk prompt cache** is the durable tier under Skippy's radix +cache: exported continuation (KV) state is cut into content-addressed segments +and committed under manifests so a later run can restore an exact prefix from +disk instead of paying cold prefill. This guide covers operating the tier that +ships today (the exact L1/L3 stack): configuration, defaults, status, restart +behavior, prune/clear, corruption handling, and troubleshooting. + +This is a different cache from the in-memory OpenAI prompt-prefix cache +documented in [PROMPT_CACHE.md](PROMPT_CACHE.md). That one lives in the serving +process; this one is the on-disk L3 store described here. + +**Fail-closed configuration, fail-open runtime.** Invalid disk-cache +*configuration* fails **closed**: config validation and resolution reject bad +values — fixed mode without a budget, a non-IEC or zero size, a relative +directory, or a minimum-free reserve below 1 GiB — by returning an error, so a +node refuses to start on a broken cache setting rather than silently ignoring +it. A *valid* configuration is then **fail-open for inference**: if the store +cannot open, reaches low space, or cannot admit a write, the node logs a warning +and serves with cold prefill. Valid runtime unavailability never blocks +generation. + +## Modes and safe defaults + +The tier has three modes, selected by `mode` / `--kv-cache-disk`: + +| Mode | Meaning | +|---|---| +| `off` | Disabled. **This is the default** — no disk cache unless you opt in. | +| `auto` | Enabled with an automatically computed budget from live filesystem free space (see [Auto budget](#auto-budget)). | +| `fixed` | Enabled with an explicit hard byte cap you set. A budget is **required** in this mode. | + +Because the default is `off`, an out-of-the-box node writes nothing to disk for +the prompt cache. Turn it on deliberately. + +## Configuration surface + +The same four settings are expressible through a config file, environment +variables, and CLI flags. Sizes everywhere use **explicit IEC suffixes** +(`KiB`, `MiB`, `GiB`, `TiB`) on a positive whole number — for example `32GiB`. +Bare numbers and decimal/`GB`-style units are rejected. + +| Setting | Config (`[runtime.kv_cache.disk]`) | Environment | CLI flag | +|---|---|---|---| +| Mode / fixed budget | `mode` (`off`/`auto`/`fixed`) + `budget_mib` | `MESH_LLM_KV_CACHE_DISK` (`off`/`auto`/`SIZE`) | `--kv-cache-disk off\|auto\|SIZE` | +| Directory | `directory` (absolute) | `MESH_LLM_KV_CACHE_DISK_DIR` (absolute) | `--kv-cache-disk-dir ABSOLUTE_PATH` | +| Minimum free reserve | `minimum_free_mib` | `MESH_LLM_KV_CACHE_MIN_FREE` (SIZE) | `--kv-cache-min-free SIZE` | + +Notes: + +- In the config file, `budget_mib` is a plain MiB integer and is **only valid + when `mode = "fixed"`**; it is **required** there. Setting it under `off`/`auto` + is a config error. +- On the CLI and in the environment, mode and budget share one value: + `--kv-cache-disk 32GiB` selects fixed mode at 32 GiB; `--kv-cache-disk auto` + and `--kv-cache-disk off` select those modes. +- Directories must be absolute. Relative paths (including bare-drive forms like + `C:\cache` on non-Windows hosts) fail closed rather than resolving under the + working directory. + +### Precedence + +Each of the four settings is resolved **independently, field by field**, with +later sources overriding earlier ones: + +```text +CLI flag > MESH_LLM_KV_CACHE_* env > config file > built-in default +``` + +Legacy `SKIPPY_L3_*` variables only fill a field for which none of those public +sources supplied a value; they never override a public config field. + +So you can pin the directory in the config file and still override just the +budget with `--kv-cache-disk`, without disturbing the other fields. `mesh-llm +kv-cache status` reports the winning source per field (`default`, `config`, +`environment`, `cli`, or `legacy_environment`). + +### Legacy environment variables (deprecated) + +`SKIPPY_L3_DIR` and `SKIPPY_L3_BUDGET_BYTES` are still honored but **only as a +field-level fallback** when no public setting is present, and they emit a +deprecation warning. Migrate to `[runtime.kv_cache.disk]` or the +`MESH_LLM_KV_CACHE_*` variables. Two behaviors to know: + +- Presence of `SKIPPY_L3_DIR` (without a public mode) implies `fixed` mode, and + the budget defaults to the legacy **32 GiB** when unset. +- `SKIPPY_L3_BUDGET_BYTES=0` **no longer means unbounded** — it is treated as + the 32 GiB legacy default, with a warning. `SKIPPY_L3_BUDGET_BYTES` is ignored + entirely if `SKIPPY_L3_DIR` is not set. + +## Directory, budget, and minimum-free behavior + +**Directory.** When unset, the root resolves to `$MESH_LLM_HOME/kv-cache`, or +`~/.mesh-llm/kv-cache` when `MESH_LLM_HOME` is unset. On startup the node +creates the root and its store subdirectories (`segments/`, `manifests/`, +`prefixes/`, and the packed-store dirs; `quarantine/` is used when an object +fails verification), restricts them to owner-only permissions (`0700`), and +takes an **exclusive lock** on `.owner.lock` in the root. One process owns a cache root at a time; a second node pointed at the same +root fails to acquire the lock rather than corrupting it. The root must not +contain symlinks — a symlinked entry under the root is refused. + +**Minimum free reserve.** `minimum_free` is the free space the store preserves +for everything else on the filesystem. Default **16 GiB** +(`DEFAULT_KV_DISK_MINIMUM_FREE_MIB`); the floor is **1 GiB** +(`MIN_KV_DISK_MINIMUM_FREE_MIB`) and a smaller value is rejected. When the +filesystem sits at the reserve, the store goes **read-only**: existing entries +still serve restores, but new writes are refused (`read_only_low_space`) so the +cache never eats into the reserve. + +**Fixed budget.** The hard whole-node cap. Eviction runs oldest-manifest-first +and evicts to ~85% of the budget (a low-water margin that amortizes the +O(manifests×segments) eviction scan); the newest manifest is never evicted. + +**Auto budget.** In `auto` mode the budget is resolved +from live filesystem facts immediately before the root opens, as the **minimum** +of: + +- 20% of the capacity basis (current filesystem available + bytes already + managed under this root), +- what is actually allocatable after honoring `minimum_free`, and +- a hard **64 GiB** ceiling. + +If that resolves to `0` (for example, the filesystem is already at or below the +minimum-free reserve), the disk cache stays disabled and the node logs a warning +— again, cold prefill still serves. + +## Status and observability + +Inspect a node with: + +```bash +mesh-llm kv-cache status # human-readable +mesh-llm kv-cache status --json # machine-readable +``` + +Without `--endpoint`, the command talks to the local node's loopback control +API (default port `3131`; the control endpoint is loopback-only). To inspect +nodes you own remotely, pass one or more `--endpoint ` values (repeatable). + +The status payload reports: + +- `configured` — the resolved `mode`, `directory`, `budget_bytes`, + `minimum_free_bytes`, and the winning **source** for each field. +- `effective.state` — one of: + - `off` — mode is off. + - `active` — store is open and admitting writes. + - `read_only_low_space` — at the minimum-free reserve; reads serve, writes + refused. + - `degraded` — configured on but no manager: `reason` is + `storage_unavailable` (couldn't open the root), + `budget_below_entry_floor` (auto budget resolved to zero), or a runtime + storage error. +- `usage` — `budget_bytes`, `used_bytes`, `reserved_inflight_bytes`, + `filesystem_available_bytes`, `minimum_free_bytes`, `manifests`, + `unique_segments`, `evicted_manifests`, and `quarantined_objects`. +- `activity`, `reconciliation` (see below), and `inventory` (per-model entries). + +At **startup**, any resolution warnings (deprecated legacy vars, zero auto +budget, unavailable store) are emitted as `Warning` events in the node log. +Check the log first when the cache "isn't caching." + +## Restart vs. live-apply semantics + +On a config reload, disk-cache changes split into two classes: + +| Change | Applied | +|---|---| +| `budget_mib`, `minimum_free_mib` | **Live.** Limits update in place; shrinking evicts inactive entries immediately, pinned entries stay valid, and writes stay refused until usage fits. | +| `mode`, `directory` | **Restart required.** These are preserved across a live reload and only take effect when the node restarts. | + +If a reload changes only `mode`/`directory`, those fields are held at their +previous values (and the mode-coupled budget with them) until restart. Plan a +node restart when you move the cache directory or turn the tier on/off via +config reload. + +## Prune, clear, and shutdown + +Both operations act on **inactive** entries only — pinned/in-use state is never +removed — and both are gated behind an explicit confirmation (or `--yes`): + +```bash +# Evict least-recently-used inactive entries, optionally down to a target size +# and/or scoped to one exact model identity. +mesh-llm kv-cache prune [--target 16GiB] [--model-identity ] [--yes] + +# Remove inactive entries entirely (inference falls back to cold prefill), +# optionally scoped to one model identity; omit the filter to clear the root. +mesh-llm kv-cache clear [--model-identity ] [--yes] +``` + +- `--model-identity` takes an **exact numerical model identity**; display names + are not accepted, and a blank/whitespace filter is rejected (it would match + nothing and report a no-op success). +- `prune` without `--target` trims toward the low-water margin of the budget. + On an uncapped (legacy budget `0`) store this default would remove everything, + so that case is refused and steered to `clear`. +- Both accept `--endpoint`/`--port` for owner-controlled remote nodes and + `--json`. + +**Shutdown.** There is no flush step to run. Manifests only commit after every +referenced segment is present and the assembled payload digest matches, so a +process that stops mid-write leaves temp files, never a partial cache entry. +The root lock is released on exit; leftover temp files are reconciled on the +next start (below). A hard kill is safe — you lose only in-flight writes. + +## Corruption, quarantine, and cold fallback + +Integrity is content-addressed end to end. Every segment is addressed by the +BLAKE3 digest of its bytes and reads verify that digest, so corruption is +**detected, never silently imported**. A manifest is only loadable once all its +segments are present and the reassembled payload digest matches; partial state +is unreadable by construction. + +On startup the node **reconciles** the root before serving and records what it +repaired in `reconciliation`: + +- `removed_temporary_files` — abandoned in-flight temp files. +- `quarantined_manifests` — manifests that failed verification, moved aside into + `quarantine/`. +- `removed_prefix_links` — dangling prefix index links. +- `removed_orphan_bytes` — unreferenced segment bytes garbage-collected. + +Objects that fail verification while serving are moved to `quarantine/` and +counted in `usage.quarantined_objects`. A quarantined or missing entry is simply +a cache miss: the request falls back to cold prefill. Rising +`quarantined_objects` points at underlying storage trouble (bad disk, truncation, +external tampering) — investigate the filesystem; the cache itself stays safe. + +## Troubleshooting + +| Symptom | Likely cause | What to do | +|---|---|---| +| No disk caching at all | Mode is `off` (the default) | Set `--kv-cache-disk auto` (or `fixed SIZE`), or `[runtime.kv_cache.disk] mode`. | +| Node logs "disk prompt cache is unavailable; inference will use cold prefill" | Root couldn't be opened (permissions, missing parent, symlink in root, lock held by another process) | Check the directory exists and is writable, contains no symlinks, and no other node owns `.owner.lock`. `status` shows `degraded`/`storage_unavailable`. | +| `auto` mode caches nothing | Auto budget resolved to `0` — filesystem at/below `minimum_free` | Free space or lower `--kv-cache-min-free` (floor 1 GiB). `status` shows `degraded`/`budget_below_entry_floor`. | +| Writes refused, reads still work | At the minimum-free reserve | `status` state `read_only_low_space`; free disk, `prune`, or lower `minimum_free`. | +| Startup fails to open the cache | Another process holds the root lock | Only one node per cache root; point the second node at a different `directory`. | +| Node refuses to start with a config/resolution error (e.g. "must use an explicit IEC suffix", "must be an absolute path", "fixed disk prompt-cache mode requires a positive budget", minimum-free below 1 GiB) | **Invalid configuration** — rejected up front, not a runtime fallback | Fix the value: use `32GiB` (not `32`/`32GB`), absolute directories, a positive `budget_mib` under `fixed`, and `minimum_free ≥ 1GiB`. These are config errors, distinct from the store being unavailable at runtime. | +| Deprecation warning about `SKIPPY_L3_*` | Legacy env vars in use | Migrate to `[runtime.kv_cache.disk]` or `MESH_LLM_KV_CACHE_*`. | +| Config reload didn't move the cache / toggle the mode | `mode`/`directory` are restart-only | Restart the node; only `budget`/`minimum_free` apply live. | +| `quarantined_objects` climbing | Segments failing digest verification | Inspect the underlying storage; entries fall back to cold prefill, cache stays safe. | +| `prune` reports it freed nothing | Wrong/blank `--model-identity`, or only pinned entries present | Use the exact numerical model identity; pinned/in-use entries are never pruned. | + +## Reference + +- Config resolution and precedence: `crates/mesh-llm-host-runtime/src/runtime/kv_disk_config.rs` +- Config schema and constants: `crates/mesh-llm-config/src/model.rs`, `crates/mesh-llm-config/src/validate.rs` +- L3 store, integrity, reconciliation, status contract: `crates/skippy-cache/src/l3.rs` +- CLI and control API: `crates/mesh-llm-cli/src/parser/commands.rs`, `crates/mesh-llm-commands/src/kv_cache.rs`, `crates/mesh-llm-host-runtime/src/api/routes/kv_cache.rs` From 3c18a33ea12e41265530399f60556e1c16c2bcc3 Mon Sep 17 00:00:00 2001 From: scama Date: Sat, 12 Sep 2026 18:48:13 +1000 Subject: [PATCH 2/4] docs(skippy): add KV caching website guide Explain the default cache experience, supported controls, durable disk opt-in, backend scope, and experimental CacheGen boundary. --- website/src/_data/docs.js | 1 + website/src/docs/pages/config-defaults.md | 5 +- website/src/docs/pages/config-reference.md | 11 +- website/src/docs/pages/kv-caching.md | 239 ++++++++++++++++++ .../src/docs/pages/openai-compatible-api.md | 5 + 5 files changed, 255 insertions(+), 6 deletions(-) create mode 100644 website/src/docs/pages/kv-caching.md diff --git a/website/src/_data/docs.js b/website/src/_data/docs.js index b8b435e00f..2c37496446 100644 --- a/website/src/_data/docs.js +++ b/website/src/_data/docs.js @@ -40,6 +40,7 @@ export default [ links: [ ["OpenAI-compatible API", "/docs/pages/openai-compatible-api/"], ["Automatic routing", "/docs/pages/automatic-routing/"], + ["KV caching", "/docs/pages/kv-caching/"], ["Streaming", "/docs/pages/openai-compatible-api/#streaming"], ["Tool calling", "/docs/pages/openai-compatible-api/#tool-calling"], ["Structured outputs", "/docs/pages/openai-compatible-api/#structured-outputs"] diff --git a/website/src/docs/pages/config-defaults.md b/website/src/docs/pages/config-defaults.md index d69d7ec76f..97f6b63306 100644 --- a/website/src/docs/pages/config-defaults.md +++ b/website/src/docs/pages/config-defaults.md @@ -14,8 +14,9 @@ Shared default settings applied to every model. Individual model entries can ove ctx_size = 0 # Context window (0 = auto) batch = 0 # Batch size (0 = auto) ubatch = 0 # Micro-batch size (0 = auto) -cache_type_k = "f16" # Key cache dtype -cache_type_v = "f16" # Value cache dtype +cache_type_k = "auto" # Key cache dtype (resolved by policy) +cache_type_v = "auto" # Value cache dtype (resolved by policy) +kv_cache_policy = "balanced" # "balanced", "quality", or "saver" kv_offload = "auto" # KV-cache offload policy prompt_cache = "auto" # Prompt-cache policy flash_attention = "auto" # Flash-attention policy diff --git a/website/src/docs/pages/config-reference.md b/website/src/docs/pages/config-reference.md index cca86e0c7f..965ba8f292 100644 --- a/website/src/docs/pages/config-reference.md +++ b/website/src/docs/pages/config-reference.md @@ -137,14 +137,14 @@ for the activity policy and privacy boundary. | `model_fit.ctx_size` | integer | `0` = auto | both | model reload | wired | `--ctx-size` on the ad-hoc single-model path | | `model_fit.batch` | integer | `0` = auto (`n_batch`) | both | model reload | wired | none | | `model_fit.ubatch` | integer | `0` = auto (`n_ubatch`); should not exceed `batch` | both | model reload | wired | none | -| `model_fit.cache_type_k`
`model_fit.cache_type_v` | enum (dtype) | `auto`, `f32`, `f16` (default), `bf16`, `q8_0`, `q4_0`, `q4_1`, `iq4_nl`, `q5_0`, `q5_1`; explicit value overrides `kv_cache_policy` | both | model reload | wired | none | -| `model_fit.kv_cache_policy` | enum | `auto`, `quality`, `balanced`, `saver`; expands into cache dtypes | both | model reload | wired | none | +| `model_fit.cache_type_k`
`model_fit.cache_type_v` | enum (dtype) | `auto` (default), `f16`, `q8_0`, `q4_0`; schema validation also accepts `f32`, `bf16`, `q4_1`, `iq4_nl`, `q5_0`, `q5_1`, but the normal serving runtime rejects them at model load; explicit value overrides `kv_cache_policy` | both | model reload | partial | none | +| `model_fit.kv_cache_policy` | enum | `balanced` (default), `auto`, `quality`, `saver`; expands into cache dtypes | both | model reload | wired | none | | `model_fit.kv_offload` | bool-or-`auto` | `auto` | both | model reload | wired | none | | `model_fit.kv_unified` | bool-or-`auto` | `auto` | both | model reload | wired (recurrent/hybrid architectures still force this true natively) | none | | `model_fit.cache_ram_mib` | integer | unset (no cap) | both | model reload | unwired (any positive value fails at model load) | none | -| `model_fit.cache_idle_slots` | integer | unset (unbounded) | both | model reload | wired | none | +| `model_fit.cache_idle_slots` | integer | unset uses the runtime lane count; `0` drops every reset lane, positive values cap retained idle sessions | both | model reload | wired | none | | `model_fit.prompt_cache` | bool-or-`auto` | `auto` | both | model reload | wired | none | -| `model_fit.prefix_cache.enabled` | boolean | unset (disabled) | both | model reload | wired | none | +| `model_fit.prefix_cache.enabled` | boolean | unset uses family defaults; `false` disables | both | model reload | wired | none | | `model_fit.prefix_cache.max_entries` | integer | runtime default | both | model reload | wired | none | | `model_fit.prefix_cache.max_bytes` | integer | `0`/unset = no cap | both | model reload | wired | none | | `model_fit.prefix_cache.min_tokens` | integer | runtime default | both | model reload | wired | none | @@ -162,6 +162,9 @@ Missing TOML for this group: GGUF metadata `kv_overrides`. There is no schema key for it yet; do not expect an override path until a later PR adds one. +See [KV Caching](/docs/pages/kv-caching/) for the default behavior and common +configuration recipes. + ## Group 4: device selection, GPU offload, multi-GPU, CPU MoE, and loading behavior | Key path | Type | Allowed values / default (`auto`) | `[defaults]` / `[[models]]` | Restart | Status | CLI equivalent | diff --git a/website/src/docs/pages/kv-caching.md b/website/src/docs/pages/kv-caching.md new file mode 100644 index 0000000000..2e088acf28 --- /dev/null +++ b/website/src/docs/pages/kv-caching.md @@ -0,0 +1,239 @@ +--- +title: KV Caching +description: Reuse prompt work in memory or on disk and tune KV memory use +--- + +# KV Caching + +KV caching lets Mesh reuse the model state created while reading a prompt. +When a later request starts with the same tokens, Mesh can restore that state +and skip some or all of the repeated prefill work. This lowers time to first +token for repeated system prompts, long documents, and multi-turn chats. + +Applications do not need a cache-specific API. Keep using the +[OpenAI-compatible API](/docs/pages/openai-compatible-api/) and send the full +conversation or prompt on every request. Mesh identifies matching prefixes and +reports reused prompt tokens in `usage.prompt_tokens_details.cached_tokens`. +Callers that already use OpenAI's `prompt_cache_key` may keep sending it; Mesh +trims the value and uses it as a cache and routing namespace. Keep it stable for +the requests that should share cached prefixes. It does not force a cache hit, +and requests without it share the default namespace. Do not put secrets in the +key because Mesh records it verbatim in request telemetry. + +For the Responses API, an explicit `prompt_cache_key` wins. When it is absent, +Mesh uses `previous_response_id`, then the conversation ID, as the cache key. + +## Default experience + +With no cache settings: + +- Mesh enables the in-memory prefix cache with family-aware limits. Prefixes of + at least 256 tokens are eligible, and Mesh selects the stored state format + from the model architecture. +- The balanced KV memory policy uses Q8_0 key and value caches for models under + 50 GiB and Q4_0 for models at or above 50 GiB. If metadata proves that the + selected quantized format cannot load, the automatic policy falls back to + F16. An explicit incompatible override fails instead of changing silently. +- KV offload and unified-cache behavior remain automatic. +- The durable disk cache is off, so a process restart starts with an empty + prompt cache and Mesh writes no prompt state to disk. + +These defaults apply across the supported CPU, Metal, CUDA, and ROCm runtimes. + +## The controls at a glance + +| Goal | Setting | Default | Where to set it | +|---|---|---|---| +| Choose KV memory/quality policy | `model_fit.kv_cache_policy` | `balanced` | `[defaults.model_fit]` or `[models.model_fit]` | +| Pin key/value cache formats | `model_fit.cache_type_k`, `model_fit.cache_type_v` | `auto` | config file | +| Control KV device offload | `model_fit.kv_offload` | `auto` | config file | +| Cap retained idle native sessions | `model_fit.cache_idle_slots` | lane count | config file | +| Disable all prompt-prefix reuse | `model_fit.prompt_cache` | `auto` | config file | +| Tune or disable in-memory prefix reuse | `model_fit.prefix_cache.*` | family defaults | config file | +| Persist prompt state across restarts | `runtime.kv_cache.disk.*` | `off` | config, environment, or `serve` flags | +| Inspect or remove disk entries | `mesh-llm kv-cache ...` | n/a | CLI | + +Model-level cache controls do not currently have CLI equivalents. Use +`~/.mesh-llm/config.toml`, or pass a different file with `mesh-llm serve +--config PATH`. The disk tier has CLI overrides for one-off runs. + +## Tune KV memory use + +Set a policy for every model under `[defaults.model_fit]`: + +```toml +[defaults.model_fit] +kv_cache_policy = "balanced" +``` + +The policies expand to these runtime choices: + +| Policy | Key/value format | KV offload | Use when | +|---|---|---|---| +| `balanced` or `auto` | Q8_0 below 50 GiB, Q4_0 at or above 50 GiB, with compatibility fallback | `auto` | General use; this is the default | +| `quality` | F16/F16 | `auto` | You prefer maximum KV precision over memory savings | +| `saver` | Q8_0/Q8_0 | enabled | You want a fixed quantized cache and forced KV offload | + +You can override one model without changing the others: + +```toml +[[models]] +model = "org/model-GGUF" + +[models.model_fit] +kv_cache_policy = "quality" +``` + +For direct control, set `cache_type_k` and `cache_type_v` to a supported dtype +of `f16`, `q8_0`, or `q4_0`. Explicit types take precedence over +`kv_cache_policy`: + +```toml +[defaults.model_fit] +cache_type_k = "q8_0" +cache_type_v = "f16" +kv_offload = "auto" +``` + +## Tune in-memory prefix reuse + +The automatic prefix cache is usually the right choice. To disable it for a +model, set: + +```toml +[defaults.model_fit] +prompt_cache = false +``` + +To keep it enabled but set explicit bounds: + +```toml +[defaults.model_fit.prefix_cache] +enabled = true +payload_mode = "auto" +min_tokens = 512 +max_entries = 256 +max_bytes = 8589934592 +shared_stride_tokens = 128 +shared_record_limit = 4 +``` + +`payload_mode = "auto"` stores resident KV for known dense models and KV plus +recurrent state for known recurrent or hybrid models. Unknown architectures do +not cache automatically. The size fields are byte counts; `max_bytes = 0` +means no explicit byte cap. Setting `prompt_cache = false` disables prefix +caching and conflicts with an explicitly enabled `prefix_cache` block. + +Per-model `[[models]]` values override `[defaults]`; explicit cache dtypes +override policy-derived dtypes; unresolved values fall back through family +policy and built-ins. Model cache changes apply when the model reloads. + +The OpenAI request field `prompt_cache_retention` accepts `in_memory` and +`24h`. Mesh records it as telemetry, but neither value currently enforces a +cache lifetime. Use the runtime limits above and the disk maintenance commands +below to control retention. + +`cache_idle_slots` limits how many reset native sessions remain available for +reuse. Unset means the runtime lane count is the bound, `0` drops every reset +lane, and a positive value adds a lower cap. `cache_ram_mib` is reserved; any +positive value currently fails model loading, so there is no configurable +host-RAM L2 tier. + +Cache matches require the exact token prefix and exact runtime identity. Mesh +does not use fuzzy or semantic prompt matching. + +## Enable durable disk caching + +Disk caching preserves reusable prompt state across process restarts. It is +node-local and disabled until you opt in. + +For an automatically sized cache: + +```toml +[runtime.kv_cache.disk] +mode = "auto" +directory = "/var/lib/mesh-llm/kv-cache" +minimum_free_mib = 16384 +``` + +For a fixed 32 GiB cap: + +```toml +[runtime.kv_cache.disk] +mode = "fixed" +directory = "/var/lib/mesh-llm/kv-cache" +budget_mib = 32768 +minimum_free_mib = 16384 +``` + +The directory must be absolute. If omitted, it is +`$MESH_LLM_HOME/kv-cache`, or `~/.mesh-llm/kv-cache`. Auto mode uses at most +20% of the filesystem capacity basis, never consumes the configured free-space +reserve, and is capped at 64 GiB. + +The same settings can be supplied for one run: + +```bash +mesh-llm serve --kv-cache-disk auto + +mesh-llm serve \ + --kv-cache-disk 32GiB \ + --kv-cache-disk-dir /var/lib/mesh-llm/kv-cache \ + --kv-cache-min-free 16GiB +``` + +Environment equivalents are `MESH_LLM_KV_CACHE_DISK`, +`MESH_LLM_KV_CACHE_DISK_DIR`, and `MESH_LLM_KV_CACHE_MIN_FREE`. CLI values win +over environment values, which win over the config file, independently for +each field. + +Invalid disk settings stop startup with a configuration error. Once a valid +configuration is running, storage trouble fails open: Mesh logs the problem +and serves the request with cold prefill. + +## Inspect and maintain the disk cache + +```bash +mesh-llm kv-cache status +mesh-llm kv-cache status --json + +mesh-llm kv-cache prune --target 16GiB --yes +mesh-llm kv-cache clear --yes +``` + +Human-readable `status` shows the effective state, configured mode, root, and +used/budget bytes. `status --json` also includes the free-space reserve, entry +counts, degradation reason, activity, reconciliation, and per-model inventory. +`prune` and `clear` only remove inactive entries. Both can be limited to an +exact numeric model identity with `--model-identity ID`. + +Budget and minimum-free changes apply live. Changing the mode or directory +requires a node restart. + +## Advanced environment controls + +These environment variables are intended for incident response and controlled +experiments: + +- Setting `SKIPPY_KV_CACHE=off` or `SKIPPY_PREFIX_CACHE=off` disables worker + prefix-cache storage even when the stage plan enables it. Other + `SKIPPY_KV_CACHE_*` tuning variables only construct cache settings when the + stage plan did not supply them. +- `MESH_LLM_DISABLE_PREFIX_AFFINITY` disables routing toward a peer with known + prefix state. `MESH_LLM_DISABLE_STICKY_ROUTING` disables sticky routing, and + `MESH_LLM_PREFIX_ONLY=1` uses the request's prefix hash as the deterministic + fallback when neither cache evidence nor a session route applies. These + change peer selection; they do not disable or resize cache storage. + +## CacheGen status + +CacheGen is an experimental compressed representation for saved KV state. The +implementation targets CPU, Metal, CUDA, and ROCm, but it is still undergoing +backend qualification, including completion and verification of the ROCm path, +and is not selected by the normal serving config or CLI today. Use the stable +resident/exact-state cache paths described above for production operation; do +not add an invented `cachegen` setting to the config. + +For every field and allowed value, see the [Config Reference](/docs/pages/config-reference/). +For disk storage details, failure modes, and recovery procedures, see the +[KV-cache disk operator guide](https://github.com/Mesh-LLM/mesh-llm/blob/main/docs/skippy/KV_CACHE_DISK.md). diff --git a/website/src/docs/pages/openai-compatible-api.md b/website/src/docs/pages/openai-compatible-api.md index b1eb9e73c0..bafd2674af 100644 --- a/website/src/docs/pages/openai-compatible-api.md +++ b/website/src/docs/pages/openai-compatible-api.md @@ -48,6 +48,11 @@ curl -s http://localhost:9337/v1/chat/completions -H "Content-Type: application/ Clients that support streamed OpenAI-compatible responses can use the same base URL. +Repeated prompt prefixes are reused automatically when the selected model +supports them. Responses report reused prompt tokens in +`usage.prompt_tokens_details.cached_tokens`. See [KV Caching](/docs/pages/kv-caching/) +for defaults, memory policy, and durable disk-cache controls. + ## Tool calling Tool-calling support depends on the selected model and the agent client. Start with console chat, then test the specific agent workflow you plan to use. From 9deb47cbf895ffec2e155a9134926e84d858bc7c Mon Sep 17 00:00:00 2001 From: scama Date: Sat, 12 Sep 2026 20:50:11 +1000 Subject: [PATCH 3/4] docs(skippy): document exact KV controls --- website/src/docs/pages/kv-caching.md | 81 +++++++++++++++++++--------- 1 file changed, 56 insertions(+), 25 deletions(-) diff --git a/website/src/docs/pages/kv-caching.md b/website/src/docs/pages/kv-caching.md index 2e088acf28..1030f4dbc7 100644 --- a/website/src/docs/pages/kv-caching.md +++ b/website/src/docs/pages/kv-caching.md @@ -30,10 +30,11 @@ With no cache settings: - Mesh enables the in-memory prefix cache with family-aware limits. Prefixes of at least 256 tokens are eligible, and Mesh selects the stored state format from the model architecture. -- The balanced KV memory policy uses Q8_0 key and value caches for models under - 50 GiB and Q4_0 for models at or above 50 GiB. If metadata proves that the - selected quantized format cannot load, the automatic policy falls back to - F16. An explicit incompatible override fails instead of changing silently. +- With no explicit K/V dtype, the resolver selects Q8_0 for both caches when + the model is smaller than 50 GiB and Q4_0 for both caches at 50 GiB or above. + It checks GGUF architecture and head dimensions before applying that default + and falls back to F16 when metadata proves the quantized layout is invalid. + An explicit incompatible dtype fails instead of changing silently. - KV offload and unified-cache behavior remain automatic. - The durable disk cache is off, so a process restart starts with an empty prompt cache and Mesh writes no prompt state to disk. @@ -44,9 +45,10 @@ These defaults apply across the supported CPU, Metal, CUDA, and ROCm runtimes. | Goal | Setting | Default | Where to set it | |---|---|---|---| -| Choose KV memory/quality policy | `model_fit.kv_cache_policy` | `balanced` | `[defaults.model_fit]` or `[models.model_fit]` | | Pin key/value cache formats | `model_fit.cache_type_k`, `model_fit.cache_type_v` | `auto` | config file | | Control KV device offload | `model_fit.kv_offload` | `auto` | config file | +| Control unified KV allocation | `model_fit.kv_unified` | `auto` | config file | +| Control the attention kernel required by quantized V | `model_fit.flash_attention` | derived from V dtype | config file | | Cap retained idle native sessions | `model_fit.cache_idle_slots` | lane count | config file | | Disable all prompt-prefix reuse | `model_fit.prompt_cache` | `auto` | config file | | Tune or disable in-memory prefix reuse | `model_fit.prefix_cache.*` | family defaults | config file | @@ -57,22 +59,60 @@ Model-level cache controls do not currently have CLI equivalents. Use `~/.mesh-llm/config.toml`, or pass a different file with `mesh-llm serve --config PATH`. The disk tier has CLI overrides for one-off runs. -## Tune KV memory use +## Choose the KV representation -Set a policy for every model under `[defaults.model_fit]`: +Configure the representation directly when you need deterministic behavior: ```toml [defaults.model_fit] -kv_cache_policy = "balanced" +cache_type_k = "q8_0" +cache_type_v = "q8_0" +kv_offload = "auto" +kv_unified = "auto" +flash_attention = "enabled" ``` -The policies expand to these runtime choices: +The relevant controls are: -| Policy | Key/value format | KV offload | Use when | -|---|---|---|---| -| `balanced` or `auto` | Q8_0 below 50 GiB, Q4_0 at or above 50 GiB, with compatibility fallback | `auto` | General use; this is the default | -| `quality` | F16/F16 | `auto` | You prefer maximum KV precision over memory savings | -| `saver` | Q8_0/Q8_0 | enabled | You want a fixed quantized cache and forced KV offload | +| Setting | Loadable embedded-runtime values | Runtime effect | +|---|---|---| +| `cache_type_k` | `auto`, `f16`, `q8_0`, `q4_0` | Storage and compute dtype for attention keys | +| `cache_type_v` | `auto`, `f16`, `q8_0`, `q4_0` | Storage and compute dtype for attention values | +| `kv_offload` | `auto`, `true`, `false` | Whether KV tensors may reside on the selected accelerator rather than host memory | +| `kv_unified` | `auto`, `true`, `false` | Whether runtime slots use the backend unified KV allocation | +| `flash_attention` | `auto`, `enabled`, `disabled` | Selects the fused attention path; a quantized V cache requires the enabled path | + +Q8_0 and Q4_0 encode values in 32-element blocks. A model whose KV head +dimension cannot satisfy that block layout cannot use the corresponding +quantized cache. Automatic selection can detect that from GGUF metadata and +fall back to F16. Explicit dtype selection bypasses that fallback so invalid +combinations fail during model load. A backend Flash Attention capability +failure is only known when the runtime loads, so metadata validation alone +cannot prove that every quantized combination will start. + +K and V may use different dtypes. Resolution order is per-model explicit +dtype, per-model preset expansion, global explicit dtype, global preset +expansion, family default, then the built-in size rule. + +The config validator currently recognizes additional GGML dtype labels that +the pinned embedded runtime does not load. The table above lists the values +accepted by `skippy_runtime::parse_cache_type`; use those values for a serving +configuration. `auto` is consumed by the resolver and does not reach that +parser. + +### Compatibility presets + +`kv_cache_policy` is an older Mesh macro layered over the technical controls. +It is still accepted, but it does not add another cache implementation: + +| Preset | Exact expansion before explicit overrides | +|---|---| +| `quality` | `cache_type_k = "f16"`, `cache_type_v = "f16"`, `kv_offload = "auto"` | +| `saver` | `cache_type_k = "q8_0"`, `cache_type_v = "q8_0"`, `kv_offload = true` | +| `auto` or `balanced` | Apply the built-in size-derived K/V dtype and leave `kv_offload = "auto"` | + +`auto` and `balanced` are aliases in the current resolver. New deployments +that need a reproducible layout should set the technical fields directly. You can override one model without changing the others: @@ -81,18 +121,9 @@ You can override one model without changing the others: model = "org/model-GGUF" [models.model_fit] -kv_cache_policy = "quality" -``` - -For direct control, set `cache_type_k` and `cache_type_v` to a supported dtype -of `f16`, `q8_0`, or `q4_0`. Explicit types take precedence over -`kv_cache_policy`: - -```toml -[defaults.model_fit] -cache_type_k = "q8_0" +cache_type_k = "f16" cache_type_v = "f16" -kv_offload = "auto" +kv_offload = false ``` ## Tune in-memory prefix reuse From 637d1d307f76ebb929f768bd17e8c6963a66badc Mon Sep 17 00:00:00 2001 From: scama Date: Sat, 12 Sep 2026 20:54:14 +1000 Subject: [PATCH 4/4] docs(skippy): remove unshipped KV presets --- website/src/docs/pages/kv-caching.md | 20 +++----------------- 1 file changed, 3 insertions(+), 17 deletions(-) diff --git a/website/src/docs/pages/kv-caching.md b/website/src/docs/pages/kv-caching.md index 1030f4dbc7..df79ac2005 100644 --- a/website/src/docs/pages/kv-caching.md +++ b/website/src/docs/pages/kv-caching.md @@ -90,9 +90,9 @@ combinations fail during model load. A backend Flash Attention capability failure is only known when the runtime loads, so metadata validation alone cannot prove that every quantized combination will start. -K and V may use different dtypes. Resolution order is per-model explicit -dtype, per-model preset expansion, global explicit dtype, global preset -expansion, family default, then the built-in size rule. +K and V may use different dtypes. For these technical fields, per-model values +override global values, which override family defaults and finally the built-in +size rule. The config validator currently recognizes additional GGML dtype labels that the pinned embedded runtime does not load. The table above lists the values @@ -100,20 +100,6 @@ accepted by `skippy_runtime::parse_cache_type`; use those values for a serving configuration. `auto` is consumed by the resolver and does not reach that parser. -### Compatibility presets - -`kv_cache_policy` is an older Mesh macro layered over the technical controls. -It is still accepted, but it does not add another cache implementation: - -| Preset | Exact expansion before explicit overrides | -|---|---| -| `quality` | `cache_type_k = "f16"`, `cache_type_v = "f16"`, `kv_offload = "auto"` | -| `saver` | `cache_type_k = "q8_0"`, `cache_type_v = "q8_0"`, `kv_offload = true` | -| `auto` or `balanced` | Apply the built-in size-derived K/V dtype and leave `kv_offload = "auto"` | - -`auto` and `balanced` are aliases in the current resolver. New deployments -that need a reproducible layout should set the technical fields directly. - You can override one model without changing the others: ```toml