Skip to content

Dev/leasing controller - #9

Merged
edlwang merged 96 commits into
mainfrom
dev/leasing-controller
Jun 23, 2026
Merged

edlwang merged 96 commits into
mainfrom
dev/leasing-controller

Conversation

@Erotemic

Copy link
Copy Markdown
Collaborator

No description provided.

Erotemic and others added 30 commits June 16, 2026 15:01
Begin the infer-stack leasing/controller redesign. Adds a new
`infer_stack.leasing` subpackage, kept cleanly separate from the existing
stateless compiler (resolver/validator/renderer):

- models.py: EndpointRequest / Lease / DeploymentGroup, compatibility_key
  (structural identity), vllm_structural/ollama_structural, and
  capacity_satisfies (subsumption, not equality).
- store.py: SqliteStore (WAL + busy_timeout + BEGIN IMMEDIATE) so the
  find-or-create-group critical section is race-safe across processes.
- ledger.py: acquire/release/renew/sweep/reclaimable_groups/status. Demand
  is a COUNT(DISTINCT lease_id) over protecting leases (computed, not stored);
  groups flip LIVE<->IDLE on demand; soft-TTL is the crash backstop.

Coalescing grain is "the thing that gets a process": per-model for vLLM,
per-daemon for Ollama (host config is the structural key, tags accumulate in
the group's served map). 17 unit tests + 3 xdoctests pass; ruff clean.

Bumps version 0.6.1 -> 0.7.0. Design: aiq-eval-runner
dev/infer-stack-redesign-critique.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add `infer_stack.leasing.catalog`: the declarative input side of the controller
and the replacement for "profiles" as the primary unit.

- New schema: models / endpoints / runtime_hosts / bundles. `Catalog.from_dict`
  / `Catalog.load` parse and validate (vllm endpoint -> model, ollama endpoint
  -> host, bundle -> endpoints, known engines).
- `resolve_endpoint` / `resolve_names` turn endpoint and bundle names into the
  ledger's EndpointRequest objects (vLLM per-model, Ollama per-daemon), so the
  catalog composes directly with the ledger from the previous commit.

Alias endpoints (two names, same model+runtime) resolve to the same compat key
and therefore coalesce onto one deployment, validating "compat key = deployment
identity, not endpoint name". 15 catalog tests (incl. ledger integration),
32 leasing tests total, 4 xdoctests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add the seam between the ledger and real serving, plus the orchestrator:

- backend.py: a 4-method `Backend` Protocol (realize/teardown/observe/
  probe_ready, all idempotent) + `Readiness` + a `MemoryBackend` for tests and
  dry-runs. This Protocol is the only surface a real backend (Compose, KubeAI)
  must implement.
- controller.py: `Controller(ledger, backend)` with the desired-vs-actual
  reconcile loop, `wait_ready`, and thin `acquire`/`release`.

Reconcile sweeps the ledger first so TTL self-enforces (a crashed job's lease
stops protecting once its TTL elapses). Desired = LIVE groups + keep-warm idle
groups; teardown falls out of the diff, no separate STOPPED state needed.
Readiness waits are scoped to the endpoints a lease requested, so a node
acquiring one Ollama tag doesn't block on a sibling tag's health.

11 controller tests (incl. catalog integration), 43 leasing tests total, 5
xdoctests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add the user-facing surface for the leasing model:

- leasing/envfile.py: the endpoint descriptor (aligned with the contracts.py
  shape: base_url / api_key_env / models) and its sourceable shell env-file;
  read_session_id recovers the session so `release --env-file` works.
- leasing/backend.py: NullBackend (dry-run) — observe() is empty so the
  controller no-op-realizes and never tears down; readiness immediate; no
  in-memory state, so it stays coherent across CLI invocations.
- cli/commands_leasing.py: acquire / serve / release / renew / run / leases,
  registered on ManageCLI. `run -- <cmd>` acquires, injects the endpoint env
  into the child, runs, releases on exit (TTL backstop), propagates exit code.

Default --backend null exercises the entire surface without docker, locking the
verb + env-file contract the consumer repos and the Compose backend build on.
Status verb is `leases` (the legacy profile model owns `status`).

13 CLI tests + an end-to-end shell smoke; 66 leasing/CLI tests total, 6
xdoctests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
First slice of the Compose backend: leasing/placement.py places GPUs across the
whole live set of deployment groups (the assignment the ledger deliberately
defers). plan_placement(groups, inventory, *, allowed_gpus, reserved, pinned,
skip_display) reuses the resolver's _first_fit / _available_gpu_indices.

Deterministic three-tier order (pinned, explicit, first-fit; groups sorted by
created_at,id): `pinned` keeps already-running groups put so reconciles don't
reshuffle live models; Ollama daemons pin explicit host gpu_indices (or [] for
CPU); vLLM first-fits tp*dp GPUs. Honors allowed_gpus, reserved (Phase-2 raw
GPUs), and display-GPU skipping. VRAM-fit deferred until ModelSpec carries
memory. Multi-node/bin-packing stay out of scope (KubeAI/Slurm).

15 placement tests, 71 leasing/CLI tests total, 7 xdoctests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Compose backend slice 2: leasing/compose.py renders a docker-compose project
straight from the live set of deployment groups (reusing profile_runtime
.vllm_args), and ComposeBackend converges the whole union on each reconcile
(docker compose up -d --remove-orphans).

- Focused renderer over reusing the legacy resolve()/render_compose_artifacts:
  avoids forcing the new model through the resolved-deployment v5 schema.
- Placement persisted in a sidecar and fed back as `pinned`, so reconciles never
  migrate a running model's GPUs; --remove-orphans is teardown.
- observe() parses `docker compose ps --format json`; docker is an injected
  run() seam (unit-tested with a stateful fake; real docker/GPU validated on a
  host).
- Controller.reconcile now prefers backend.converge(desired) when present,
  falling back to the per-group realize/teardown loop for Memory/Null backends.
- Wired `infer-stack ... --backend compose` (detect_inventory + --allowed-gpus).

Readiness this slice is "container running"; the HTTP generation probe, Ollama
pull/warmup, the LiteLLM front door (real descriptor base_url), and a converge
file-lock land in slice 3.

90 leasing/CLI tests, 7 xdoctests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Compose backend slice 3:
- LiteLLM front door (default on): render a gateway service + config that routes
  each endpoint alias to its upstream vLLM (openai/<served>) or Ollama service,
  so clients hit one stable base_url and request the public endpoint name.
- ComposeBackend.access() supplies the real base_url + per-endpoint request names
  to the env-file descriptor; the CLI's _descriptor_for prefers it over the
  --base-url placeholder. build_descriptor gained a request_names override.
  NullBackend has no access() -> CLI falls back unchanged.
- probe_ready now checks the gateway's /v1/models listing (model is routable)
  through an injected http_get seam (tested offline against a fake that lists
  whatever the rendered LiteLLM config declares).
- converge serialized with an fcntl file lock (cross-process compose-file race).

Ollama tag pull/warmup remains a readiness follow-up. 94 leasing/CLI tests, 7
xdoctests, ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The leasing placement planner was importing the resolver's *private*
_first_fit / _available_gpu_indices. Move those plus _resolve_gpu_indices into
infer_stack.hardware as public functions (available_gpu_indices / first_fit /
resolve_gpu_indices) and have both the resolver and the leasing planner reuse
them. One home for "which GPUs are available" and "first-fit N of them"; no
cross-module private imports.

Behavior-preserving: legacy resolver/profile tests (76) and leasing placement
tests pass; resolve() smoke unchanged. ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
cli/probes.py held _ready_openai_probe / _ready_ollama_probe (pure HTTP but
trapped in the cli layer); the leasing Compose backend had its own weaker
/v1/models check. Move the probes to a new layer-neutral infer_stack.probe
(openai_ready / ollama_ready) over an injectable `http` client:

- cli/probes re-exports them under the original names; the legacy wait-ready /
  switch / smoke callers are unchanged (tests patch the shared requests module
  object, so relocating the body still intercepts correctly).
- ComposeBackend.probe_ready now calls openai_ready(require_listed=True,
  require_generation=...), gaining the advertised-alias check and an optional
  real-generation rung. Its HTTP seam is now a requests-like client (.get/.post)
  matching the probe and the rest of the codebase.

Schema-tied helpers (_default_model_for_deployment,
_resolve_smoke_protocol_from_deployment) stay in cli/probes. One probe
implementation instead of two; neither consolidation adds an upward import.

Full suite: 197 passed; ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
An Ollama daemon serves tags lazily, so "container up" and "alias in the LiteLLM
config" don't mean the tag is resident. ComposeBackend.probe_ready now, for
ollama groups:
- pulls the endpoint's tag into its daemon (`docker compose exec -T ollama-<gid>
  ollama pull <tag>`), idempotently (tracked per process; failures retry via the
  poll loop), then
- forces a real generation through the LiteLLM front door, which confirms the
  tag loads and warms it.

vLLM readiness is unchanged (alias-listed by default); a new --require-generation
flag opts vLLM into the same real-generation check. Everything stays behind the
one front-door base_url; the only ollama-specific step is the pull (by service
name, no host-port tracking).

17 compose tests (3 new ollama), 88 leasing/CLI + 7 xdoctests; ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
F1 (found testing on real 2-GPU hardware): `docker compose up` rejected the
rendered file — _gpu_reservation emitted `capabilities: [["gpu"]]` but the
Compose schema wants a list of strings (`capabilities: ["gpu"]`). Fixed and
locked with a test asserting the capabilities shape.

Add dev/leasing-test-plan.md: a living runbook for exercising the Compose
backend on yardrat (9 self-contained test blocks, a findings/fixes log, debug
capture, cleanup), tailored to Turing GPUs (--dtype=half) and the display-GPU
constraint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a parametrized test that writes the rendered compose project and runs
`docker compose config -q` (schema validation only — no image pull, no
containers, no GPU), skipped where docker compose is unavailable. Covers
vllm-single / vllm-tp2 / vllm+ollama / no-litellm. This catches the class of bug
the FakeDocker tests missed (they checked the dict we build, not whether docker
accepts it — e.g. F1's capabilities shape). Sub-second per case; passes against
real docker compose.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
8 legacy CLI tests (test_cli_meta / test_cli_setup) failed on a box where
INFER_STACK_DATA_DIR was exported in the shell (from manual leasing testing).
Their subprocess run_cli helpers anchored tmp dirs with
env.setdefault('INFER_STACK_CONFIG_DIR'/'DATA_DIR', tmp_path) — setdefault is a
no-op when the var is already exported, so the tests read the real data dir
(which also held a live `infer-stack` compose project, hence the docker-ps
leak into the status test).

Force the vars to tmp_path instead. Reproduced locally by exporting the var
(identical failure), and verified all of test_cli_meta + test_cli_setup pass
*with the ambient var set* (29 passed) and clean (29 passed). Not a leasing
regression — a latent isolation hole in the subprocess helpers (the in-process
_anchor_paths already used monkeypatch.setenv, which forces). Also isort-cleaned
the two touched files.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
F8 (found on yardrat): a re-run hit the capabilities error again, this time
during `ps` — a stale docker-compose.yml from run #1 (pre-fix) was still on
disk, and reconcile's converge branch calls observe() (which runs `docker
compose ps`, validating the file) BEFORE converge() rewrites it. So a bad file
bricked acquire before the renderer fix could overwrite it.

observe() is now best-effort: catch any docker/parse error and return an empty
set, so converge overwrites the file and self-heals. A reconciler's read-side
must never be able to prevent the write that fixes the file it chokes on.

Test added with a runner that raises on `ps`. 22 compose tests; ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The Setup block's catalog heredoc had been duplicated and mangled during
editing (a broken `placement: {models)` line, a stray mkdir/second `cat <<YAML`
injected mid-YAML, and an `echo` mashed into a leftover fragment), so the
copy-pasted catalog.yaml was invalid. Collapse it back to one valid heredoc.
Verified the result parses and resolves through Catalog.load (all 4 endpoints +
the bundle).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…vLLM flag

Restore the legacy "infer-stack owns the secret, you fetch it" ergonomic that
the leasing model had regressed (it made the user invent + export
LITELLM_MASTER_KEY):

- ComposeBackend.master_key() manages the key in the state dir's .env (reused if
  you pin one, else generated via env_utils.ensure_secret), bakes the literal
  into the LiteLLM service (so container + probe agree regardless of shell env),
  uses it for the readiness probe, and access() returns it.
- The --env-file descriptor now carries OPENAI_API_KEY (+ OPENAI_BASE_URL), so
  `source is.env` fully configures an OpenAI client — no manual export.
- New `infer-stack secrets [KEY]` prints the managed secrets, e.g.
  $(infer-stack secrets LITELLM_MASTER_KEY).

Also F9: drop --disable-log-requests from vllm_args (vLLM v0.19.1 rejects it ->
vLLM crash -> LiteLLM "Connection error"). Update the test plan to drop the
manual key export and source the env-file instead.

208 tests (11 new), ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Post-refactor the legacy meta commands ignored the leasing world. Now:
- `infer-stack config paths` gained a `leasing` group (lease ledger, compose
  state dir + rendered docker-compose.yml / litellm_config.yaml / secrets .env /
  sidecar), and is also exposed top-level as `infer-stack paths`.
- `infer-stack status` prints a one-line leasing summary (active leases / live
  groups) pointing at `infer-stack leases`, when a ledger exists.

13 meta tests (4 new) + the legacy status test pass; touched files ruff-clean
(the remaining cli/ ruff findings are pre-existing in untouched files).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The manual chat-completions curl omitted `-H "Content-Type: application/json"`,
so `curl -d` sent form-encoding and LiteLLM saw `model=None`. Added the header
to the three chat curls. Not an infer-stack bug — the serving path works
(vLLM healthy, /v1/models lists the alias, acquire --require-generation already
did a real chat through the gateway).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…/stop)

`infer-stack logs -f` (and ps/restart/pull/start/stop) failed for a leasing
user with "No config.yaml" — they were bound to the legacy model
(config_for_runtime + generated/docker-compose.yml). Add _day2_compose_base()
which prefers the leasing Compose deployment (data_root/leasing/compose, project
`infer-stack`, no config.yaml needed) and falls back to the legacy rendered
stack; route all six generic wrappers through it. Add a shared LEASING_PROJECT
constant. ollama-* wrappers stay legacy (fixed `ollama` service name).

Also point the test plan's inspection steps at `infer-stack ps`/`logs`.

212 tests (1 new), ruff clean; smoke: `infer-stack ps`/`logs` run against the
leasing compose with no config.yaml.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The runbook (dev/leasing-test-plan.md) was copy-paste only. This adds an
executable harness that runs the same scenarios, asserts the wiring, and
writes one self-contained results dir (report.md + per-step logs + the run's
rendered compose/litellm/.env) to rsync back.

- lib.sh: step/assert DSL; each step -> one results.jsonl record; no set -e.
- render_report.py (stdlib): per-step JSON emitter + report.md assembler with
  a Correctness/Efficiency/Ergonomics rollup and failures-first log tails.
- catalog.yaml: committed (validated via Catalog.load/resolve_names).
- tests/: non-serving tiers (env, dry-run, ergonomics, negatives) + GPU
  serving tiers (single-vllm, coalescing, dedicated/F5, TTL, run-wrapper,
  ollama, concurrency) + always-run cleanup.
- run.sh: ./run.sh [--gpu] [--only ...]; fresh data dir per run inside the
  results dir; prints the rsync line.

Non-serving tiers smoke-tested locally against the null backend (16/16).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Test scripts are run via `bash "$f"` (a subprocess), not sourced, so a
top-level `return 0` after the skip block errored ("can only return from a
function or sourced script") and execution fell THROUGH into the real
compose acquire on a non-GPU box. Use `exit 0` so the script ends cleanly
and run.sh continues to the next tier. Verified with a full non-GPU run
(18 passed, 19 skipped, 0 failed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…lkthrough

F5 was the blocker for using yardrat's display-attached GPU 1. The placer's
skip_display had no CLI knob, so only GPU 0 was placeable. Add a
_DisplayGpuMixin (--include-display-gpus) on the leasing common flags, wired to
ComposeBackend(skip_display=...). Default still skips display GPUs.

- e2e tier 45_both_gpus: acquire two distinct models with the flag, assert the
  backend sidecar placed them on two distinct GPUs and both route via the
  gateway.
- dev/leasing-demo.md: a demo-minded companion to leasing-test-plan.md — real
  user-config + storage setup, deploy a 14B model as a standing service, drive
  it from Open WebUI, switch models. Copy-pastable, no cross-block env deps.
  Ends with a ranked list of ergonomic smells the walkthrough surfaced.
- Mark F5 FIXED in the test plan; CHANGELOG + unit test for the flag wiring.

Full suite: 213 passed (was 212).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
GPU e2e found 50_coalescing/coalesce-acquire timing out: a 2nd alias coalesced
onto a live group (endpoint with public_name) was added to the rendered LiteLLM
config, but the running gateway never picked it up, so its readiness probe hung
to the 1200s timeout.

LiteLLM reads its bind-mounted model_list only at startup. converge rewrote the
file but `docker compose up -d` left the litellm container running (service spec
unchanged), so live routes went stale. Stamp a content-hash label
(infer-stack.config-hash) on the litellm service so converge recreates it iff
the routing changes; idempotent otherwise. Unit-tested.

Also tighten the e2e container-count check to count only LIVE vllm groups (a
reclaim:stop group from an earlier tier lingers as state=stopped in the ledger).

214 passed (was 213).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…rsync

- run.sh gains a finish() EXIT trap (idempotent) that assembles the report,
  writes rsync-back.sh, and prints the summary + rsync line; INT/TERM exit so
  finish runs on interrupt. So an interrupted GPU run still returns a reviewable
  report (verified normal + SIGINT-mid-run paths).
- The data dir lives inside the results dir and holds multi-GB HF/kernel caches.
  The printed rsync line + rsync-back.sh now --exclude *-cache/, ollama/,
  open-webui/, postgres-*/, runtime/ (keeping leasing/: ledger, compose, litellm
  config, .env).
- Move per-step scratch (.lastout/.asserts/.notes) out of logs/ into .scratch/
  (also excluded) so the rsync'd logs/ holds only real .log files.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ry later)

The config-hash label is an interim correctness fix that reintroduces a brief
gateway blip on model add/remove (legacy kept LiteLLM up across switches). Record
the decision durably: #1 (static superset route table) is the chosen direction;
#2 (DB-backed live model management via LiteLLM admin API) was tried before with
issues and should be retried in the future. Includes scope notes for #1 (catalog
must be threaded to render; upstream service names must become deterministic) so
it can be picked up cold.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ign notes

- run.sh prints the (cache-excluding) rsync line at the START and writes
  rsync-back.sh before any test, so partial results can be synced mid-run.
- Reset the compose stack between GPU tiers. The 80/85/90 failures were test
  isolation, not a product bug: a keep-warm qwen-15b group from tier 45 survived
  release and held GPU 0 (the only usable GPU on yardrat), so later acquires
  couldn't place and timed out. A between-tier 'compose down' isolates tiers
  (weights stay cached → reload, not re-download).
- leasing-followups.md: group-id determinism recommendation (prefer compat-key
  hash; network-alias alternative; per-user NOT recommended — fights
  coalescing), the /health vs /v1/models split to indicate undeployed models for
  the #1 superset, and the keep-warm-vs-live-demand scheduling question.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Keep-warm vs demand resolved: never reclaim a LIVE group (a held lease protects
it regardless of traffic); an IDLE group (demand 0 — all leases released or
TTL-expired/unrenewed) is reclaimable, so a new live-demand lease may evict it to
free the GPU. keep-warm = 'stay resident if there's room, yield to real demand'.

Group-id determinism: approved the compat-key hash approach (option 1), to be
done after the GPU e2e dashboards are green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The wall-clock was dominated by vLLM startup (torch.compile + CUDA graph
capture, ~model-size-independent) and a couple of deliberate waits, not weights.

- Every vLLM endpoint now runs --enforce-eager (skip compile/graph capture) with
  max_model_len 2048 (smaller KV alloc / capture). Biggest lever.
- Swap the test models to the smallest sane ungated chat models: SmolLM2-135M
  (was Qwen2.5-0.5B) and SmolLM2-360M for the 2nd distinct model; ollama uses
  smollm2:135m. Endpoint names kept (qwen-*) so no test-script churn; coalescing
  + distinctness verified.
- Trim intentional waits: tier 60 dedicated acquire --timeout 180->45 (it's an
  expected place-failure we just confirm), tier 70 TTL 60s/sleep 75 -> 30s/40s,
  tier 85 ollama --timeout 900->300 (success returns early; this is the failure
  ceiling).

Success-case acquire timeouts (40/50) left high — a timeout is a ceiling, not a
wait; a ready model returns immediately.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ComposeBackend.converge always ran 'docker compose up -d', but with zero services
(last reclaim:stop lease released) 'up' errors 'no service selected', so the
release reconcile crashed and 'infer-stack run' returned non-zero even though the
job succeeded. Converge now 'down's the project when there are no services.

Latent until the between-tier reset isolated deployments so a release could
actually converge to empty; the GPU e2e 80_run_wrapper tier caught it (only tier
that surfaces the release exit code). Unit-tested converge([]) -> down, not up.
215 passed (was 214).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
GPU tiers 80/85/90 hung because the shared ledger accumulated state across tiers
and the between-tier reset only downed containers. A tier-60 acquire left its
lease ACTIVE (a not-ready/uncleaned acquire keeps its lease), so its qwen-small
group stayed LIVE in the ledger; every later reconcile re-realized it on GPU 0
and starved ollama / run-exit-code / the concurrency tier.

- reset_between_tiers now also wipes the ledger (ledger.db + wal/shm) and downs
  by project name, so each GPU tier starts from zero leases/groups.
- run() prints a '… still running (Ns)' heartbeat every 20s (and reaps the
  heartbeat's sleep child) so long cold-starts don't look hung.
- New cleanup.sh: down the 'infer-stack' project (+ stragglers by label) and
  optionally wipe a run's ledger — for recovering after a Ctrl-C.
- leasing-followups.md: record the underlying product issues (not-ready acquire
  leaks an active lease; re-acquire spawns a new group instead of reviving an
  idle one; ledger-vs-reality drift re-realizes leaked LIVE groups).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Erotemic and others added 29 commits June 19, 2026 00:26
Add a 'deployment' column to the leases pane (the group id(s) a lease
holds, matching the groups pane id) and 'leases'/'held by' columns to the
groups pane in place of the opaque 'demand', so the many-to-one link is
visible. Cursor movement narrates the link in the status bar. Make
action-bar buttons 1 row tall.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The ledger object is a Lease; the session_id/lease dual vocabulary was a
source of confusion. Hard-rename throughout: env-file key
INFER_STACK_SESSION_ID -> INFER_STACK_LEASE_ID, release/renew --session ->
--lease, JSON key session_id -> lease_id, read_session_id -> read_lease_id,
generated id prefix sess- -> lease-. No back-compat alias (per decision).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Deployment group was an unintuitive name; the object is a Deployment and the
mental model is 'many leases -> one deployment'. Hard rename across core, CLI,
TUI, docs: DeploymentGroup->Deployment, GroupState->DeploymentState,
group_id(s)->deployment_id(s), ledger methods, leases --json key
groups->deployments, and the SQLite groups table->deployments
(claims.group_id->deployment_id). No back-compat / no migration: delete any
existing ledger DB and it is recreated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Remove the superseded 'legacy' profile/active-profile surface (no back-compat,
pre-release): the legacy command group + commands_profile/commands_smoke,
renderer/benchmark/verification/contracts, the active-profile runtime verbs
(Up/Down/Purge/Deploy/Env/Ollama*) and the _ensure_rendered/RenderCLI hook, plus
their tests (~5k LOC). Keep the ollama + kubeai concepts (kubeai_ops,
backends/kubeai_renderer, catalog engine:ollama) for when those backends land.

Rewrite commands_runtime.py: 'stack' day-2 wrappers now target only the leasing
compose project, and 'infer-stack status' becomes a holistic leasing-native
overview (locations, backend, leases/deployments summary, dig-deeper hints).

resolver/validator/old catalog.py + legacy context plan helpers are deferred to
a follow-up (they're load-bearing for context.py's kept helpers).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Phase 2a of the legacy sweep. The pre-leasing resolver + validator (and the
cli/compose, cli/probes shims that only the deleted commands used) are gone;
cli/context is carved down to the two helpers the leasing surface needs
(_apply_path_overrides, effective_inventory). config paths now reports only
current locations (config_root, settings.yaml, catalog.yaml, data_root, the
leasing ledger/compose files) instead of the old profile artifacts.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
backends/compose_renderer is replaced by leasing/compose; remove it and slim
backends/__init__ to the kubeai renderer, which is retained as the kubeai
conceptual path. Note the old top-level catalog.py + dead config.py helpers as
an internal-only follow-up (no CLI surface).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The add-endpoint wizard now exposes tensor-parallel size, max model len, GPU
memory fraction, raw extra vLLM args, and reclaim policy (mirrors
catalog endpoint add). Add an Edit action (blocked while an endpoint is actively
served) and Remove actions for endpoints and models, with a confirm dialog and
the validating writer rejecting a model still referenced by an endpoint.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Wrap the multipane monitor in a top-level TabbedContent: a Dashboard tab (the
existing panes) and a new Settings tab that reads/writes the durable settings
(backend, data dir, Open WebUI, reverse proxy, skip-display-GPUs) to
settings.yaml. The dashboard body is composed via a helper so its layout is
unchanged. Textual's command palette (ctrl+p) covers action search.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a Control tab to the docker pane showing the rendered compose-file path and
Compose up / Compose down buttons that drive the leasing project through the
captured docker runner (progress shows in the Logs tab).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Models table now shows the quantization (from the catalog) and a 'cached' flag
derived from a fast existence check against the HF hub cache dir
(state.hf_cache/hub/models--Org--Model) — no slow du.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add the two splitters Jon called out: #csplit (catalog endpoints|models) and
#lsplit (leases|deployments), reusing the fixed-height-neighbor + 1fr-neighbor
pattern. All four panes pairs are now drag-resizable.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
From live testing:
- Endpoint wizard now labels every field and adapts to engine: vLLM exposes
  tensor-parallel, data-parallel, max-model-len, gpu-mem, max-seqs,
  prefix-caching, extra-args; Ollama exposes host + free-form KEY=VALUE runtime.
  (CLI parity holds: these map to catalog runtime keys / --runtime / --extra-args.)
- Move the API tester to its own top-level tab (more room for the monitor).
- Localize Add/Edit/Remove to the endpoints + models panels; drop the bottom
  button stack; move Suggest to the Settings tab.
- Vertical splitter now drags the full width range (clamped to terminal width),
  not just the middle; fix the endpoints|models drag direction.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Convert the leases and deployments monitor panes to Collapsibles (docker and
system already were), so all four monitor panes collapse to focus. Collapse and
a fixed-height drag bar can't coexist on one pane in Textual (the explicit
height defeats collapse), so the leases|deployments drag bar (lsplit) is dropped
in favor of collapse there; the catalog stays drag-sized (endpoints|models
csplit) since you act on it constantly. Pane titles carry live counts.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Excise the superseded profile-catalog machinery from config.py (builtin_*_catalog,
the file/overlay/merge chain, normalized_catalogs, initial_config) — all dead
externally; the leasing path only uses normalized_output/normalized_state/
default_state_paths + the port/image constants. Delete the now-orphaned
infer_stack/catalog.py (leasing/catalog.py is the real one — removes the 'two
catalogs' confusion) and the profile-era templates (default-*.yaml + the legacy
.j2 renderer templates); keep suggestion-pool.yaml.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The 4 endpoint buttons used width:1fr but Button's default min-width (16) kept
them from shrinking, so Edit/Remove overflowed off the sidebar and only appeared
after widening it. Add min-width:0 to the catalog action buttons (and drop the
double-width glyphs) so Serve/Add/Edit/Remove all fit at the default width.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The API tab now shows the gateway + Open WebUI URLs (ctrl+click to open),
adds a 'List models' check (GET /v1/models on the gateway) alongside Send /
Test-all, shows a live curl preview with a 'Copy curl' button, and an
'Open WebUI' button. Clipboard via Textual OSC 52: Copy curl, plus 'y' to copy
the status line.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Textual's content markup rejects the ':' in [link=http://...], crashing
get_content_height when the gateway/Open WebUI URLs were shown. Render the URLs
as plain text; ctrl+click the line and the Open WebUI button still open it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
vLLM OOM'd at startup (KV cache too small) because suggest overrode a model's
hand-tuned gpu_memory_utilization with a bare footprint ratio
(min_vram*1.3/host_mem), which on a big GPU lands *below* the default the pool
author sized for the context's KV cache (e.g. smollm2-1.7b: 0.22 vs 0.4 → only
~1.43 GiB KV, under the ~1.5 GiB an 8192 context needs). The footprint estimate
must only raise the reservation, never lower it: util = max(footprint, default),
still clamped [0.2, 0.92]. Regression test added.

Existing catalogs keep the old value — re-run catalog suggest --apply --force
or edit the endpoint to pick up the fix.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
System pane: the polling gate trusted Collapsible.Toggled, which doesn't fire on
every toggle path, so expanding System never flipped the gate and the GPU table
stayed empty. Read the live collapse + active-tab state on the UI thread
(_sync_pane_state) before each refresh instead; the gate now self-heals every
interval. Regression test added.

Clipboard: many terminals don't honor the OSC 52 escape Textual uses (and it
blocks native select+copy), so Copy curl / 'y' silently failed. Prefer OS
clipboard tools (wl-copy / xclip / xsel / pbcopy) and fall back to OSC 52;
report success/failure in the status line.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…odels

Open WebUI was hard-wired to the LiteLLM front door with
ENABLE_OLLAMA_API=False, so the UI could only see models declared as
catalog endpoints — you could not pull/manage models from it. That is the
opposite of "the ollama daemon manages its own models."

Rework the front door so it matches a hand-run ollama + Open WebUI stack:

* Open WebUI now holds two independent connections, wired from the rendered
  services: an OpenAI connection (the LiteLLM gateway when on, else a single
  upstream's own /v1, else disabled) for chat, and a native Ollama connection
  pointed straight at any Ollama daemon for pull/run/delete + on-demand load.
* The LiteLLM gateway is optional: --litellm/--no-litellm flag, `config set
  litellm`, and _resolve_litellm in the CLI. Open WebUI renders without a
  gateway whenever there is an upstream to point at (and is a standing front
  door only when the gateway is on).
* access() reports the UI URL when the gateway is off but the UI is on.
* probe_ready pulls an Ollama tag regardless of the gateway, so a lean
  --no-litellm stack still makes its anchor model present.

GPU pinning for the daemon already works via runtime_hosts placement
(gpu_indices -> CUDA_VISIBLE_DEVICES).

Adds dev/ollama-openwebui-tutorial.md (CLI-driven, 2x1080ti worked example)
and compose tests for the three wirings (direct ollama, ollama+litellm,
ui-without-litellm) + the no-litellm pull. 243 tests pass; ruff clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Create docs/source/manual/ and wire it into the Sphinx toctree (User Manual
section). Move the two current, leasing-era user guides there:

* ollama-openwebui-tutorial.md (new) — self-managing Ollama + Open WebUI
* leasing-demo.md (from dev/) — the happy-path walkthrough

Update the live references (README.md, dev/leasing-followups.md) and the
CHANGELOG path. The remaining docs/ and docs/demos/ guides are profile-era
(pre-leasing) and were intentionally left out of the manual — they document
a model that no longer exists and need a rewrite/removal pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…n yardrat

Add real-GPU coverage for the new leasing features and fix a GPU-pin bug they
surfaced. Leverages yardrat's 2nd GPU (the display-attached GPU 1).

Fix: an Ollama daemon pinned to a non-zero GPU silently ran on CPU. The docker
device reservation (device_ids) exposes only the pinned GPU and the NVIDIA
runtime renumbers it to 0 inside the container, but _ollama_service also set
CUDA_VISIBLE_DEVICES to the *host* index, so pinning to GPU 1 pointed at a
device that didn't exist in-container. Pin by the reservation alone, like vLLM.
Updated the unit test that had enshrined the bug.

e2e harness (dev/e2e_tests):
* 86_ollama_lean (new) — --no-litellm: no gateway, Open WebUI wired straight at
  the daemon's native API, tag still pulled, chat hits the daemon directly,
  status reports the UI URL.
* 88_gpu_pinning (new) — pin a daemon to GPU 1 (--include-display-gpus): assert
  device_ids=[1], no host-index CUDA_VISIBLE_DEVICES, and it runs ON the GPU
  (ollama ps != 100% CPU). Catches the bug above on real hardware.
* 85_ollama — now asserts Open WebUI gets BOTH connections (gateway OpenAI +
  native Ollama) when the gateway is on.
* catalog.yaml — add an ollama host pinned to GPU 1 + smol-ollama-gpu1 endpoint.
* inspect.py (new) — compose-file inspector so the tiers assert wiring without
  brittle YAML grepping. README tier table updated.

243 unit tests pass; ruff clean; catalog validates.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ew tiers

Two harness bugs the first yardrat run surfaced:

* `--include-display-gpus` no longer exists — placement uses every GPU by
  default now (opt out with `--skip-display-gpus`). The flag made acquire
  fast-fail with "unrecognized arguments", so 88_gpu_pinning never rendered and
  inspect.py read a stale compose (wrong device_ids). Dropped the flag from
  88_gpu_pinning and from the pre-existing 45_both_gpus (broken the same way);
  GPU 1 is placeable without it. Updated the README GPU note.
* 86_ollama_lean asserted the Open WebUI URL appears in `infer-stack status`,
  but status doesn't print it (acquire/serve do). Assert the UI front-door port
  from the rendered compose instead (inspect.py now prints open-webui host_port)
  and just sanity-check `status` exits 0.

Verified inspect.py output matches every assertion against rendered compose.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@edlwang
edlwang merged commit bfb7dbe into main Jun 23, 2026
7 of 63 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants