Repository navigation
Dev/leasing controller - #9
Merged
Merged
Conversation
Begin the infer-stack leasing/controller redesign. Adds a new `infer_stack.leasing` subpackage, kept cleanly separate from the existing stateless compiler (resolver/validator/renderer): - models.py: EndpointRequest / Lease / DeploymentGroup, compatibility_key (structural identity), vllm_structural/ollama_structural, and capacity_satisfies (subsumption, not equality). - store.py: SqliteStore (WAL + busy_timeout + BEGIN IMMEDIATE) so the find-or-create-group critical section is race-safe across processes. - ledger.py: acquire/release/renew/sweep/reclaimable_groups/status. Demand is a COUNT(DISTINCT lease_id) over protecting leases (computed, not stored); groups flip LIVE<->IDLE on demand; soft-TTL is the crash backstop. Coalescing grain is "the thing that gets a process": per-model for vLLM, per-daemon for Ollama (host config is the structural key, tags accumulate in the group's served map). 17 unit tests + 3 xdoctests pass; ruff clean. Bumps version 0.6.1 -> 0.7.0. Design: aiq-eval-runner dev/infer-stack-redesign-critique.md. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add `infer_stack.leasing.catalog`: the declarative input side of the controller and the replacement for "profiles" as the primary unit. - New schema: models / endpoints / runtime_hosts / bundles. `Catalog.from_dict` / `Catalog.load` parse and validate (vllm endpoint -> model, ollama endpoint -> host, bundle -> endpoints, known engines). - `resolve_endpoint` / `resolve_names` turn endpoint and bundle names into the ledger's EndpointRequest objects (vLLM per-model, Ollama per-daemon), so the catalog composes directly with the ledger from the previous commit. Alias endpoints (two names, same model+runtime) resolve to the same compat key and therefore coalesce onto one deployment, validating "compat key = deployment identity, not endpoint name". 15 catalog tests (incl. ledger integration), 32 leasing tests total, 4 xdoctests, ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add the seam between the ledger and real serving, plus the orchestrator: - backend.py: a 4-method `Backend` Protocol (realize/teardown/observe/ probe_ready, all idempotent) + `Readiness` + a `MemoryBackend` for tests and dry-runs. This Protocol is the only surface a real backend (Compose, KubeAI) must implement. - controller.py: `Controller(ledger, backend)` with the desired-vs-actual reconcile loop, `wait_ready`, and thin `acquire`/`release`. Reconcile sweeps the ledger first so TTL self-enforces (a crashed job's lease stops protecting once its TTL elapses). Desired = LIVE groups + keep-warm idle groups; teardown falls out of the diff, no separate STOPPED state needed. Readiness waits are scoped to the endpoints a lease requested, so a node acquiring one Ollama tag doesn't block on a sibling tag's health. 11 controller tests (incl. catalog integration), 43 leasing tests total, 5 xdoctests, ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add the user-facing surface for the leasing model: - leasing/envfile.py: the endpoint descriptor (aligned with the contracts.py shape: base_url / api_key_env / models) and its sourceable shell env-file; read_session_id recovers the session so `release --env-file` works. - leasing/backend.py: NullBackend (dry-run) — observe() is empty so the controller no-op-realizes and never tears down; readiness immediate; no in-memory state, so it stays coherent across CLI invocations. - cli/commands_leasing.py: acquire / serve / release / renew / run / leases, registered on ManageCLI. `run -- <cmd>` acquires, injects the endpoint env into the child, runs, releases on exit (TTL backstop), propagates exit code. Default --backend null exercises the entire surface without docker, locking the verb + env-file contract the consumer repos and the Compose backend build on. Status verb is `leases` (the legacy profile model owns `status`). 13 CLI tests + an end-to-end shell smoke; 66 leasing/CLI tests total, 6 xdoctests, ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
First slice of the Compose backend: leasing/placement.py places GPUs across the whole live set of deployment groups (the assignment the ledger deliberately defers). plan_placement(groups, inventory, *, allowed_gpus, reserved, pinned, skip_display) reuses the resolver's _first_fit / _available_gpu_indices. Deterministic three-tier order (pinned, explicit, first-fit; groups sorted by created_at,id): `pinned` keeps already-running groups put so reconciles don't reshuffle live models; Ollama daemons pin explicit host gpu_indices (or [] for CPU); vLLM first-fits tp*dp GPUs. Honors allowed_gpus, reserved (Phase-2 raw GPUs), and display-GPU skipping. VRAM-fit deferred until ModelSpec carries memory. Multi-node/bin-packing stay out of scope (KubeAI/Slurm). 15 placement tests, 71 leasing/CLI tests total, 7 xdoctests, ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Compose backend slice 2: leasing/compose.py renders a docker-compose project straight from the live set of deployment groups (reusing profile_runtime .vllm_args), and ComposeBackend converges the whole union on each reconcile (docker compose up -d --remove-orphans). - Focused renderer over reusing the legacy resolve()/render_compose_artifacts: avoids forcing the new model through the resolved-deployment v5 schema. - Placement persisted in a sidecar and fed back as `pinned`, so reconciles never migrate a running model's GPUs; --remove-orphans is teardown. - observe() parses `docker compose ps --format json`; docker is an injected run() seam (unit-tested with a stateful fake; real docker/GPU validated on a host). - Controller.reconcile now prefers backend.converge(desired) when present, falling back to the per-group realize/teardown loop for Memory/Null backends. - Wired `infer-stack ... --backend compose` (detect_inventory + --allowed-gpus). Readiness this slice is "container running"; the HTTP generation probe, Ollama pull/warmup, the LiteLLM front door (real descriptor base_url), and a converge file-lock land in slice 3. 90 leasing/CLI tests, 7 xdoctests, ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Compose backend slice 3: - LiteLLM front door (default on): render a gateway service + config that routes each endpoint alias to its upstream vLLM (openai/<served>) or Ollama service, so clients hit one stable base_url and request the public endpoint name. - ComposeBackend.access() supplies the real base_url + per-endpoint request names to the env-file descriptor; the CLI's _descriptor_for prefers it over the --base-url placeholder. build_descriptor gained a request_names override. NullBackend has no access() -> CLI falls back unchanged. - probe_ready now checks the gateway's /v1/models listing (model is routable) through an injected http_get seam (tested offline against a fake that lists whatever the rendered LiteLLM config declares). - converge serialized with an fcntl file lock (cross-process compose-file race). Ollama tag pull/warmup remains a readiness follow-up. 94 leasing/CLI tests, 7 xdoctests, ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The leasing placement planner was importing the resolver's *private* _first_fit / _available_gpu_indices. Move those plus _resolve_gpu_indices into infer_stack.hardware as public functions (available_gpu_indices / first_fit / resolve_gpu_indices) and have both the resolver and the leasing planner reuse them. One home for "which GPUs are available" and "first-fit N of them"; no cross-module private imports. Behavior-preserving: legacy resolver/profile tests (76) and leasing placement tests pass; resolve() smoke unchanged. ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
cli/probes.py held _ready_openai_probe / _ready_ollama_probe (pure HTTP but trapped in the cli layer); the leasing Compose backend had its own weaker /v1/models check. Move the probes to a new layer-neutral infer_stack.probe (openai_ready / ollama_ready) over an injectable `http` client: - cli/probes re-exports them under the original names; the legacy wait-ready / switch / smoke callers are unchanged (tests patch the shared requests module object, so relocating the body still intercepts correctly). - ComposeBackend.probe_ready now calls openai_ready(require_listed=True, require_generation=...), gaining the advertised-alias check and an optional real-generation rung. Its HTTP seam is now a requests-like client (.get/.post) matching the probe and the rest of the codebase. Schema-tied helpers (_default_model_for_deployment, _resolve_smoke_protocol_from_deployment) stay in cli/probes. One probe implementation instead of two; neither consolidation adds an upward import. Full suite: 197 passed; ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
An Ollama daemon serves tags lazily, so "container up" and "alias in the LiteLLM config" don't mean the tag is resident. ComposeBackend.probe_ready now, for ollama groups: - pulls the endpoint's tag into its daemon (`docker compose exec -T ollama-<gid> ollama pull <tag>`), idempotently (tracked per process; failures retry via the poll loop), then - forces a real generation through the LiteLLM front door, which confirms the tag loads and warms it. vLLM readiness is unchanged (alias-listed by default); a new --require-generation flag opts vLLM into the same real-generation check. Everything stays behind the one front-door base_url; the only ollama-specific step is the pull (by service name, no host-port tracking). 17 compose tests (3 new ollama), 88 leasing/CLI + 7 xdoctests; ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
F1 (found testing on real 2-GPU hardware): `docker compose up` rejected the rendered file — _gpu_reservation emitted `capabilities: [["gpu"]]` but the Compose schema wants a list of strings (`capabilities: ["gpu"]`). Fixed and locked with a test asserting the capabilities shape. Add dev/leasing-test-plan.md: a living runbook for exercising the Compose backend on yardrat (9 self-contained test blocks, a findings/fixes log, debug capture, cleanup), tailored to Turing GPUs (--dtype=half) and the display-GPU constraint. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a parametrized test that writes the rendered compose project and runs `docker compose config -q` (schema validation only — no image pull, no containers, no GPU), skipped where docker compose is unavailable. Covers vllm-single / vllm-tp2 / vllm+ollama / no-litellm. This catches the class of bug the FakeDocker tests missed (they checked the dict we build, not whether docker accepts it — e.g. F1's capabilities shape). Sub-second per case; passes against real docker compose. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
8 legacy CLI tests (test_cli_meta / test_cli_setup) failed on a box where
INFER_STACK_DATA_DIR was exported in the shell (from manual leasing testing).
Their subprocess run_cli helpers anchored tmp dirs with
env.setdefault('INFER_STACK_CONFIG_DIR'/'DATA_DIR', tmp_path) — setdefault is a
no-op when the var is already exported, so the tests read the real data dir
(which also held a live `infer-stack` compose project, hence the docker-ps
leak into the status test).
Force the vars to tmp_path instead. Reproduced locally by exporting the var
(identical failure), and verified all of test_cli_meta + test_cli_setup pass
*with the ambient var set* (29 passed) and clean (29 passed). Not a leasing
regression — a latent isolation hole in the subprocess helpers (the in-process
_anchor_paths already used monkeypatch.setenv, which forces). Also isort-cleaned
the two touched files.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
F8 (found on yardrat): a re-run hit the capabilities error again, this time during `ps` — a stale docker-compose.yml from run #1 (pre-fix) was still on disk, and reconcile's converge branch calls observe() (which runs `docker compose ps`, validating the file) BEFORE converge() rewrites it. So a bad file bricked acquire before the renderer fix could overwrite it. observe() is now best-effort: catch any docker/parse error and return an empty set, so converge overwrites the file and self-heals. A reconciler's read-side must never be able to prevent the write that fixes the file it chokes on. Test added with a runner that raises on `ps`. 22 compose tests; ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The Setup block's catalog heredoc had been duplicated and mangled during
editing (a broken `placement: {models)` line, a stray mkdir/second `cat <<YAML`
injected mid-YAML, and an `echo` mashed into a leftover fragment), so the
copy-pasted catalog.yaml was invalid. Collapse it back to one valid heredoc.
Verified the result parses and resolves through Catalog.load (all 4 endpoints +
the bundle).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…vLLM flag Restore the legacy "infer-stack owns the secret, you fetch it" ergonomic that the leasing model had regressed (it made the user invent + export LITELLM_MASTER_KEY): - ComposeBackend.master_key() manages the key in the state dir's .env (reused if you pin one, else generated via env_utils.ensure_secret), bakes the literal into the LiteLLM service (so container + probe agree regardless of shell env), uses it for the readiness probe, and access() returns it. - The --env-file descriptor now carries OPENAI_API_KEY (+ OPENAI_BASE_URL), so `source is.env` fully configures an OpenAI client — no manual export. - New `infer-stack secrets [KEY]` prints the managed secrets, e.g. $(infer-stack secrets LITELLM_MASTER_KEY). Also F9: drop --disable-log-requests from vllm_args (vLLM v0.19.1 rejects it -> vLLM crash -> LiteLLM "Connection error"). Update the test plan to drop the manual key export and source the env-file instead. 208 tests (11 new), ruff clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Post-refactor the legacy meta commands ignored the leasing world. Now: - `infer-stack config paths` gained a `leasing` group (lease ledger, compose state dir + rendered docker-compose.yml / litellm_config.yaml / secrets .env / sidecar), and is also exposed top-level as `infer-stack paths`. - `infer-stack status` prints a one-line leasing summary (active leases / live groups) pointing at `infer-stack leases`, when a ledger exists. 13 meta tests (4 new) + the legacy status test pass; touched files ruff-clean (the remaining cli/ ruff findings are pre-existing in untouched files). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The manual chat-completions curl omitted `-H "Content-Type: application/json"`, so `curl -d` sent form-encoding and LiteLLM saw `model=None`. Added the header to the three chat curls. Not an infer-stack bug — the serving path works (vLLM healthy, /v1/models lists the alias, acquire --require-generation already did a real chat through the gateway). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…/stop) `infer-stack logs -f` (and ps/restart/pull/start/stop) failed for a leasing user with "No config.yaml" — they were bound to the legacy model (config_for_runtime + generated/docker-compose.yml). Add _day2_compose_base() which prefers the leasing Compose deployment (data_root/leasing/compose, project `infer-stack`, no config.yaml needed) and falls back to the legacy rendered stack; route all six generic wrappers through it. Add a shared LEASING_PROJECT constant. ollama-* wrappers stay legacy (fixed `ollama` service name). Also point the test plan's inspection steps at `infer-stack ps`/`logs`. 212 tests (1 new), ruff clean; smoke: `infer-stack ps`/`logs` run against the leasing compose with no config.yaml. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The runbook (dev/leasing-test-plan.md) was copy-paste only. This adds an executable harness that runs the same scenarios, asserts the wiring, and writes one self-contained results dir (report.md + per-step logs + the run's rendered compose/litellm/.env) to rsync back. - lib.sh: step/assert DSL; each step -> one results.jsonl record; no set -e. - render_report.py (stdlib): per-step JSON emitter + report.md assembler with a Correctness/Efficiency/Ergonomics rollup and failures-first log tails. - catalog.yaml: committed (validated via Catalog.load/resolve_names). - tests/: non-serving tiers (env, dry-run, ergonomics, negatives) + GPU serving tiers (single-vllm, coalescing, dedicated/F5, TTL, run-wrapper, ollama, concurrency) + always-run cleanup. - run.sh: ./run.sh [--gpu] [--only ...]; fresh data dir per run inside the results dir; prints the rsync line. Non-serving tiers smoke-tested locally against the null backend (16/16). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Test scripts are run via `bash "$f"` (a subprocess), not sourced, so a
top-level `return 0` after the skip block errored ("can only return from a
function or sourced script") and execution fell THROUGH into the real
compose acquire on a non-GPU box. Use `exit 0` so the script ends cleanly
and run.sh continues to the next tier. Verified with a full non-GPU run
(18 passed, 19 skipped, 0 failed).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…lkthrough F5 was the blocker for using yardrat's display-attached GPU 1. The placer's skip_display had no CLI knob, so only GPU 0 was placeable. Add a _DisplayGpuMixin (--include-display-gpus) on the leasing common flags, wired to ComposeBackend(skip_display=...). Default still skips display GPUs. - e2e tier 45_both_gpus: acquire two distinct models with the flag, assert the backend sidecar placed them on two distinct GPUs and both route via the gateway. - dev/leasing-demo.md: a demo-minded companion to leasing-test-plan.md — real user-config + storage setup, deploy a 14B model as a standing service, drive it from Open WebUI, switch models. Copy-pastable, no cross-block env deps. Ends with a ranked list of ergonomic smells the walkthrough surfaced. - Mark F5 FIXED in the test plan; CHANGELOG + unit test for the flag wiring. Full suite: 213 passed (was 212). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
GPU e2e found 50_coalescing/coalesce-acquire timing out: a 2nd alias coalesced onto a live group (endpoint with public_name) was added to the rendered LiteLLM config, but the running gateway never picked it up, so its readiness probe hung to the 1200s timeout. LiteLLM reads its bind-mounted model_list only at startup. converge rewrote the file but `docker compose up -d` left the litellm container running (service spec unchanged), so live routes went stale. Stamp a content-hash label (infer-stack.config-hash) on the litellm service so converge recreates it iff the routing changes; idempotent otherwise. Unit-tested. Also tighten the e2e container-count check to count only LIVE vllm groups (a reclaim:stop group from an earlier tier lingers as state=stopped in the ledger). 214 passed (was 213). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…rsync - run.sh gains a finish() EXIT trap (idempotent) that assembles the report, writes rsync-back.sh, and prints the summary + rsync line; INT/TERM exit so finish runs on interrupt. So an interrupted GPU run still returns a reviewable report (verified normal + SIGINT-mid-run paths). - The data dir lives inside the results dir and holds multi-GB HF/kernel caches. The printed rsync line + rsync-back.sh now --exclude *-cache/, ollama/, open-webui/, postgres-*/, runtime/ (keeping leasing/: ledger, compose, litellm config, .env). - Move per-step scratch (.lastout/.asserts/.notes) out of logs/ into .scratch/ (also excluded) so the rsync'd logs/ holds only real .log files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ry later) The config-hash label is an interim correctness fix that reintroduces a brief gateway blip on model add/remove (legacy kept LiteLLM up across switches). Record the decision durably: #1 (static superset route table) is the chosen direction; #2 (DB-backed live model management via LiteLLM admin API) was tried before with issues and should be retried in the future. Includes scope notes for #1 (catalog must be threaded to render; upstream service names must become deterministic) so it can be picked up cold. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ign notes - run.sh prints the (cache-excluding) rsync line at the START and writes rsync-back.sh before any test, so partial results can be synced mid-run. - Reset the compose stack between GPU tiers. The 80/85/90 failures were test isolation, not a product bug: a keep-warm qwen-15b group from tier 45 survived release and held GPU 0 (the only usable GPU on yardrat), so later acquires couldn't place and timed out. A between-tier 'compose down' isolates tiers (weights stay cached → reload, not re-download). - leasing-followups.md: group-id determinism recommendation (prefer compat-key hash; network-alias alternative; per-user NOT recommended — fights coalescing), the /health vs /v1/models split to indicate undeployed models for the #1 superset, and the keep-warm-vs-live-demand scheduling question. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Keep-warm vs demand resolved: never reclaim a LIVE group (a held lease protects it regardless of traffic); an IDLE group (demand 0 — all leases released or TTL-expired/unrenewed) is reclaimable, so a new live-demand lease may evict it to free the GPU. keep-warm = 'stay resident if there's room, yield to real demand'. Group-id determinism: approved the compat-key hash approach (option 1), to be done after the GPU e2e dashboards are green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The wall-clock was dominated by vLLM startup (torch.compile + CUDA graph capture, ~model-size-independent) and a couple of deliberate waits, not weights. - Every vLLM endpoint now runs --enforce-eager (skip compile/graph capture) with max_model_len 2048 (smaller KV alloc / capture). Biggest lever. - Swap the test models to the smallest sane ungated chat models: SmolLM2-135M (was Qwen2.5-0.5B) and SmolLM2-360M for the 2nd distinct model; ollama uses smollm2:135m. Endpoint names kept (qwen-*) so no test-script churn; coalescing + distinctness verified. - Trim intentional waits: tier 60 dedicated acquire --timeout 180->45 (it's an expected place-failure we just confirm), tier 70 TTL 60s/sleep 75 -> 30s/40s, tier 85 ollama --timeout 900->300 (success returns early; this is the failure ceiling). Success-case acquire timeouts (40/50) left high — a timeout is a ceiling, not a wait; a ready model returns immediately. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ComposeBackend.converge always ran 'docker compose up -d', but with zero services (last reclaim:stop lease released) 'up' errors 'no service selected', so the release reconcile crashed and 'infer-stack run' returned non-zero even though the job succeeded. Converge now 'down's the project when there are no services. Latent until the between-tier reset isolated deployments so a release could actually converge to empty; the GPU e2e 80_run_wrapper tier caught it (only tier that surfaces the release exit code). Unit-tested converge([]) -> down, not up. 215 passed (was 214). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
GPU tiers 80/85/90 hung because the shared ledger accumulated state across tiers and the between-tier reset only downed containers. A tier-60 acquire left its lease ACTIVE (a not-ready/uncleaned acquire keeps its lease), so its qwen-small group stayed LIVE in the ledger; every later reconcile re-realized it on GPU 0 and starved ollama / run-exit-code / the concurrency tier. - reset_between_tiers now also wipes the ledger (ledger.db + wal/shm) and downs by project name, so each GPU tier starts from zero leases/groups. - run() prints a '… still running (Ns)' heartbeat every 20s (and reaps the heartbeat's sleep child) so long cold-starts don't look hung. - New cleanup.sh: down the 'infer-stack' project (+ stragglers by label) and optionally wipe a run's ledger — for recovering after a Ctrl-C. - leasing-followups.md: record the underlying product issues (not-ready acquire leaks an active lease; re-acquire spawns a new group instead of reviving an idle one; ledger-vs-reality drift re-realizes leaked LIVE groups). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a 'deployment' column to the leases pane (the group id(s) a lease holds, matching the groups pane id) and 'leases'/'held by' columns to the groups pane in place of the opaque 'demand', so the many-to-one link is visible. Cursor movement narrates the link in the status bar. Make action-bar buttons 1 row tall. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The ledger object is a Lease; the session_id/lease dual vocabulary was a source of confusion. Hard-rename throughout: env-file key INFER_STACK_SESSION_ID -> INFER_STACK_LEASE_ID, release/renew --session -> --lease, JSON key session_id -> lease_id, read_session_id -> read_lease_id, generated id prefix sess- -> lease-. No back-compat alias (per decision). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Deployment group was an unintuitive name; the object is a Deployment and the mental model is 'many leases -> one deployment'. Hard rename across core, CLI, TUI, docs: DeploymentGroup->Deployment, GroupState->DeploymentState, group_id(s)->deployment_id(s), ledger methods, leases --json key groups->deployments, and the SQLite groups table->deployments (claims.group_id->deployment_id). No back-compat / no migration: delete any existing ledger DB and it is recreated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Remove the superseded 'legacy' profile/active-profile surface (no back-compat, pre-release): the legacy command group + commands_profile/commands_smoke, renderer/benchmark/verification/contracts, the active-profile runtime verbs (Up/Down/Purge/Deploy/Env/Ollama*) and the _ensure_rendered/RenderCLI hook, plus their tests (~5k LOC). Keep the ollama + kubeai concepts (kubeai_ops, backends/kubeai_renderer, catalog engine:ollama) for when those backends land. Rewrite commands_runtime.py: 'stack' day-2 wrappers now target only the leasing compose project, and 'infer-stack status' becomes a holistic leasing-native overview (locations, backend, leases/deployments summary, dig-deeper hints). resolver/validator/old catalog.py + legacy context plan helpers are deferred to a follow-up (they're load-bearing for context.py's kept helpers). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Phase 2a of the legacy sweep. The pre-leasing resolver + validator (and the cli/compose, cli/probes shims that only the deleted commands used) are gone; cli/context is carved down to the two helpers the leasing surface needs (_apply_path_overrides, effective_inventory). config paths now reports only current locations (config_root, settings.yaml, catalog.yaml, data_root, the leasing ledger/compose files) instead of the old profile artifacts. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
backends/compose_renderer is replaced by leasing/compose; remove it and slim backends/__init__ to the kubeai renderer, which is retained as the kubeai conceptual path. Note the old top-level catalog.py + dead config.py helpers as an internal-only follow-up (no CLI surface). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The add-endpoint wizard now exposes tensor-parallel size, max model len, GPU memory fraction, raw extra vLLM args, and reclaim policy (mirrors catalog endpoint add). Add an Edit action (blocked while an endpoint is actively served) and Remove actions for endpoints and models, with a confirm dialog and the validating writer rejecting a model still referenced by an endpoint. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Wrap the multipane monitor in a top-level TabbedContent: a Dashboard tab (the existing panes) and a new Settings tab that reads/writes the durable settings (backend, data dir, Open WebUI, reverse proxy, skip-display-GPUs) to settings.yaml. The dashboard body is composed via a helper so its layout is unchanged. Textual's command palette (ctrl+p) covers action search. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a Control tab to the docker pane showing the rendered compose-file path and Compose up / Compose down buttons that drive the leasing project through the captured docker runner (progress shows in the Logs tab). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Models table now shows the quantization (from the catalog) and a 'cached' flag derived from a fast existence check against the HF hub cache dir (state.hf_cache/hub/models--Org--Model) — no slow du. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add the two splitters Jon called out: #csplit (catalog endpoints|models) and #lsplit (leases|deployments), reusing the fixed-height-neighbor + 1fr-neighbor pattern. All four panes pairs are now drag-resizable. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
From live testing: - Endpoint wizard now labels every field and adapts to engine: vLLM exposes tensor-parallel, data-parallel, max-model-len, gpu-mem, max-seqs, prefix-caching, extra-args; Ollama exposes host + free-form KEY=VALUE runtime. (CLI parity holds: these map to catalog runtime keys / --runtime / --extra-args.) - Move the API tester to its own top-level tab (more room for the monitor). - Localize Add/Edit/Remove to the endpoints + models panels; drop the bottom button stack; move Suggest to the Settings tab. - Vertical splitter now drags the full width range (clamped to terminal width), not just the middle; fix the endpoints|models drag direction. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Convert the leases and deployments monitor panes to Collapsibles (docker and system already were), so all four monitor panes collapse to focus. Collapse and a fixed-height drag bar can't coexist on one pane in Textual (the explicit height defeats collapse), so the leases|deployments drag bar (lsplit) is dropped in favor of collapse there; the catalog stays drag-sized (endpoints|models csplit) since you act on it constantly. Pane titles carry live counts. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Excise the superseded profile-catalog machinery from config.py (builtin_*_catalog, the file/overlay/merge chain, normalized_catalogs, initial_config) — all dead externally; the leasing path only uses normalized_output/normalized_state/ default_state_paths + the port/image constants. Delete the now-orphaned infer_stack/catalog.py (leasing/catalog.py is the real one — removes the 'two catalogs' confusion) and the profile-era templates (default-*.yaml + the legacy .j2 renderer templates); keep suggestion-pool.yaml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The 4 endpoint buttons used width:1fr but Button's default min-width (16) kept them from shrinking, so Edit/Remove overflowed off the sidebar and only appeared after widening it. Add min-width:0 to the catalog action buttons (and drop the double-width glyphs) so Serve/Add/Edit/Remove all fit at the default width. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The API tab now shows the gateway + Open WebUI URLs (ctrl+click to open), adds a 'List models' check (GET /v1/models on the gateway) alongside Send / Test-all, shows a live curl preview with a 'Copy curl' button, and an 'Open WebUI' button. Clipboard via Textual OSC 52: Copy curl, plus 'y' to copy the status line. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Textual's content markup rejects the ':' in [link=http://...], crashing get_content_height when the gateway/Open WebUI URLs were shown. Render the URLs as plain text; ctrl+click the line and the Open WebUI button still open it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
vLLM OOM'd at startup (KV cache too small) because suggest overrode a model's hand-tuned gpu_memory_utilization with a bare footprint ratio (min_vram*1.3/host_mem), which on a big GPU lands *below* the default the pool author sized for the context's KV cache (e.g. smollm2-1.7b: 0.22 vs 0.4 → only ~1.43 GiB KV, under the ~1.5 GiB an 8192 context needs). The footprint estimate must only raise the reservation, never lower it: util = max(footprint, default), still clamped [0.2, 0.92]. Regression test added. Existing catalogs keep the old value — re-run catalog suggest --apply --force or edit the endpoint to pick up the fix. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
System pane: the polling gate trusted Collapsible.Toggled, which doesn't fire on every toggle path, so expanding System never flipped the gate and the GPU table stayed empty. Read the live collapse + active-tab state on the UI thread (_sync_pane_state) before each refresh instead; the gate now self-heals every interval. Regression test added. Clipboard: many terminals don't honor the OSC 52 escape Textual uses (and it blocks native select+copy), so Copy curl / 'y' silently failed. Prefer OS clipboard tools (wl-copy / xclip / xsel / pbcopy) and fall back to OSC 52; report success/failure in the status line. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…odels Open WebUI was hard-wired to the LiteLLM front door with ENABLE_OLLAMA_API=False, so the UI could only see models declared as catalog endpoints — you could not pull/manage models from it. That is the opposite of "the ollama daemon manages its own models." Rework the front door so it matches a hand-run ollama + Open WebUI stack: * Open WebUI now holds two independent connections, wired from the rendered services: an OpenAI connection (the LiteLLM gateway when on, else a single upstream's own /v1, else disabled) for chat, and a native Ollama connection pointed straight at any Ollama daemon for pull/run/delete + on-demand load. * The LiteLLM gateway is optional: --litellm/--no-litellm flag, `config set litellm`, and _resolve_litellm in the CLI. Open WebUI renders without a gateway whenever there is an upstream to point at (and is a standing front door only when the gateway is on). * access() reports the UI URL when the gateway is off but the UI is on. * probe_ready pulls an Ollama tag regardless of the gateway, so a lean --no-litellm stack still makes its anchor model present. GPU pinning for the daemon already works via runtime_hosts placement (gpu_indices -> CUDA_VISIBLE_DEVICES). Adds dev/ollama-openwebui-tutorial.md (CLI-driven, 2x1080ti worked example) and compose tests for the three wirings (direct ollama, ollama+litellm, ui-without-litellm) + the no-litellm pull. 243 tests pass; ruff clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Create docs/source/manual/ and wire it into the Sphinx toctree (User Manual section). Move the two current, leasing-era user guides there: * ollama-openwebui-tutorial.md (new) — self-managing Ollama + Open WebUI * leasing-demo.md (from dev/) — the happy-path walkthrough Update the live references (README.md, dev/leasing-followups.md) and the CHANGELOG path. The remaining docs/ and docs/demos/ guides are profile-era (pre-leasing) and were intentionally left out of the manual — they document a model that no longer exists and need a rewrite/removal pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…n yardrat Add real-GPU coverage for the new leasing features and fix a GPU-pin bug they surfaced. Leverages yardrat's 2nd GPU (the display-attached GPU 1). Fix: an Ollama daemon pinned to a non-zero GPU silently ran on CPU. The docker device reservation (device_ids) exposes only the pinned GPU and the NVIDIA runtime renumbers it to 0 inside the container, but _ollama_service also set CUDA_VISIBLE_DEVICES to the *host* index, so pinning to GPU 1 pointed at a device that didn't exist in-container. Pin by the reservation alone, like vLLM. Updated the unit test that had enshrined the bug. e2e harness (dev/e2e_tests): * 86_ollama_lean (new) — --no-litellm: no gateway, Open WebUI wired straight at the daemon's native API, tag still pulled, chat hits the daemon directly, status reports the UI URL. * 88_gpu_pinning (new) — pin a daemon to GPU 1 (--include-display-gpus): assert device_ids=[1], no host-index CUDA_VISIBLE_DEVICES, and it runs ON the GPU (ollama ps != 100% CPU). Catches the bug above on real hardware. * 85_ollama — now asserts Open WebUI gets BOTH connections (gateway OpenAI + native Ollama) when the gateway is on. * catalog.yaml — add an ollama host pinned to GPU 1 + smol-ollama-gpu1 endpoint. * inspect.py (new) — compose-file inspector so the tiers assert wiring without brittle YAML grepping. README tier table updated. 243 unit tests pass; ruff clean; catalog validates. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ew tiers Two harness bugs the first yardrat run surfaced: * `--include-display-gpus` no longer exists — placement uses every GPU by default now (opt out with `--skip-display-gpus`). The flag made acquire fast-fail with "unrecognized arguments", so 88_gpu_pinning never rendered and inspect.py read a stale compose (wrong device_ids). Dropped the flag from 88_gpu_pinning and from the pre-existing 45_both_gpus (broken the same way); GPU 1 is placeable without it. Updated the README GPU note. * 86_ollama_lean asserted the Open WebUI URL appears in `infer-stack status`, but status doesn't print it (acquire/serve do). Assert the UI front-door port from the rendered compose instead (inspect.py now prints open-webui host_port) and just sanity-check `status` exits 0. Verified inspect.py output matches every assertion against rendered compose. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.