From 2cc1a3fc7878661bb9b54cbcc11eb7a942101cbe Mon Sep 17 00:00:00 2001 From: JadenFK Date: Tue, 15 Sep 2026 07:07:24 -0400 Subject: [PATCH 1/2] ndif skills: what the second fresh-user round found on a shared 8-GPU box Both routes passed with Llama-3.1-8B and gemma-3-27b-it over two cards. What the runs surfaced: nothing said how to pin GPUs or that the ledger ignores other tenants (a new "on a GPU box you share" section, with the arithmetic for --gpus N in the operate skill and what --gpus means on the default actor); the Ray temp-dir disk rule was a config-table footnote, now a prerequisite; the 4 MiB result threshold is on compressed bytes; any comprehension or .append(), not just a list one, loses the value; multimodal checkpoints keep the decoder under .language_model; export loses gpus on 0.1.1; the tag table pinned a release number; and one paragraph still said doctor's MinIO hint pointed nowhere. Co-Authored-By: Claude Fable 5.1 --- plugins/ndif/skills/ndif-operate/SKILL.md | 19 ++++++- plugins/ndif/skills/ndif-selfhost/SKILL.md | 63 ++++++++++++++++------ 2 files changed, 64 insertions(+), 18 deletions(-) diff --git a/plugins/ndif/skills/ndif-operate/SKILL.md b/plugins/ndif/skills/ndif-operate/SKILL.md index 4139a3b..38b53e3 100644 --- a/plugins/ndif/skills/ndif-operate/SKILL.md +++ b/plugins/ndif/skills/ndif-operate/SKILL.md @@ -120,7 +120,10 @@ silently dropped. `--sync` evicts every HOT model key not in the file, trims replicas above the requested count, and deploys only the shortfall. Pair it with `ndif export` for a snapshot/restore loop — but note `padding_factor` does **not** round-trip: it is a deploy-time sizing input, never stored on the deployment, so -a restored model falls back to `NDIF_DEFAULT_PADDING_FACTOR`. +a restored model falls back to `NDIF_DEFAULT_PADDING_FACTOR`. `gpus` round-trips +from 0.1.2; on 0.1.1 an export drops it, and restoring a two-card model from that +file lets the placer put it on one card that cannot hold it — check the YAML +before `--sync`. ## Sizing: how the controller decides a model fits @@ -134,6 +137,10 @@ disagreed. padded = ceil(base + base * padding_factor + padding_bias) ``` +`ndif status`'s "GPU Memory ... free" line is this ledger (0.1.2 labels it +"unreserved by NDIF"); COLD on 0.1.1 also lists datasets and adapters found in +the HF cache, none of them deployable. + `base` is parameters + buffers at the target dtype. Defaults: `padding_factor` 0.15, `padding_bias` 500 MiB. That padding is the **entire** budget for activations, KV cache, CUDA workspaces and the CUDA context — and the actor @@ -143,7 +150,15 @@ allocator cap at run time. This is why a block can die with Placement charges each card the replica's **share**, `ceil(size / gpus_needed)`, not the whole card — so a model 1% over one card's capacity takes two cards at -about half each and the rest stays usable. +about half each and the rest stays usable. On the default actor, `--gpus N` +loads the model across N cards with an accelerate device map; it is not tensor +parallelism and needs no `NDIF_TP_MODEL_ACTOR_CLASS`. + +Worked example, a 27B bf16 model on cards with 56 GB actually free of 80 GB: +base = 27.4e9 × 2 B = 54.8 GB, padded = 54.8 × 1.15 + 0.5 = 63.5 GB. The ledger +sees 80 GB free per card and would place it on one; the card cannot hold it. +`ndif deploy google/gemma-3-27b-it --gpus 2` charges 31.8 GB to each card and +loads 26.8 + 26.4 GB. An 8B model (16.06 GB base → 18.99 GB padded) fits one. What the ledger cannot see, in order of how often it bites: **anything NDIF did not place** (a stray training job, an actor left behind by a killed controller), diff --git a/plugins/ndif/skills/ndif-selfhost/SKILL.md b/plugins/ndif/skills/ndif-selfhost/SKILL.md index 2b0dd79..ecfc61a 100644 --- a/plugins/ndif/skills/ndif-selfhost/SKILL.md +++ b/plugins/ndif/skills/ndif-selfhost/SKILL.md @@ -48,16 +48,19 @@ single-host default. Full per-route detail: `docs/operating/quickstart.md`. | An NVIDIA GPU and a CUDA driver | The controller only manages Ray nodes that report a `GPU` resource. A CPU-only node joins Ray and is then ignored; `/ping` answers but `/request` cannot be served. | | NVIDIA Container Toolkit (routes 1 and 2) | `docker run --gpus all` must work, or the container fails to create. | | `--shm-size 4g` (routes 1 and 2) | Ray's plasma object store lives in `/dev/shm`; Docker's 64 MB default makes Ray spill to disk or fail, and the error never mentions shm. | +| Room on the filesystem under `NDIF_RAY_TEMP_DIR` (default `/tmp/ray`) | Ray's raylet stops scheduling once that filesystem passes 95 % full, and the only symptom is a server that comes up and never runs anything. On a full `/`, point it elsewhere (route 1: `-e NDIF_RAY_TEMP_DIR=/ndifray -v /big/disk/ray:/ndifray`). `ndif doctor` checks this from 0.1.2. | | Disk for weights, plus host RAM | Checkpoints download at deploy time (gpt2 ~0.5 GB, a 70B ~140 GB). Evicted models are held in host RAM as WARM. | | Python 3.12 or 3.13 (route 3) | `requires-python = ">=3.12,<3.14"`; `ndif doctor` fails below 3.12. | -**Match the tag to your driver.** `ndif/ndif:0.1.0-cu126` (= `0.1.0` = `latest`) -is the default line and runs on any CUDA 12.x driver >= 525; -`ndif/ndif:0.1.0-cu130` is for CUDA 13 drivers (580+) and Blackwell (RTX 50xx, B200). There is no cu128 tag — PyTorch's cu130 index stopped at torch 2.11. A wheel built for a +**Match the tag to your driver.** Every release is published as +`ndif/ndif:-cu126` and `-cu130`; the bare `` and +`latest` point at the newest cu126. cu126 runs on any CUDA 12.x driver >= 525; +cu130 is for CUDA 13 drivers (580+) and Blackwell (RTX 50xx, B200). There is no +cu128 tag — PyTorch's cu128 index stopped at torch 2.11. A wheel built for a CUDA line your driver predates does not fail loudly — `torch.cuda.is_available()` just returns `False` and Ray starts with `cuda_memory_bytes: 0`. -`0.1.0` carries nnsight 0.8.0rc1, torch 2.14, transformers 5.17, ray 2.55.1, +The 0.1 line carries nnsight 0.8.0rc1, torch 2.14, transformers 5.17, ray 2.55, Python 3.12. Ask an image directly, with no GPU or volume needed: ```bash @@ -65,6 +68,25 @@ docker run --rm ndif/ndif version docker run --rm --gpus all ndif/ndif doctor ``` +## On a GPU box you share + +Every example here assumes the machine is yours. On a shared host three things +change: + +- **Pin the cards.** `docker run --gpus '"device=0,1"'` (that exact quoting) for + routes 1 and 2, `CUDA_VISIBLE_DEVICES=0,1` for route 3. `ndif doctor` should + then report torch seeing exactly that many GPUs. +- **The controller sizes from each card's *total* memory and never reads what + other people hold.** `ndif status` will say a 80 GB card is 80 GB free while a + colleague's job holds 25 GB of it, and the placer will put a model there. + Subtract other tenants' `nvidia-smi` usage yourself and force the placement + with `ndif deploy --gpus N` (or `--size-bytes`); see the `ndif-operate` skill's + sizing section for the arithmetic. +- **The container runs as root**, so anything it writes into a bind mount — the + Ray temp dir, new files in the HF cache — is root-owned afterwards. Clean up + through a container (`docker run --rm -v /path:/v alpine rm -rf /v/...`) or + keep those mounts on directories you don't need to delete as yourself. + ## Route 1 — the published image ```bash @@ -166,9 +188,8 @@ Two traps on this route that nothing else warns about: **The MinIO binary is the awkward part.** `ndif doctor` checks for `redis-server` and `minio` on `PATH`, and MinIO publishes no standalone server binaries any more -(`dl.min.io` returns 410, the GitHub releases carry no assets), so doctor's -"install the MinIO server binary" hint has nothing to point at. Two options that -work: +(`dl.min.io` returns 410, the GitHub releases carry no assets). Doctor's hint +names the conda-forge package; the two options that work: ```bash conda install --override-channels -c conda-forge redis-server minio-server # verified; --override-channels skips the anaconda ToS prompt a stock miniconda raises @@ -259,11 +280,14 @@ and exhaustively in `docs/reference/env-vars.md` and `docs/reference/ports.md`. ## Gotchas -- **Results over 4 MiB come back as a presigned MinIO URL.** Under - `NDIF_MAX_SOCKET_RESULT_BYTES` they ride on the response itself and port 9000 - is never touched; above it — and for *every* non-blocking request — the client - fetches the URL directly, so it needs to reach 9000 and the signature has to - name a host it can resolve. +- **Results over 4 MiB *after compression* come back as a presigned MinIO + URL.** Under `NDIF_MAX_SOCKET_RESULT_BYTES` they ride on the response itself + and port 9000 is never touched; above it — and for *every* non-blocking + request — the client fetches the URL directly, so it needs to reach 9000 and + the signature has to name a host it can resolve. The threshold is on the + serialized, compressed payload, not on the tensor bytes you count: 4.7 MiB of + bf16 activations can still ride the socket. The client prints a + `Downloading result` bar when MinIO was used. - **A block that allocates a lot of GPU memory dies with `CUDA out of memory ... N MiB allowed` even on an almost-empty card.** The actor caps per-process GPU memory at the model's size × `NDIF_DEFAULT_PADDING_FACTOR` (plus @@ -290,10 +314,17 @@ and exhaustively in `docs/reference/env-vars.md` and `docs/reference/ports.md`. - **A prompt longer than the model's context dies as `CUDA error: device-side assert triggered` in `masking_utils`,** not as an index error naming the limit (gpt2: 1024 positions). The actor survives it; the next request runs. -- **A list built by comprehension or `.append()` inside a trace block is not - bound after the block** — only plain `name = value.save()` assignments come - back. That is nnsight behaviour, not a server fault; send users to the nnsight - plugin's `debugging` skill. +- **Anything but a plain `name = value.save()` at block scope is not bound after + the block** — a list, dict or set comprehension, `.append()` into a list, a + value tucked into a container. The request reports `COMPLETED` with nothing + downloaded and the client then hits `NameError`. Write one assignment per + saved value. That is nnsight behaviour, not a server fault; send users to the + nnsight plugin's `debugging` skill. +- **Multimodal checkpoints reshape the module tree.** A `*ForConditionalGeneration` + model such as `google/gemma-3-27b-it` keeps its decoder under + `model.model.language_model.layers[i]`, not `model.model.layers[i]`; print the + model once before naming a layer. Decoder layers on transformers 5.x return a + plain tensor, so `.output[0]` selects batch element 0, not a tuple slot. ## References From 3ef1450db2564087d80b383dd73296c0c3a69e37 Mon Sep 17 00:00:00 2001 From: JadenFK Date: Tue, 15 Sep 2026 08:05:17 -0400 Subject: [PATCH 2/2] ndif-selfhost: the save-a-container rule, stated correctly Tested on gpt2: .save() on the assigned object (the comprehension, an empty list, a dict) or an outer list of saved values all bind; only saving the elements of a container built inside the block does not. Co-Authored-By: Claude Fable 5.1 --- plugins/ndif/skills/ndif-selfhost/SKILL.md | 15 +++++++++------ 1 file changed, 9 insertions(+), 6 deletions(-) diff --git a/plugins/ndif/skills/ndif-selfhost/SKILL.md b/plugins/ndif/skills/ndif-selfhost/SKILL.md index ecfc61a..1cbe6fb 100644 --- a/plugins/ndif/skills/ndif-selfhost/SKILL.md +++ b/plugins/ndif/skills/ndif-selfhost/SKILL.md @@ -314,12 +314,15 @@ and exhaustively in `docs/reference/env-vars.md` and `docs/reference/ports.md`. - **A prompt longer than the model's context dies as `CUDA error: device-side assert triggered` in `masking_utils`,** not as an index error naming the limit (gpt2: 1024 positions). The actor survives it; the next request runs. -- **Anything but a plain `name = value.save()` at block scope is not bound after - the block** — a list, dict or set comprehension, `.append()` into a list, a - value tucked into a container. The request reports `COMPLETED` with nothing - downloaded and the client then hits `NameError`. Write one assignment per - saved value. That is nnsight behaviour, not a server fault; send users to the - nnsight plugin's `debugging` skill. +- **`.save()` goes on the object you assign, not on what you put inside it.** + `acts = [h[i].output.save() for i in ...]` and `acts = []` + `.append(...save())` + leave `acts` unbound after the block — the request reports `COMPLETED`, nothing + downloads, and the client hits `NameError`. Any of these work: save the + container (`acts = [h[i].output for i in ...].save()`, `{i: ... }.save()`), + save an empty one and append to it (`acts = list().save()` then + `acts.append(h[i].output)`), or create the list *before* the block and append + `.save()`d values into it. That is nnsight behaviour, not a server fault; the + nnsight plugin's `debugging` skill covers it. - **Multimodal checkpoints reshape the module tree.** A `*ForConditionalGeneration` model such as `google/gemma-3-27b-it` keeps its decoder under `model.model.language_model.layers[i]`, not `model.model.layers[i]`; print the