feat(execution): sample local VRAM peak via nvidia-smi during guarded execution - #265
Merged
Conversation
… execution Fills the 1A gap from the v0.14.7 optimization spec audit: PeakVRAMMB was previously only populated from *remote* execution responses; local execution recorded zero VRAM even on GPU hosts. Adds a background sampler (internal/execution/vram.go) that polls nvidia-smi --query-gpu=memory.used at a 500ms interval while the local workload runs, sums across all GPUs, and returns the max total observed. runLocal starts the sampler right before runWithReservationHeartbeat and stops it immediately after, setting resp.PeakVRAMMB on both success and failure paths. Degrades cleanly to 0 on hosts without nvidia-smi (Apple Silicon, Intel-only, or PATH without the binary) — the sampler's query returns (0, false) and the goroutine stops without error. VRAM is best-effort telemetry, not load-bearing. Verified: - Unit tests pass on the no-GPU path (PATH=/nonexistent). - Live smoke test on cranium (2x NVIDIA GPUs) captured 17896 MiB peak over a 500ms window. - Full execution + mcp suites pass; go build ./... clean.
Contributor
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
toasterbook88
added a commit
that referenced
this pull request
Jul 27, 2026
…266) ## Summary Fills the **1B gap** from the v0.14.7 optimization spec audit: the spec described a \`PairwiseLinkMatrix\` / \`LinkMetric\` data structure that did not exist anywhere in the codebase. This adds it for real — a first cut that measures what's actually measurable from a single vantage. ## Types (\`internal/models/types.go\`) \`\`\`go type PairwiseLinkMatrix struct { Links []LinkMetric } type LinkMetric struct { SourceNode string TargetNode string OverlayType string // lan | tailscale | thunderbolt RTTLatencyP95 time.Duration ThroughputMBps float64 // 0 = unmeasured in this cut } \`\`\` \`ClusterSnapshot\` gains an optional \`Topology *PairwiseLinkMatrix\` field (zero-value-safe; existing consumers unaffected). ## Builder (\`internal/facts/pairwise.go\`) \`BuildPairwiseLinkMatrix(ctx, localName, nodes)\`: - Probes directional RTT from the local node to each remote node's \`ResolvedDialTarget\` (fallback \`SSHTarget\`) via \`ping -c 4 -W 1\`. - Parses the **max** RTT from the ping summary line as a P95 stand-in (4 samples). - Infers overlay from address shape: \`100.*\` → \`tailscale\`, \`169.254.*\` → \`thunderbolt\`, else \`lan\`. - **Single-vantage model**: only local→remote is probed; cross-remote edges are not measured. Throughput left at 0 (unmeasured) in this first cut — RTT is the primary signal. - **Best-effort**: unreachable nodes or missing \`ping\` are omitted, not fatal. No-local-node returns an empty matrix. \`TopologySummary(m)\` helper for logging. ## Verification | Check | Result | |---|---| | \`TestParseRTTP95\` (linux summary / no-summary / malformed) | PASS | | \`TestInferOverlay\` (tailscale / thunderbolt / lan / 10.x / hostname) | PASS | | \`TestBuildPairwiseLinkMatrixNoLocalNode\` | PASS — empty matrix, no panic | | \`TestBuildPairwiseLinkMatrixSkipsLocalAndEmptyTargets\` | PASS — self + empty skipped | | \`TestTopologySummaryEmpty\` / \`NonEmpty\` | PASS | | \`TestBuildPairwiseLinkMatrixLiveOnCranium\` (loopback) | PASS — captured \`cranium→loopback-target lan 27µs\` | | Full \`internal/facts\` + \`internal/models\` suites | PASS | | \`go build ./...\` | clean | ## Scope notes - This is a **first cut**: throughput is unmeasured, and only local→remote edges are probed. Cross-remote measurement would require distributed probing (each node probing the others) — out of scope for a single-PR addition, and the spec didn't require it. - The builder is **not yet wired into the snapshot pipeline** — it's available for callers but \`ClusterSnapshot.Topology\` stays nil by default. Wiring it into \`axis status\` / daemon refresh is a deliberate follow-up so this PR stays focused on the type + builder. ## Related Follows #264 (MCP tools, 4B) and #265 (local VRAM sampling, 1A). Completes the three real gaps flagged in the spec audit. --------- Co-authored-by: AXIS Contributor <axis-dev@example.invalid>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fills the 1A gap from the v0.14.7 optimization spec audit: `PeakVRAMMB` was previously only populated from remote execution responses — local execution on GPU hosts recorded zero VRAM. This adds real local VRAM telemetry.
Implementation
New `internal/execution/vram.go`:
Wired into `runLocal` (`internal/execution/guarded.go`):
Degradation
On hosts without nvidia-smi (Apple Silicon, Intel-only, or PATH without the binary), `queryTotalVRAMUsed` returns `(0, false)` and the sampler's peak stays 0. No error surfaced — VRAM is best-effort telemetry, not load-bearing.
Verification
Related
Follows #264 (MCP tools, 4B). The remaining audit gap (1B pairwise link matrix) will follow as a separate PR.