Skip to content
MiaAI-LabPublic

About

sparkDash ⚡ — Multi-DGX Spark Monitoring Dashboard

Resources

Stars

658 stars

Watchers

7 watching

Forks

Latest commit

 

History

389 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

sparkDash ⚡ — Multi-unit monitoring dashboard for NVIDIA DGX Spark

Platform: ARM64 React 19 Express 5 Apache License 2.0
by Mia'a AI Lab

Follow Mia on X

macOS (darwin) SSH collectors authored by C. Michael Gibbs and Nyx Voss (Onyx AI Labs) — commits. sparkDash by Mia's AI Lab.

sparkDash is a real-time web dashboard for one or more NVIDIA DGX Spark (GB10) machines in a single browser window. It streams GPU, CPU, unified memory, storage, network, and local LLM metrics — and lets you add, edit, reorder, or remove Sparks from the UI without restarts or code changes.

It also supports non-Spark units: any Linux machine with an NVIDIA GPU (e.g. a workstation with a dedicated RTX/L-series card) can be added as a dedicated GPU host and monitored the same way via SSH and nvidia-smi. For these units the dashboard correctly separates RAM (system memory) from VRAM (discrete GPU memory).

Monitored units run Linux, macOS, or Windows (a Windows PC with an NVIDIA GPU — or an AMD-iGPU Windows host like a Strix Halo, whose GPU usage comes from Windows' GPU engine counters; see Windows PCs).

sparkDash Overview page with multiple DGX Spark units, GPU metrics, and LLM status

LLM Prompt Showcase

LLM Prompt Showcase — multi-terminal streaming demo (click for MP4)

Watch the MP4


Table of contents


Latest version changelog

Version 2.40.0

  • Hide any Overview panel: Settings has switches for Recent activity and for each of the four tiles (Fleet decode, Fleet power, Memory in use, Tokens today), next to Fleet energy and LLM token totals.
  • LM Studio and Ollama are detected by their own endpoints, so they are no longer labelled SGLang or vLLM; LLM panels match their port, and the Add wizard and LLM panels have LM Studio / Ollama / TensorFold port presets.
  • Windows PCs with an AMD iGPU (Strix Halo) show GPU use and memory from Windows' GPU counters. Windows setup docs now cover the administrators_authorized_keys gotcha.

Version 2.32.0

  • LLM panel: an engine that only SSH can reach (a remote unit behind a NAT that forwards just SSH) is now monitored through an SSH forward, with a via SSH chip; TensorFold prefill tok/s counts only computed tokens, so prefix-cache hits no longer read as 100k+.
  • Decode benchmark: live generation tok/s is measured from the run's own streams, so it no longer shows 0.
  • Health: a finding is held for 90 s, so a reading sitting on a threshold does not flap and spam Activity.
  • iPhone app: the header clears the status bar and the home-screen icon is the logo. Dropping a Spark after dragging it in the sidebar no longer opens it.

Version 2.31.0

  • RoCE / RDMA panel for Sparks: link state and rate, traffic, drops, PFC and QoS settings and the RDMA counters per port, with two health findings (a link went down, packets being dropped). See RoCE / RDMA monitoring.
  • Time zone setting for the Tokens, Energy and Activity charts, so the 24 h charts line up with your local time.
  • /metrics now shows when each collector last returned real data, so a stuck GPU poll can be alerted on while SSH liveness still passes.
  • Spark page: Unified memory, CPU and Models share one row; CPU and RAM sit side by side on GPU hosts; the GPU history chart fills its panel; Windows PCs got a first real-hardware pass.

Version 2.30.0

  • Windows PCs as units: a Windows PC with an NVIDIA GPU can be monitored over its built-in OpenSSH server (GPU, RAM, CPU load, disks, network, shutdown). See Windows PCs.
  • Mobile app: sparkDash installs as a PWA over HTTPS (for example Tailscale Serve).
  • Monthly energy history with a month/year view, Prometheus /metrics (opt-in), per-GPU graphs on multi-GPU hosts, a model picker in the decode benchmark, per-Spark llmHost, and Restart sparkDash from the UI after a fleet change.
  • The Settings poll interval now really sets the polling rate; the Fleet energy card explains partial coverage.
  • Fixes: "Frontend not built" after a Docker upgrade, the SSH tunnel for benchmarks, model launchers inheriting PORT=5555, and a token-protected remote bind showing a blank page. Everything since 2.0.0 is in the CHANGELOG.

Version 2.0.0 — a new sparkDash

  • New shell: sidebar, command palette (Ctrl/⌘ K), four themes (White, Light, Dark, OLED), smooth page transitions and real links everywhere (right-click, new tab).
  • Rebuilt Overview: per-Spark cards with the model row, head/worker relations, a Model/System memory bar, and a launcher to load a model from the card.
  • Benchmarks section: Decode, Prefill, Quality and the new Tool Eval Bench (trials, side-by-side compare, one-click upgrade).
  • Showcase is now an in-app page with much smoother streaming; Token totals, Fleet energy and Activity get their own pages.
  • Start and stop your own LLMs from the dashboard; live prefill tok/s during long prefills; Quality bench with GSM8K and MMLU.
  • Health findings per Spark (hot GPU, low unified memory under 3 GB, Xid/OOM errors, slow links) and Activity events for them; Clear / Reset for Activity, Fleet energy and Token totals.
  • Mobile: docked tab bar with iOS-style motion, a Stats tab (Token totals, Fleet energy, Activity) and a rebuilt Spark page.

Full history: CHANGELOG.md


Features

Area What you get
Multi-unit Any number of units; each has a tabbed detail page plus a shared Overview
Non-Spark GPU hosts Linux boxes with a dedicated NVIDIA GPU are first-class units: same nvidia-smi collectors over SSH, detected hardware summary, and separate RAM / VRAM panels. Detail page: GPU (left) + RAM → Network → Storage (right column); Overview cards show RAM and VRAM bars
Windows hosts Windows machines with an AMD iGPU (Strix Halo and friends) are first-class units too: agentless PowerShell over OpenSSH, CPU/RAM/disks/network collectors, and GPU usage straight from the engine performance counters (reads like Task Manager). The Add/Edit wizard has a Strix Halo / Windows host unit type
Live streaming WebSocket metrics with configurable poll intervals; central history store for sparklines across tab switches
Local + remote Host metrics via sysfs/proc/nvidia-smi; remotes over SSH (key or password)
LLM probe Auto-detects llama.cpp, vLLM, sglang, LM Studio, Ollama, ds4-server, EXL3, TensorFold, FreeToken, or q27; live decode/prefill tok/s from the engine's own counters where it publishes them (not LM Studio or Ollama); cached vs uncached prefill on ds4, llama.cpp, SGLang, and q27; daily peak history on the LLM card
LLM server presets One-click LM Studio :1234 / Ollama :11434 / TensorFold :8888 presets in the Add wizard's LLM-ports step, the "+ Add LLM port" panel, and each LLM panel's settings (they fill the port number — they install nothing)
ComfyUI Opt-in probe: queue/jobs, progress, cancel, Open link, inventory, overview chip
Hermes Agent Opt-in per unit: background update check (10 min), status badges, one-click or batch hermes update
Tailnet Opt-in probe: flags a unit that is healthy on the LAN but off its tailnet
Decode benchmark Multi-concurrency streaming decode tok/s; type picker (Structured / Prose / Code / JSON). Code is a different Python task per stream. Lab protocol (temp 0, thinking off); persisted last run. Remote units: LAN HTTP, or SSH tunnel to loopback. Remote button for an on-demand HTTPS/host:port target
Prefill benchmark Context-size sweep (1k–300k) of prefill tok/s and TTFT; unique prefix per size; persisted last run. Same remote targeting as decode
Quality benchmark Fixed, seeded quality suite (QA, reasoning, arithmetic, state tracking, GSM8K, MMLU, plus optional instruction following and long-context recall) at temperature 0; per-category and overall score; last 30 runs kept; paired compare with McNemar p-value. Same remote targeting as decode
Prompt Showcase Full-page multi-terminal LLM streaming demo (up to 32 prompts) with live tok/s and copy-out
LLM inference health KV cache free/pool (vLLM cache_config_info token and byte size; usage % when the pool size is absent), run/wait queue, TTFT/E2E/ITL p95, preemptions, prefix cache, MTP accept from Prometheus /metrics (vLLM and q27; q27 FIFO-queues so the Requests tile reads “N run” without a wait gauge)
Multiple LLM ports Monitor several LLM servers on different ports simultaneously — each gets its own panel with independent backend detection and metrics
GPU processes See the top GPU processes by VRAM usage directly in the GPU panel, including process name and memory allocation
Multi-GPU hosts A dedicated GPU host with several NVIDIA cards reports each one: the header names every card, the GPU panel adds a block per card (throttle chip, usage and temperature sparklines, power, VRAM bar) and the API exposes gpu.gpus[]. The headline gpu numbers stay an aggregate of all cards, so Overview cards and alerts need no change
Spark uptime System uptime displayed inline on each Spark header for at-a-glance availability
Power controls Graceful shutdown (SSH host script). Wake-on-LAN is for dedicated GPU hosts whose NIC supports it. DGX Spark does not wake from a magic packet
Spark roles Head / Worker / Standalone — worker label + head link; standalone can disable LLM monitoring; optional hide workers from Overview and tabs
Unified memory GB10 128 GB LPDDR5X pool (~273 GB/s), GPU/CPU split, bandwidth via nvidia-smi dmon. Non-Spark hosts show discrete VRAM (nvidia-smi) and system RAM separately
Health findings Small explainable rules over metrics already collected: GPU running hot (85 / 90 °C), GPU stuck in a low-power state, unified memory low (under 3 GB warns, under 1.5 GB critical), NVIDIA Xid errors and kernel OOM kills, several models loaded while memory is busy, a link slower than 1 Gb/s, and RoCE link / loss. Shown as chips on the cards and in full on the unit page; changes are logged to Activity
Fleet pages Token totals (per model, hourly / daily history), Fleet energy (rolling 24 h / 31 d plus a permanent monthly archive, optional cost), and Activity (event log with filters). Each page can be cleared or reset from its Clear / Reset… menu
RoCE / RDMA Per-port link state and rate, traffic, drops, PFC and QoS (trust mode, DSCP), and RDMA counters, read from sysfs / ethtool / mlnx_qos on the Spark (see RoCE / RDMA monitoring)
Windows PCs A Windows PC with an NVIDIA GPU is a first-class unit over the built-in OpenSSH server; an AMD-iGPU Windows host (Strix Halo) works too — its GPU usage comes from Windows' GPU engine counters (see Windows PCs)
Mobile + PWA A phone layout with a docked tab bar (Overview, Sparks, Stats, Settings); installable to the home screen over HTTPS (see Mobile PWA)
Time zone Settings → Time zone sets hour and day boundaries and labels on the Tokens, Energy and Activity charts, whatever device you view them on
Prometheus Opt-in GET /metrics export (see Prometheus)
Restart from the UI When the fleet changes and Energy needs a restart, a Restart sparkDash button restarts the Docker container (only offered when it can)
Themes Dark, light, cool white, OLED — neutral palettes, persisted in localStorage
Secrets SSH passwords AES-256-GCM encrypted; never in sparks.json or API responses
Docker-first Single privileged container for host metrics; prod and dev Compose files
Hot config Add / edit / remove / reorder Sparks from the UI with no process restart

ComfyUI monitoring

sparkDash can optionally monitor a ComfyUI instance on each Spark — the same way it probes local LLMs, but focused on jobs and queue, not a second copy of GPU/RAM bars (those stay on the GPU / CPU panels).

What is supported

Capability Details
Opt-in per Spark comfyMonitoring (default off) + comfyPort (default 8188)
Any role Head, worker, and standalone can enable ComfyUI independently of LLM cluster role
Liveness GET /system_stats — online, ComfyUI / PyTorch version, device type (cpu/cuda)
Queue / jobs GET /queue — running + pending items; workflow title, model/LoRA filenames from the graph, footprint (resolution · steps · sampler · batch · node count)
Progress Progress bar on the active job — Comfy WebSocket when events are available; otherwise elapsed / average-duration estimate
Last finished job Status + duration via /api/jobs (fallback /history)
Queue ETA Estimate from recent job durations × pending (+ progress remainder when known)
Cancel / remove From the Comfy card: interrupt a running job or dequeue a pending one (POST /api/sparks/:id/comfy/cancel)
Open ComfyUI One-click link to http://{lanIp}:{comfyPort} (LAN IP preferred so remote browsers do not hit localhost)
Model inventory Checkpoints + LoRAs from /models/* (UI section only when at least one file is listed)
Overview chip When monitoring is on: Comfy · idle / run / Nq / muted if unreachable
Layout Under Services: primary LLM + Comfy side-by-side when both are enabled; collapsible Resources / Services sections

Not claimed: true per-job VRAM (Comfy does not expose that cleanly over HTTP). Host GPU/VRAM remains on the GPU panel. Live step progress depends on Comfy broadcasting WS events; stock Comfy often scopes detailed progress to the client that submitted the prompt.

How to enable (per Spark)

  1. Open the Spark tab → Edit (pencil).
  2. Enable ComfyUI monitoring.
  3. Set port if needed (default 8188).
  4. Save.

The Spark page Services section shows the ComfyUI card. On Overview, a small Comfy chip appears for that unit.

Connectivity Test (in Edit) includes ComfyUI when monitoring is enabled.

ComfyUI side requirements

  • ComfyUI must be reachable from the sparkDash server on the probe host:
    • Local Spark (isLocal): sparkDash probes 127.0.0.1:{port} (use Docker network_mode: host if the dashboard runs in a container).
    • Remote Spark: probe uses the Spark LAN IP (same as LLM probes).
  • For Open from another machine’s browser, Comfy should listen on a reachable interface (e.g. --listen 0.0.0.0), not only loopback, and the Spark’s LAN IP must be set correctly in Edit.

Config fields (persisted on the Spark)

Field Default Description
comfyMonitoring false Probe ComfyUI and show the card / overview chip
comfyPort 8188 ComfyUI HTTP port

Related API

Method Path Purpose
POST /api/sparks/:id/comfy/cancel Cancel a job ({ "promptId": "<uuid>" }) — interrupt running and/or remove from queue

Env (optional): COMFY_PORT (default 8188), COMFY_PROBE_TIMEOUT_MS, POLL_INTERVAL_COMFY.


Hermes Agent monitoring

sparkDash can optionally monitor Hermes Agent (nousresearch/hermes-agent) on each unit and run one-click updates for you over SSH.

What is supported

Capability Details
Opt-in per Spark hermesMonitoring (default off) in Edit Spark
Auto update check Background hermes update --check over SSH (default every 10 min) — returns update availability + pending commits
Status badges In the Spark header: Hermes (installed version), Hermes not found if the binary is missing
One-click update Update Hermes button opens a dialog with live status, real pending commits, and release notes; Update now runs hermes update via SSH (non-interactive go)
Update state Running / success / error surfaced live (button turns into a “Hermes updating… / failed” state)
Batch update Update Hermes on Overview runs hermes update on every monitored unit, with a live per-unit progress bar

How to enable (per Spark)

  1. Open the Spark tab → Edit (pencil).
  2. Enable Hermes Agent.
  3. Save — background checks start immediately.

The Update Hermes button appears in the Spark header/mobile action row; it turns warning-yellow with a commit-count badge only when an update is actually available. It also appears on Overview (batch) when at least one unit has Hermes enabled.

Connectivity check note: local units run the check as the host user (via setpriv/nsenter, never as container root); remote units run it over SSH. Either way, the logged-in user needs permission to read the Hermes repo.

Side requirements

  • Hermes Agent must be installed on the target machine — sparkDash only checks & updates; it does not install it. The binary is looked up in ~/.local/bin and /usr/local/bin.
  • SSH user must be able to run hermes update --check / hermes update non-interactively (key auth recommended).
  • An update can take a few minutes (repo pull + dependency reinstall); a stale *.lock file from a crashed run is cleared before each attempt.

Config fields (persisted on the Spark)

Field Default Description
hermesMonitoring false Check/update Hermes Agent on this machine

Related API

Method Path Purpose
POST /api/sparks/hermes/update-all Batch hermes update on every monitored Spark (Overview button)
POST /api/sparks/:id/hermes/check Force hermes update --check now
POST /api/sparks/:id/hermes/update Run hermes update in the background (202)
GET /api/sparks/:id/hermes/updates Update preview: latest release + installed version + real pending commits + resolved view

Env (optional): POLL_INTERVAL_HERMES (default 600000 ms), HERMES_UPDATE_TIMEOUT_MS (default 600000 ms).


Tailnet monitoring

Opt-in per unit (default off). Runs tailscale status --json on the host and shows a Tailnet card under Resources.

This closes a blind spot every LAN-based check shares, including sparkDash's own SSH liveness. When tailscaled loses its session with the coordination server, SSH/GPU/LLM can all stay healthy while the box is unreachable from off-LAN.

What is supported

Capability Details
Opt-in per unit tailscaleMonitoring (default off) in Edit Spark
Off-tailnet detection Self.Online — the node's own view of the coordination server
Reason, not just state Tailscale Health messages, backend state, tailnet IP, DERP relay, version, expired-key warning

Asked of each node about itself. Peer state is never the verdict. The probe is read-only (tailscale up / down / login are never run).

How to enable

  1. Open Edit Spark.
  2. Tick Tailnet monitoring.
  3. Save. The Tailnet card appears under Resources.

Host requirements

  • tailscale CLI on the monitored host, and tailscaled running.
  • Remote units: existing SSH. Local Docker: nsenter into the host mount namespace (same as nvidia-smi; /host/proc is already bind-mounted).

Config fields

Field Default Description
tailscaleMonitoring false Run tailscale status --json and show the Tailnet card

Env (optional): POLL_INTERVAL_TAILSCALE (default 30000), TAILSCALE_PROBE_TIMEOUT_MS (default 8000).


Tailscale Serve (HTTPS)

Tailscale Serve publishes the loopback-only dashboard over HTTPS to your tailnet — no ports opened, no reverse proxy, and a real certificate (https://<node>.<tailnet>.ts.net). It is also the easiest way to satisfy the HTTPS requirement for the Mobile PWA, which browsers refuse to install over plain HTTP.

One-time setup

  1. Enable HTTPS certificates and Serve for the tailnet (admin does this once): open the approval link Tailscale prints on first use, or in the admin console enable HTTPS (MagicDNS → HTTPS Certificates) and Tailscale Serve.

  2. On the sparkDash host, publish the dashboard (the dashboard must be running — Docker or npm start):

    sudo tailscale serve --bg --https=443 http://127.0.0.1:5555

    If the container binds a specific address instead of loopback, proxy that address (check docker logs sparkDash for the bind= line), e.g. sudo tailscale serve --bg --https=443 http://100.x.y.z:5555.

  3. Verify and get your URL:

    tailscale serve status
    # https://your-node.your-tailnet.ts.net  ->  proxy http://127.0.0.1:5555

Notes:

  • The Tailscale hostname works with the default loopback bind — no SPARKDASH_ALLOWED_HOSTS entry needed (that list is only for custom reverse-proxy domains).
  • Access is tailnet-only. Any device signed in to the tailnet (phone, laptop) can open the URL; nobody else can. For public exposure use tailscale funnel and an authenticated front door (see Remote access) — but prefer keeping it tailnet-only.
  • The proxy forwards /api/* and /ws including WebSocket upgrades; nothing else to configure.
  • To remove it later: sudo tailscale serve --https=443 off.
  • Snap installs run Tailscale as root; if the plain command reports Access denied, use sudo tailscale serve ….

Mobile PWA (install on a phone)

sparkDash is a PWA: from a supported browser it installs to the home screen and runs fullscreen like a native app, with its own icon and no browser chrome.

What ships

File Purpose
public/manifest.webmanifest App name, standalone display, theme/background colors (#0a0c0f), icons
public/icons/ 192/512 px icons, maskable variants (Android adaptive icons), apple-touch-icon.png
public/sw.js Service worker: caches the app shell; never caches /api/* or /ws (live data always comes from the network)
src/components/shell/InstallPrompt.tsx In-app install banner (see below)

Requirements

  • HTTPS is mandatory for install (or localhost for local testing). Use Tailscale Serve or an authenticated TLS reverse proxy — a plain http://<lan-ip>:5555 page cannot be installed.
  • The manifest and icons are static files served by the same Express server (dist/ after npm run build); no extra configuration.

Installing

The app shows its own banner when opened in a browser it can be installed from (hidden once installed, or after dismissing):

  • Android / desktop Chrome: an Install banner appears near the bottom of the page; tapping it opens the native confirm dialog. Also available in the ⋮ menu → Install app / Add to Home screen.
  • iPhone / iPad (Safari): iOS allows no programmatic prompt; the app shows a one-time hint, then: Share ⬆︎ → Add to Home Screen → Add.
  • The banner only appears on a fresh page load — reload once if you don't see it.

After installing, the icon sits on the home screen and the dashboard opens fullscreen with live data whenever the device can reach the server (on the tailnet, with Tailscale connected).

Updating the installed app

Service worker caching: hashed assets/* are cache-first (immutable), index.html and / are network-first with a cache fallback. After deploying a new build, the installed app picks it up on the next open; no manual cache clearing is needed. If you ever need to force it, bump the CACHE version in public/sw.js.


Glance integration

Glance can show the cluster on a self-hosted dashboard with the community widget sparkdash-dgx-cluster (contributed by @linxichen). One custom-api card renders GPU temperature and usage, VRAM, generation tok/s, KV-cache usage, and uptime for a head node plus one worker, reading /api/sparks/:id/metrics directly:

- type: custom-api
  title: sparkDash · DGX Cluster
  cache: 30s
  url: ${SPARKDASH_HEAD_URL}
  subrequests:
    node2:
      url: ${SPARKDASH_NODE2_URL}
  options:
    dashboardUrl: ${SPARKDASH_DASHBOARD_URL}
  template: |
    # paste the template from the widget's README
Variable Value
SPARKDASH_HEAD_URL Metrics URL of one unit — https://sparkdash.example.com/api/sparks/<id>/metrics
SPARKDASH_NODE2_URL Metrics URL of a second unit. For a single Spark, drop the subrequests map and the widget's second node block
SPARKDASH_DASHBOARD_URL Dashboard base URL, used for the card's title link

/api/sparks/:id/metrics is unauthenticated on loopback, the same as the rest of the dashboard. Read Security before exposing it beyond a trusted network.


Quality bench

The Quality button on the LLM card scores whatever model the port is serving on a fixed suite, so you can compare models, quantizations and KV-cache formats over time. Every item comes from a seeded PRNG and has a stable id: two runs are the same questions, paired item by item.

Every request is an OpenAI chat completion at temperature 0 with a fixed seed. Thinking is switched with the same flags the other benches use.

Category Items Thinking max_tokens Scored by
QA 150 (30 fixed + 120 generated: multiplication, add/subtract, string reversal, letter counts, sorting, binary, date offsets, LCM) off ≥ 64 Accepted answer as a case-insensitive substring; counts, binary and LCM must be the last integer in the reply
Reasoning 40 multi-step word problems on 8192 Last Answer: <jar>, <per box> line in the final answer (not the reasoning)
Arithmetic chain 40 ten-step chains (multiply, add, subtract, remainder, floor divide) on 12288 Last Answer: <number> (markdown and thousands commas allowed)
State tracking 40 token-transfer stories, 5 people, 25 events on 12288 Last Answer: <number>
GSM8K 200 grade-school maths problems, a fixed seeded sample of the public GSM8K test set (MIT) on 8192 Last Answer: <number> ($, markdown and thousands commas allowed)
MMLU 285 multiple-choice questions, 5 per subject across all 57 MMLU subjects, a fixed seeded sample of the public test set (MIT) off 512 Last Answer: <letter> (or a bare letter)
Instruction following 40 prompts with 2–4 verifiable formatting rules (case, length, bullets, paragraphs, keywords, endings…) off 1024 Each rule is checked by code; an item passes only when every rule holds. Off by default so older overall scores stay comparable
Long-context recall Items per size (default 2) at 8k–256k (default 32k, off by default) off 512 16 animal codes spread through filler text; 4 are corrected near the end. Passes when all 16 latest codes come back; overwritten codes are counted as stale

The overall score is the mean of the category percentages. Long sizes that do not fit the model context (from the live probe or /v1/models) are skipped and listed.

The dialog shows live progress, the per-category table (mean completion tokens and how many replies hit max_tokens for the thinking categories, keys found and stale answers for long recall), and a collapsible per-item table with reply excerpts. Pick an earlier run under Compare with to see both scores, how many replies were identical, how many items only one run got right, and the exact two-sided McNemar p-value. When p ≥ 0.05 the row reads difference within noise. Copy results copies a plain-text summary.

The last 30 runs per unit are kept in config/quality-bench-history.json, with the label, model id, settings, and per-item pass/fail and reply hash. One quality run per unit at a time; it cannot overlap a decode or prefill bench or a showcase. A full default run is about 755 requests, and the thinking categories can take a while on slow models.

Quality bench by Mia's AI Lab.


Prometheus

Opt-in (default off). Settings → Prometheus metrics serves every unit's metrics at GET /metrics in the Prometheus text format (0.0.4), so Prometheus, VictoriaMetrics or Grafana Agent can scrape sparkDash and keep history beyond what the dashboard shows. While the setting is off the path answers 404.

/metrics goes through the same auth as the REST API: open on a loopback install, and on a remote bind with SPARKDASH_TOKEN set it needs the token as a bearer header.

scrape_configs:
  - job_name: sparkdash
    scrape_interval: 15s
    static_configs:
      - targets: ["sparkdash.lan:5555"]
    # Only when SPARKDASH_TOKEN is set on a non-loopback bind:
    authorization:
      type: Bearer
      credentials_file: /etc/prometheus/sparkdash.token

The exporter reads the snapshot sparkDash already holds, so scraping adds no SSH or nvidia-smi calls; values are as fresh as the poll interval. Names follow node_exporter / DCGM-exporter practice: base units (bytes, seconds, °C, watts, 0–1 ratios), _total only on counters. Every series carries unit (the unit id), unit_name, kind (spark / host) and role (head / worker / standalone). Memory and disk sizes are collected in MiB and converted to bytes.

Family Type Extra labels Notes
sparkdash_up gauge 1 when the unit answered its last liveness check. An unreachable unit exports sparkdash_up 0 and nothing else
sparkdash_uptime_seconds gauge Host uptime
sparkdash_collector_last_success_timestamp_seconds gauge collector Unix time of the last poll of each collector (gpu, cpu, ram, memory, network, storage, llm, roce) that returned real data, not zeroed defaults. A stuck collector stops advancing while sparkdash_up stays 1, so alert on time() - sparkdash_collector_last_success_timestamp_seconds > 60. Absent until a collector has succeeded
sparkdash_gpu_info gauge gpu, name, uuid Always 1
sparkdash_gpu_utilization_ratio gauge gpu 0–1
sparkdash_gpu_temperature_celsius gauge gpu
sparkdash_gpu_power_watts, sparkdash_gpu_power_limit_watts gauge gpu
sparkdash_gpu_memory_used_bytes, sparkdash_gpu_memory_total_bytes gauge gpu GB10: the GPU's share of / the whole unified pool
sparkdash_gpu_throttled gauge gpu, reason 1 while clocks are limited for thermal, power or hw
sparkdash_memory_available_bytes gauge Headroom for GPU work: unified-pool MemAvailable on a GB10, free VRAM on a discrete GPU
sparkdash_cpu_utilization_ratio gauge 0–1
sparkdash_cpu_temperature_celsius gauge sensor A GB10 has no CPU package sensor; sensor="acpitz" says it is a board zone
sparkdash_ram_used_bytes, sparkdash_ram_total_bytes gauge
sparkdash_network_receive_bytes_per_second, …_transmit_bytes_per_second gauge interface Interfaces hidden in sparkDash are skipped
sparkdash_disk_used_bytes, …_available_bytes, …_total_bytes gauge mount, device Devices hidden in sparkDash are skipped
sparkdash_llm_up gauge port, backend, model One per LLM endpoint (head and standalone units)
sparkdash_llm_generation_tokens_per_second, sparkdash_llm_prefill_tokens_per_second gauge port, backend, model
sparkdash_llm_kv_cache_usage_ratio gauge port, backend, model Where the backend reports it
sparkdash_llm_requests_running, sparkdash_llm_requests_waiting gauge port, backend, model Where the backend reports it
sparkdash_llm_generated_tokens_total, sparkdash_llm_prompt_tokens_total, sparkdash_llm_cached_prompt_tokens_total counter port, backend, model The engine's own lifetime counters — they reset when the engine restarts, which rate() handles

Unknown is left out rather than written as zero: a failed GPU read, a sensor that does not exist, or a field the LLM backend does not publish produces no sample.


RoCE / RDMA monitoring

Spark units with RDMA devices (the CX7 ports) get a RoCE / RDMA panel on their page, with no switch integration and no root. Per port: link state and rate, MTU, RX/TX traffic, interface errors and drops, PFC priorities, QoS trust mode (PCP or DSCP) and the DSCP map, global pause settings, and the RDMA counters (out of buffer, sequence errors, ACK timeouts, ECN-marked packets, CNPs sent and handled) plus ethtool discards, CRC errors and pause frames, with the change since the last sample.

Everything comes from the Spark itself: /sys/class/infiniband, /sys/class/net, ethtool -S / -a and mlnx_qos -i. The fast part (link, counters, traffic) is read about every 5 s; the slow part (ethtool counters, PFC and trust) every 30 s. Units without RDMA devices are skipped and re-checked every 10 minutes. The API exposes it as metrics.roce in /api/sparks/:id/metrics.

Two health findings use it: a RoCE link went down (a port that was up, critical) and RoCE is dropping packets (RDMA loss counters such as out_of_buffer or sequence errors rising for three samples in a row, or port discards / CRC errors rising on two consecutive 30 s readings; a single dropped frame is not reported). Ports that never came up are not reported. Switch telemetry (PFC and ECN statistics on the switch side) is not part of this; it could be added later as an optional integration.


Windows PCs

A Windows PC with an NVIDIA GPU can be monitored with no agent. On the PC:

  1. Install the OpenSSH Server (Settings → Apps → Optional features → OpenSSH Server), then start it and set it to start automatically: Start-Service sshd; Set-Service sshd -StartupType Automatic in an elevated PowerShell. Windows Firewall opens port 22 for it.
  2. Make sure the NVIDIA driver is installed (nvidia-smi works in a terminal).
  3. In sparkDash choose Add Spark / GPU host, set Unit type to Windows PC with an NVIDIA GPU, enter the PC's LAN IP and your Windows user (a password works; for key login see the administrator-account note below).

Keys work everywhere but this PC? For a Windows user in the Administrators group, sshd ignores ~\.ssh\authorized_keys and reads C:\ProgramData\ssh\administrators_authorized_keys instead, and refuses it unless only Administrators and SYSTEM can write to it. Put the public key in that file, then fix its permissions in an elevated PowerShell:

icacls.exe "$env:ProgramData\ssh\administrators_authorized_keys" /inheritance:r /grant "Administrators:F" /grant "SYSTEM:F"

A refused key on a Windows unit also shows this hint in the connection error.

sparkDash then runs two short PowerShell scripts over SSH per poll (no files are installed): GPU temperature, utilisation, power, VRAM and GPU processes from nvidia-smi, and RAM, uptime, CPU load, disks (fixed drives) and network adapters from Windows' CIM classes. Shutdown works (shutdown.exe /s). Not available on Windows: CPU temperature (Windows exposes no unprivileged sensor; CPU power is estimated from load), automatic Wake-on-LAN MAC detection (enter the MAC by hand), Hermes and Tailnet checks, model launchers, and the kernel Xid/OOM events. LLM servers on the PC (llama.cpp, Ollama, LM Studio, vLLM) are probed over HTTP as for any unit; if the server only listens on 127.0.0.1, benchmarks fall back to an SSH tunnel.

AMD-iGPU Windows hosts (Strix Halo): if nvidia-smi is not installed, the same GPU script falls back to Windows' GPU Engine / GPU Adapter Memory performance counters — utilization is the busiest engine per physical adapter (three samples per poll, clamped to 100%), and memory comes from the adapter counters, with the total read from the driver's reported memory size (on a unified-memory iGPU this is the BIOS carve-out, e.g. 96 GB on a 128 GB Strix Halo; a MemTotal/MemAvailable fallback applies when the driver does not report it). The unit type is the same Windows option. An NVIDIA PC whose nvidia-smi is installed but failing does not use this fallback; it reports the failure instead. AMD publishes no GPU temperature or power to Windows, so those read 0. The model launchers are bash/systemd and stay unsupported on Windows; HTTP probing and benchmarks work as usual.


LLM server presets & backend auto-detection

Server presets

LM Studio (:1234), Ollama (:11434) and TensorFold (:8888) are one-click presets in three places — the Add wizard's LLM-ports step, the unit page's "+ Add LLM port" panel (clicking a preset adds the port immediately), and each LLM panel's settings (the preset fills the port field, then Save). Presets fill a port number only: they do not install, start or switch anything, and the dashboard will show a port as unavailable until a server actually answers there.

How each backend is detected

Engine Signal
llama.cpp /slots with per-slot timings
vLLM Prometheus /metrics (the default for OpenAI-compatible servers)
SGLang native /server_info / /get_server_info + /model_info
LM Studio /api/v0/models — model entries carry a load state (a contract no other engine has)
Ollama native /api/tags
ds4-server ds4_* Prometheus series
EXL3 /health with backend + token totals
q27 Prometheus exposition
TensorFold owned_by on /v1/models, totals from /health
FreeToken /v1/stats throughput contract

Detection is positive-signature first: some desktop engines (LM Studio) answer HTTP 200 with a JSON body on every path, which would trip naive heuristics ("does /server_info return JSON?" → SGLang, or the vLLM default). sparkDash probes each engine's unique endpoint and never lets a catch-all-200 server downgrade an identified backend.

Decode / prefill tok/s on engines without counters

vLLM, SGLang and friends publish cumulative token counters, and the dashboard diffs them per poll. LM Studio and Ollama publish none, so their tok/s tiles stay at 0 in the LLM panel. sparkDash never sends a request of its own to a model server to measure speed (that would load models, keep them resident and add traffic you did not ask for). The decode and prefill benchmarks and the Prompt Showcase measure real speed on demand.

Multiple LLM ports

Any unit can monitor several LLM servers on different ports at once; each port gets its own panel with independent backend detection, metrics and history. Panels are matched to their port explicitly, so adding, removing or switching ports never shows one port's data on another's panel.

Quick start

git clone https://github.com/MiaAI-Lab/sparkDash.git
cd sparkDash

# Production (Docker; loopback-only by default)
docker compose up --build -d

# Or development (host, with hot reload)
npm install
npm run dev
  • Docker: open http://127.0.0.1:5555 on the host (arm64 image, auto-restart, host mounts for GPU/metrics access). The compose file serves the frontend from the host's ./dist when it has been built (npm run build) and from the copy baked into the image otherwise, so a fresh clone works; a host build that is older than the image's copy is ignored, so after git pull one docker compose up --build -d is enough
  • Dev: Vite on http://localhost:5173 (proxies API/WS to Express)

For another computer, keep the server on loopback and use an SSH tunnel:

ssh -N -L 5555:127.0.0.1:5555 user@sparkdash-host

Then open http://127.0.0.1:5555 on that computer. For shared access, use an authenticated TLS reverse proxy, Tailscale Serve, or set BIND_HOST=0.0.0.0 and SPARKDASH_TOKEN. A direct LAN bind without a token is open by default: anyone who can reach the port can change settings and power units off, and the header shows an Open access warning. Set SPARKDASH_TOKEN to require a token, or SPARKDASH_ALLOW_OPEN_REMOTE=0 to make a tokenless LAN bind refuse to start.

When the server has SPARKDASH_TOKEN set, the dashboard asks for it: the first request or live-telemetry connection the server turns away opens an Access token dialog. Enter the token once; it is checked against the server, stored in this browser only, and the live connection reconnects with it — no reload, no devtools. Settings → Access token shows whether one is stored and lets you change or clear it.

For development with Docker (source-mounted, HMR):

docker compose -f docker-compose.dev.yml up --build

Remote units + SSH keys (Docker): SSH is executed inside the container on the sparkDash host (typically the head DGX). Configured LAN IPs are from that host’s point of view, not your laptop. OpenSSH looks for keys under /root/.ssh in the container — the host user’s ~/.ssh is not used unless you bind-mount it. Uncomment this volume in docker-compose.yml (and recreate the container):

- ${HOME}/.ssh/id_ed25519:/root/.ssh/id_ed25519:ro

If the key file has a non-default name (e.g. id_ed25519_shared), mount it as id_ed25519, or set SSH_IDENTITY_FILE to the path inside the container. Keep the file mode 600. The unit that runs sparkDash itself should be added with This host (local collectors — no SSH for metrics).


Running it on a machine that is not a Spark

The dashboard does not have to run on a Spark. Every unit, including the machine running the dashboard, can be monitored over SSH, so you can run sparkDash on any Linux box (an x86 mini PC, a NAS, a VM) and add all your Sparks as remote units.

  • The shipped Dockerfile and docker-compose.yml target arm64 DGX Spark hosts (platform pin and aarch64 NVIDIA library mounts). On x86 either edit those lines or run it bare: npm install && npm run build && npm start (Node 22).
  • Add each Spark with Add Spark using its LAN IP and an SSH user. The machine running the dashboard does not need a GPU; skip adding it as a unit.
  • An x86 Docker build is untested. Reports and PRs are welcome.

Architecture

Design principle: one Spark model, N instances. Every unit is a record in config/sparks.json with a kind field (spark or host). The same SparkMonitor, SystemCollector, and LlmProbe code runs for all of them. Adding a unit is a config change, not a code change.

┌────────────────────── Docker container (sparkDash) ────────────────────────┐
│  Express (server/)                                                         │
│  ├─ config/sparks.json        Spark registry (API read/write)              │
│  ├─ SparkRegistry             load/persist Sparks; change events           │
│  ├─ SparkMonitor (per Spark)  collector + LLM probe + rate baselines       │
│  │   ├─ SystemCollector       local sysfs/proc OR remote SSH               │
│  │   └─ LlmProbe              HTTP to host:LLM_PORT, backend autodetect    │
│  ├─ REST /api/*                                                            │
│  └─ WebSocket /ws             snapshot stream to browsers                  │
│  React SPA (src/)  — Overview + per-Spark pages, themes, dialogs           │
└────────────────────────────────────────────────────────────────────────────┘
         │ SSH (key or sshpass)                    │ HTTP :8888
         ▼                                         ▼
    remote Spark(s)                         each Spark’s LLM server

Data flow

Browser  ←→  WebSocket /ws   ←→  SparkMonitor.snapshot()  ←→  collectors
Browser  ←→  REST /api/*     ←→  SparkRegistry + SparkMonitor

Poll loops run in the background (even with no clients) so rate metrics — tokens/s, network bytes/s, disk I/O — stay correct.


Tech stack

Layer Stack
Frontend React 19, TypeScript, Vite 8, Tailwind CSS v4
Backend Node.js (ESM), Express 5, ws
Platform The dashboard targets ARM64 — DGX Spark GB10 (Neoverse V2). Units it monitors can be Linux, macOS or Windows machines over SSH
Deploy Docker multi-stage (arm64), Compose
Secrets AES-256-GCM SSH password store
Ports 5555 dashboard/API; 5173 Vite (dev only)

Repository layout

sparkDash/
├── src/                 React + TypeScript SPA
│   ├── api/             REST client + shared types
│   ├── components/      Overview, Spark pages, dialogs, UI primitives
│   ├── hooks/           WebSocket snapshot, routing
│   └── theme / CSS      Tailwind v4 + four themes
├── server/              Express + WebSocket (plain JS ESM)
│   ├── sparks/          SparkRegistry, SparkMonitor
│   ├── collectors/      SystemCollector, LlmProbe, ssh, RoCE, Windows (PowerShell)
│   ├── health/          HealthEvaluator (findings)
│   ├── energy/          Fleet energy tracker + monthly archive
│   ├── llmtokens/       Token ledger · events/ Activity log · metrics/ GPU history
│   ├── llmlaunch/       Model launchers · tooleval/ Tool Eval Bench
│   ├── secretsStore.js  Encrypted password persistence
│   └── validate.js      Host/user validation (SSRF-minded)
├── config/              Runtime state (volume; secrets gitignored)
├── assets/              Logo (bolt.svg) and legacy media
├── .github/             README screenshot and Showcase video
├── Dockerfile           Production multi-stage arm64
├── docker-compose.yml   Production
├── docker-compose.dev.yml
└── deploy.sh            Rebuild / recreate helpers

REST API

Method Path Purpose
GET /api/sparks List Sparks (passwords redacted)
POST /api/sparks Add Spark and start its monitor
PATCH /api/sparks/:id Update Spark (hot-swap config)
DELETE /api/sparks/:id Remove Spark and drain monitor
PUT /api/sparks/order Persist tab order
GET /api/sparks/:id/metrics One-shot metrics snapshot
GET /api/fleet-energy Estimated fleet watts, rolling energy, coverage, and Wh/output-token
GET /api/fleet-energy/history Hourly / daily energy buckets for the Fleet energy page
GET/DELETE /api/fleet-energy/monthly Permanent per-UTC-month energy totals; DELETE needs ?month=YYYY-MM or ?all=true&confirm=delete-all-history
GET /api/llm-token-totals · /api/llm-token-totals/history Cumulative LLM tokens by model, and the hourly / daily history behind the Token totals page
GET /api/sparks/:id/gpu-history Last hours of GPU utilization, temperature and power % (windowMs, up to 8 h; parallel arrays)
GET/POST /api/restart GET says whether the server can restart itself (Docker with a restart policy); POST shuts down cleanly so Docker starts it again
DELETE /api/events · /api/fleet-energy Clear Activity or Fleet energy: everything, or only entries older than ?olderThanMs=
DELETE /api/llm-token-totals Reset Token totals (all Sparks, or one with ?sparkId=)
GET /api/health · /api/auth/status Bind / auth mode; whether a token is required and accepted
POST /api/sparks/:id/shutdown · /wake · /api/sparks/shutdown-all · /wake-all Power controls (see Power controls)
POST /api/sparks/:id/refresh/:domain Re-poll one domain now (storage or llm only)
GET/POST/DELETE /api/sparks/:id/llm/showcase[/:sessionId] Prompt Showcase sessions
GET /api/events Fleet event log (limit, sparkId, sinceId, beforeId; 2000 kept)
GET/POST/PUT/DELETE /api/sparks/:id/llm-launchers[/:lid] Registered start.sh / stop.sh model launchers; POST …/:lid/start and …/stop run them, GET …/jobs/:jobId streams the shell output
GET/POST/DELETE /api/sparks/:id/tool-eval/… Tool Eval Bench on a Spark: status, install, preview, probe, runs (start / list / stream / attach / stop / refresh / result / delete); GET /api/tool-eval/spec serves the option spec
POST /api/sparks/test Ephemeral SSH + LLM (+ Comfy if enabled) test (no persist)
POST /api/sparks/:id/test Connectivity test (can save password)
POST /api/sparks/:id/comfy/cancel Cancel ComfyUI job by promptId
PUT /api/sparks/:id/password Save SSH password (works offline)
PUT /api/sparks/:id/disabled-devices Hide storage devices (hot)
PUT /api/sparks/:id/disabled-interfaces Hide network interfaces (hot)
PUT /api/sparks/:id/llm-ports Replace all LLM ports (hot)
POST /api/sparks/:id/llm-ports Add an LLM port (hot)
DELETE /api/sparks/:id/llm-ports/:port Remove an LLM port (hot)
PUT /api/sparks/:id/llm-port LLM port — backward-compat (hot)
GET /api/sparks/:id/llm/daily Daily busy decode/prefill tok/s (port, days)
POST /api/sparks/:id/llm/bench Start decode benchmark (202); poll/cancel/clear on the same path
POST /api/sparks/:id/llm/prefill-bench Start prefill + TTFT context sweep (202); poll/cancel/clear on the same path
POST /api/sparks/:id/llm/quality-bench Start quality suite (202); GET lists active / last / history summaries, GET :benchId returns a full run, DELETE :benchId cancels, DELETE clears history
GET /metrics Prometheus exposition (opt-in, see Prometheus)
GET /api/settings Global settings
PUT /api/settings Update global settings
WS /ws Real-time metrics stream

There is no application authentication on the HTTP/WebSocket API. sparkDash therefore binds to loopback and refuses direct LAN binding. Use an SSH tunnel, authenticated TLS reverse proxy, or Tailscale Serve; see Remote access.

/api/fleet-energy samples the configured fleet independently every two seconds. It estimates each node as GPU board draw + a CPU utilization model (5.2–65 W) + 23 W of memory/network/base overhead, clamped to the DGX Spark power envelope. Current and hourly fleet watts require fresh, simultaneous telemetry from every node; coverage fields make gaps explicit. Minute buckets are persisted at mode 0600 for rolling 24-hour and 31-day windows. Wh/output-token is computed from one observation per available LLM endpoint on each head or standalone node (workers are skipped, since they front their head's engine). These values are estimates, not wall-meter measurements. Restart sparkDash after changing fleet membership so the persisted series has one stable node set.

Finished minutes are also rolled up once into a permanent monthly archive (config/fleet-energy-monthly.json, read through /api/fleet-energy/monthly): per UTC month and node watt-hours and coverage, the fleet total, output tokens and Wh/output-token. A minute is added 10 seconds after it ends, a stored watermark guarantees it is never counted twice, and the first start backfills from the existing 31-day file. The archive is not touched by Reset…, restarts or fleet membership changes (the previous scope's file is rolled up, after the same plausibility checks as a normal load, before it is set aside). Only DELETE /api/fleet-energy/monthly removes months (all of them only with confirm=delete-all-history). A clock that jumps ahead cannot hide later minutes: nothing is folded, and no watermark is kept, beyond the clock plus one day (a stored watermark beyond that is clamped back on load, totals unchanged).


Configuration

Global settings (UI or API)

Gear icon in the header, or GET/PUT /api/settings:

Setting Default Description
Refresh rate (pollIntervalMs) 2000 ms How often every unit is polled for GPU, CPU/RAM, network, memory bandwidth, LLM and ComfyUI metrics, and how often the dashboard is pushed an update (the dialog offers 1 s to 10 s; the server accepts 0.5 s to 60 s). Remote units are polled over SSH, so 1 s costs the most. Memory bandwidth (nvidia-smi dmon, which blocks ~1 s) never goes below 2 s. Storage, liveness, Tailnet, Hermes and the NV_ERR scan keep their own cadences. A POLL_INTERVAL_* env var, when set, pins its domain instead
Default LLM port 8888 Default for new Sparks
Hide offline Sparks false Hide offline Sparks on Overview
Hide worker nodes false Hide Worker-role Sparks from Overview and the tab bar
Temperature unit Celsius Display GPU temperature in °C or °F
Time zone Browser Hour and day boundaries and labels on the Tokens, Energy and Activity charts (IANA name such as Europe/Paris; empty follows each browser). Daily token buckets and monthly energy stay UTC
Benchmark share image true Decode/prefill Copy results becomes a split button: the label copies the text summary, the caret offers Copy as text / Copy as image on hover or click. Turn it off to keep the plain button. The image copies where the page has an image clipboard (HTTPS or localhost); over plain http on a LAN IP the card downloads instead
Detailed VRAM breakdown true The VRAM bar on the Overview cards and the GPU panel is split by what holds the memory — LLM engine (largest GPU process while an endpoint is serving), system/CPU (GB10 unified pool), other GPU use — over a free track, and turns amber/red on low free memory (GB10: under 8 / 4 GB; discrete GPU: under 2 / 1 GB) rather than on a high percentage. Hover or focus for the breakdown, including the engine's KV fill where the backend reports it. Turn it off for the single percentage bar
Show Fleet Energy true Overview card with rolling fleet power estimates (the Fleet energy page is always in the sidebar)
Electricity price / currency not set / $ Per kWh, for the estimated cost on the Fleet energy page; empty hides cost
Show LLM Token Totals true Overview card with cumulative tokens per model
Show Recent activity true Overview card with the latest health and kernel events (the Activity page is always in the sidebar)
Show Fleet decode / Fleet power / Memory in use / Tokens today tile true Each of the four Overview tiles can be hidden on its own; the rest share the row (Tokens today also covers the Online count that replaces it when no tokens are tracked)
Show active fleet exceptions false Overview strip for offline units, throttling, disk and LLM alerts
Show search and status filters false Overview search field and status dropdown
Save benchmark debug traces false Store prompts, HTTP ids and GPU samples in bench history (larger files)
Prometheus metrics false Serve GET /metrics for Prometheus / Grafana (see Prometheus); 404 while off

Environment variables

Copy .env.example to .env if needed:

Variable Default Description
BIND_HOST 127.0.0.1 HTTP and WebSocket listen address. A non-loopback bind without SPARKDASH_TOKEN is open to anyone who can reach it (see SPARKDASH_ALLOW_OPEN_REMOTE).
SPARKDASH_TOKEN (empty) Bearer token. When set, it is required for every mutation and WebSocket connection, and for REST reads on a non-loopback bind. The browser prompts for it when needed (Settings → Access token to change it).
SPARKDASH_ALLOW_OPEN_REMOTE 1 Unset, empty, or 1: a non-loopback bind without SPARKDASH_TOKEN stays open. 0: refuse to start without a token (fail closed).
SPARKDASH_ALLOWED_HOSTS (empty) Only for a reverse proxy on a custom domain: comma-separated names a loopback bind should also answer to. localhost, IP addresses, this machine's hostname and its Tailscale name work without it.
PORT 5555 HTTP + WebSocket listen port
LLM_PORT 8888 Default LLM probe port
COMFY_PORT 8188 Default ComfyUI probe port
POLL_INTERVAL_GPU (setting) GPU poll (ms). Unset: follows Settings → Refresh rate; set: pins GPU polling regardless of the setting
POLL_INTERVAL_COMFY (setting) ComfyUI probe poll (ms); same rule
POLL_INTERVAL_CPU (setting) CPU / RAM poll (ms); same rule
POLL_INTERVAL_NETWORK (setting) Network poll (ms); same rule
POLL_INTERVAL_STORAGE 5000 Storage poll (ms): disk usage and the read/write speed on the Storage panel
POLL_INTERVAL_LLM (setting) LLM probe poll (ms); same rule
POLL_INTERVAL_BANDWIDTH (setting, ≥ 2000) Memory bandwidth / dmon poll (ms); same rule, and when it follows the setting it never drops below 2000
POLL_INTERVAL_HERMES 600000 Hermes Agent update check poll (ms)
POLL_INTERVAL_TAILSCALE 30000 Tailnet probe poll (ms)
TAILSCALE_PROBE_TIMEOUT_MS 8000 Timeout for tailscale status --json (ms)
POLL_INTERVAL_NVERR 60000 Kernel journal scan for NVRM NV_ERR_NO_MEMORY (ms)
HERMES_UPDATE_TIMEOUT_MS 600000 Hard timeout for running hermes update over SSH (ms)
POLL_INTERVAL_LIVENESS 5000 Online/SSH liveness check (ms)
SPARKDASH_SECRETS_KEY (auto) Passphrase or 64-char hex for secret encryption
HOST_PROC_PATH /host/proc Host proc mount inside container
HOST_SYS_PATH /host/sys Host sys mount
HOST_ROOT_PATH /host/root Host root mount
SSH_IDENTITY_FILE (unset) Path inside the process to a private key (ssh -i). Use when the bind-mount is not a default OpenSSH name.
SSH_CONTROL_PERSIST_SECONDS 60 Idle SSH transport persistence in seconds, capped at 3600. Set to 0 to disable multiplexing.
FLEET_ENERGY_JSON_PATH config/fleet-energy.json Rolling fleet-energy persistence path
FLEET_ENERGY_MONTHLY_JSON_PATH config/fleet-energy-monthly.json Permanent monthly energy archive

For compatibility, SSH_CONTROL_PERSIST is accepted as a seconds-based fallback when SSH_CONTROL_PERSIST_SECONDS is unset. The existing SSH_MULTIPLEX=0 switch also disables reuse. SSH tunnels always use an independent connection.

The listener and both Compose files default to 127.0.0.1. Existing Docker users who opened http://<host-ip>:5555 must migrate to an SSH tunnel, authenticated reverse proxy, Tailscale Serve, or BIND_HOST=0.0.0.0 SPARKDASH_TOKEN=... (without the token a 0.0.0.0 bind is open to the network unless SPARKDASH_ALLOW_OPEN_REMOTE=0). Recovery: BIND_HOST=127.0.0.1 docker compose up -d --force-recreate.

Adding a unit

  1. Click + next to Sparks in the sidebar (or Add Spark / GPU host in the Sparks menu on a phone).
  2. Choose Unit type:
    • NVIDIA DGX Spark — the default; hardware summary shows DGX Spark specs and the CX7 IP field is available.
    • Dedicated GPU host — any Linux machine with an NVIDIA GPU. It is monitored exactly like a Spark (SSH + nvidia-smi) but is not reported as a DGX Spark: the header shows a detected hardware summary (GPU model, CPU, RAM) instead of fixed GB10 specs, and the page shows separate RAM and VRAM panels (VRAM from nvidia-smi, RAM from system memory). On the unit page, RAM → Network → Storage stack in the right column with GPU filling the left column. A host with more than one GPU needs nothing extra: every card nvidia-smi lists is collected, the header names them all, the GPU panel shows a block per card, and metrics.gpu stays the aggregate (hottest / busiest card, summed power and VRAM) with the per-card detail under gpu.gpus[].
    • Windows PC with an NVIDIA GPU — a Windows machine with the OpenSSH Server and the NVIDIA driver; see Windows PCs. Always a remote unit; no model launchers, CPU temperature or Hermes / Tailnet checks.
  3. Set Name and choose whether this is This host. Local units do not require a LAN IP or SSH; their optional LAN IP enables browser links. On a dedicated GPU host it also directs Wake-on-LAN. Remote units require a LAN IP/host, SSH user, and key or password. Key auth in Docker needs a key mounted into the container (see Quick start).
    • LLM host (optional, llmHost in the unit's config): when the model API listens on one specific address of a multi-interface Spark, pin it here; probes and benchmarks then use it instead of the LAN IP (SSH keeps using the SSH host). An unreachable llmHost is never tunnelled to the SSH host's loopback.
  4. Test shows pass/fail/skipped for host collectors/SSH and each enabled service (LLM, ComfyUI, Hermes Agent, Tailnet). Every enabled capability must pass; disable an unavailable optional service before saving if it should not be monitored.
  5. Save — a tab appears and metrics start streaming.

Power controls (shutdown / Wake-on-LAN)

  • Shutdown (per Spark or Shutdown All on Overview) runs the host helper /usr/local/bin/spark-shutdown with passwordless sudo. The helper contract is two invocations:

    Invocation Expected behaviour
    spark-shutdown Schedule the graceful shutdown
    spark-shutdown --check Print an acknowledgement, exit 0, change nothing

    --check is what proves authorization before anything is scheduled, so a sudoers rule scoped to the helper is enough:

    sparky ALL=(root) NOPASSWD: /usr/local/bin/spark-shutdown
    

    A helper without --check still works when sudo is granted more broadly (the authorization probe falls back to sudo -n true), but a rule limited to the helper path needs --check support.

  • On a local unit, the helper runs on the Spark itself. When the dashboard is in Docker that means the invocation first enters the host mount namespace (nsenter --mount=/host/proc/1/ns/mnt -- sudo -n …), using the same HOST_PROC_PATH mount and privileged: true the collectors already need. A bare-host install calls sudo directly. The helper always resolves against the host filesystem, so it does not need to exist inside the container.

  • Wake / Wake All send a UDP magic packet (port 9) to dedicated GPU hosts whose NIC and firmware support Wake-on-LAN. The MAC is taken from the enP7s7 interface while the host is online (persisted as detectedMacAddress), or from a MAC override in Edit Spark. Broadcast is derived as /24 from LAN IP, or 255.255.255.255 if LAN IP is missing. DGX Spark does not wake this way — the GB10 onboard NIC has no Wake-on-LAN — so the Wake control is not offered on Spark units.

  • On a Windows PC shutdown runs shutdown.exe /s /t 5 over SSH; Wake-on-LAN needs a MAC address entered by hand.

  • Batch shutdown only targets online Sparks; offline nodes are skipped.

  • Power APIs are mutations: with SPARKDASH_TOKEN set they require it; without it they are open on loopback (local trust) and on a remote bind, unless SPARKDASH_ALLOW_OPEN_REMOTE=0 makes that bind fail closed.

Themes

Choose a theme in Settings → Appearance (or the sun / moon button in the sidebar):

Theme Notes
Dark (default) Neutral grays, true black base, muted amber accent
Light Warm paper whites
White Cool neutral whites
OLED True black for OLED panels

Choice is stored in localStorage.


Security

  • SSH passwords are not stored in sparks.json and are never returned by the API.
  • Passwords are encrypted with AES-256-GCM in config/sparks-secrets.json (survives restarts).
  • Encryption key: config/.secrets-key (auto-generated) or SPARKDASH_SECRETS_KEY. Do not delete the key file or encrypted secrets become unreadable.
  • Target validation rejects clearly unsafe IPv4 targets (link-local 169.254.0.0/16, 0.0.0.0/8, multicast/reserved ≥ 224). Private, loopback, and public addresses are allowed so LAN and remote Sparks work.
  • SSH and HTTP probes use short timeouts (about 5 s SSH connect, 3 s HTTP) so a hung host cannot stall the poll loop.
  • Prefer SSH keys over passwords. In Docker, mount the private key into /root/.ssh (see Quick start); passwords are the only SSH secret the app stores itself.
  • Loopback installs remain local-trust. A remote bind (BIND_HOST not loopback) without SPARKDASH_TOKEN is open by default: anyone who can reach the port can read telemetry, change settings and power units off, and the header shows an Open access warning (dismissible per browser). Set SPARKDASH_TOKEN to require a bearer token for mutations and remote telemetry/WebSocket, and SPARKDASH_ALLOW_OPEN_REMOTE=0 to refuse to start a remote bind without one. GET /api/health reports which applies as authMode: loopback-open, bearer, open-remote, or required-missing.
  • One-off remote benchmark hosts must be listed in SPARKDASH_BENCH_HOSTS.
  • Tested operator capacity for this remediation: 12 units.

Scripts

Command Purpose
npm run dev Vite (5173) + Express (5555) together
npm run dev:server Express only (node --watch)
npm run dev:client Vite only
npm run build Production frontend → dist/
npm run build:watch Rebuild dist/ on every save
npm run typecheck tsc --noEmit
npm test Server tests (node --test) and frontend tests (vitest); also npm run test:server / test:frontend
npm start Production server (node server/index.js)
npm run docker:up docker compose up -d
npm run docker:prod Same as docker:up
npm run docker:rebuild docker compose up --build -d
npm run docker:dev Dev Compose
npm run docker:dev:build Dev Compose with rebuild
./deploy.sh Recreate container; --build, --frontend flags

How it works

Local vs remote Sparks

One SystemCollector path for both modes. When spark.isLocal is true, metrics come from host sysfs/proc and nvidia-smi (often via nsenter into the host namespace). Remote Sparks wrap the same commands in a shared sshExec() helper (key agent or sshpass). The helper reuses an authenticated OpenSSH transport by default so frequent metric polls do not create a new SSH/PAM login lifecycle each time. Set SSH_CONTROL_PERSIST_SECONDS=0 to disable reuse. For kind: "host" units, actual hardware (GPU model, driver version, CPU, RAM) is detected once and cached in place of the static DGX Spark specs, and GPU VRAM comes straight from nvidia-smi while system RAM is read from /proc/meminfo.

Remote SSH sessions and host memory

Older sparkDash versions could create hundreds of SSH/PAM login sessions per minute on each remote host. Issue #73 documents the resulting session churn and observed polkitd memory growth. Connection reuse reduces this churn while retaining the collector refresh cadence. After updating, verify that metrics keep advancing and that new SSH authentications/PAM session opens fall after the initial connection; a new SSH client process for each collector command is still expected.

If host memory remains low, compare Linux MemAvailable and per-process resident/swap usage. Memory retained by polkitd requires separate OS investigation: polkit PR #653 fixes a reference leak in NoNewPrivileges queries. Check whether your distribution's polkit package includes that fix. SSH reuse neither applies the OS patch nor releases memory already retained by another process.

Health findings

server/health/HealthEvaluator.js runs a handful of rules over metrics sparkDash already collects, once per Spark, and reports them as health[] on the unit's snapshot. A rule never guesses: a missing metric raises nothing. Findings carry a severity, a detail and a hint, show up as chips on the Overview cards and in full on the unit page, and the monitor writes an Activity event when one appears or clears (never for the first baseline after a restart). The kernel Xid / OOM counters come from one cached journalctl -k scan per minute.

Windows units

A unit with platform: "windows" is reached over the same SSH helper, but every command is a PowerShell script sent as -EncodedCommand (so it works whatever the OpenSSH server's default shell is). Two scripts per poll cover the GPU (nvidia-smi, or the GPU Engine/Adapter Memory performance counters on an AMD iGPU) and the CIM classes for RAM, CPU load, disks and adapters; their output is parsed into the same shapes the Linux collectors return. See Windows PCs.

Graceful degradation

Collectors catch errors and return zero/default metrics instead of crashing the loop. After sustained liveness failures, a Spark is marked offline; the UI shows stale or empty states rather than hard errors.

Hot configuration

Name, IP, SSH credentials, LLM port, and device/interface filters update the running SparkMonitor without tearing down poll loops or losing rate baselines. Registry writes are atomic (temp file + rename).

LLM probe

Each configured LLM port gets its own LlmProbe instance running in parallel. Probes auto-detect backends:

  • llama.cpp — /slots for live decode rates; model from /props
  • ds4-server (Entrpi/ds4-on-spark) — /v1/models (owned_by: ds4.c) + Prometheus ds4_* token counters for live tok/s
  • EXL3 (ExLlamaV3 tools/serve_openai.py) — /v1/models (owned_by: exl3) or /health {ok, busy}; live tok/s from /health cumulative counters
  • q27 (signalnine/q27 engine) — /v1/models (owned_by: q27) or Prometheus q27_* series; live tok/s from q27_*_processed counter diffs (completion-based totals as fallback), exact computed-only prefill with the cached/uncached split doubling as the prefix-cache hit rate, TTFT/E2E/ITL p95 histograms, and constant-0 preemptions (FIFO admission, no wait queue)
  • TensorFold (ashhart/TensorFold) — /v1/models (owned_by: tensorfold). No /metrics. Live tok/s comes from /health when it publishes prompt_tokens_total, completion_tokens_total, and prefill_seconds_total (decode is the completion-counter diff; prefill is prompt tokens ÷ prefill time). The slots tile reads streams.max and requests_running on CUDA 0.6.0, and max_batch_size on MLX. A stock server that only returns {ok: true} stays at 0. Decode/prefill benches and the showcase work regardless.
  • FreeToken — /v1/models (owned_by: FreeToken), or a /v1/stats document with throughput.decode_tps / prefill_tps and lifetime prompt/completion totals. Those rates are FreeToken's own 5-second window (0 when idle). requests.ttft_mean_ms is a mean; requests.p95_ms is end-to-end latency, not TTFT p95. Queue length, slot capacity, and preemptions stay unset.
  • vLLM / sglang — /v1/models; sglang via /server_info (last_gen_throughput when metrics off; /get_server_info fallback), vLLM via Prometheus /metrics counters (scientific notation supported)

Rates are derived from per-probe cumulative counter diffs (or SGLang sticky throughput while it moves). Multiple ports can be added or removed at runtime without restarting the monitor.

Live probes use the LAN IP on remote units, and fall back to an SSH forward to the unit's loopback after two failed direct probes (the LLM panel then shows a via SSH chip). Decode, prefill and quality benches try that same HTTP target first; if it is closed they open an SSH local-forward onto the remote’s 127.0.0.1 so loopback-bound servers (ds4 start.sh default) can still be measured. The tunnel is torn down when the job finishes or is cancelled.


Contributing

Contributions are welcome. Conventions:

  • Server: plain JavaScript ESM
  • Client: TypeScript + React
  • Prefer extending the shared Spark model over per-unit special cases
  • Run npm test and npm run typecheck before a pull request. Tests that start the server must point every config file at a temp directory (server/__tests__/isolatedEnv.js)
  • Read AGENTS.md and CODEBASE.md first

License

Apache License 2.0 from version 2.0.0 onward. Copyright 2026 Mia's AI Lab. See NOTICE.

Versions up to 1.9.0 were released under the MIT License (LICENSE-MIT), and those releases stay MIT-licensed.


Acknowledgements

sparkDash is built and maintained by Mia's AI Lab.

Contributors. Thank you to everyone who sent code, fixes and ideas: @MikeGibbsOnyx, @Lesilva, @vincenzopalazzo, @Acermax, @danielkuykendall23-boop, @0xdfi, @BHCC2025, @ayylemao, @SashaMIT, @krunkosaurus, @andrei-dotdna, @0xWhiteMage, @kesslerio, @nkavassalis, @Olyno, @willy92wins and @saitakarcesme, plus Nyx Voss (who wrote the macOS collectors with C. Michael Gibbs, @MikeGibbsOnyx), Liao Shiwu and Philip Eriksson, and to everyone who reported issues. The full list is on the contributors page. The sparkdash-dgx-cluster Glance widget was contributed by @linxichen.

Projects and data sparkDash builds on

  • tool-eval-bench by SeraphimSerapis (MIT) powers Tool Eval Bench. It adapts the scenario methodology of ToolCall-15 by stevibe (MIT) and credits the Typed Decisions dataset from the LocalLLaMA organization (Apache 2.0).
  • GSM8K (OpenAI, MIT) and MMLU (Dan Hendrycks, MIT) supply the maths and knowledge questions in the Quality bench. The instruction-following category is inspired by IFEval (Zhou et al., Google).
  • Hermes Agent by Nous Research can be monitored and updated from sparkDash.
  • The health findings were inspired by spark-doctor by joeynyc (MIT). No code from it is used.
  • Fonts: Geist and Geist Mono by Vercel (SIL Open Font License 1.1). The bolt in the logo follows the "zap" icon from Feather (MIT).
  • Built with React, Vite, Tailwind CSS, Express, ws, undici, dnd kit and dotenv. Licenses and copyright notices are in THIRD_PARTY_NOTICES.md.

Built for the NVIDIA DGX Spark (GB10) on ARM64.

About

sparkDash ⚡ — Multi-DGX Spark Monitoring Dashboard

Resources

Stars

658 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages