Lightweight Heterogeneous Accelerator Monitoring & Workload History
Monitor today. Review tomorrow.
English | 简体中文
Like stars in a constellation, Constella gathers independent accelerator nodes into one observable cluster.
Constella is a lightweight accelerator monitoring platform for labs, AI teams, and personal compute servers. It natively supports heterogeneous clusters; version 0.1.4 supports NVIDIA GPUs and Ascend NPUs.
Unlike terminal tools that only show the current state, Constella automatically records accelerator workload history, making it easy to review completed training and inference jobs. It supports standalone servers and small clusters without requiring a heavyweight Prometheus/Grafana stack.
| Cluster Overview | GPU & Process Detail |
|---|---|
![]() |
![]() |
Workload Curves
Workload History
- Automatically record GPU curves for completed workloads.
- Review training and inference jobs from the last 7 days.
- Prefer high-resolution memory cache for recent short jobs, with SQLite rollups for persisted history.
Accelerator Monitoring
- Monitor a standalone server or a small GPU cluster from one Web UI.
- Track GPU utilization, memory, power, temperature, clocks, processes, users, PIDs, and command fingerprints.
- On supported NVIDIA GPUs, inspect SM activity, occupancy, aggregate Tensor Core activity, DRAM bandwidth, and non-Tensor FP16/FP32/FP64 pipelines through NVML GPM.
- Use NVML with
nvidia-smifallback for NVIDIA, and DCMI withnpu-smifallback for Ascend.
Multi-User Analytics
- See user GPU usage rankings, job duration rankings, node trends, and range-aware heatmaps.
- Detect low-utilization reservations and off-hour activity.
- Keep realtime monitoring available even when historical analytics are disabled.
Lightweight Deployment
- No root privileges, system service, Prometheus, or Grafana required.
- One manager process receives data from local and remote accelerator agents.
- Remote nodes only need Python, their vendor driver/runtime, and SSH access.
| Capability | nvitop | Prometheus/Grafana | Constella |
|---|---|---|---|
| Realtime GPU view | Yes | Yes | Yes |
| Workload history | No | Requires setup | Yes |
| Small cluster view | Limited | Yes | Yes |
| Lightweight setup | Yes | No | Yes |
| Web UI | No | Yes | Yes |
| User/job analytics | No | Custom dashboards | Built in |
Constella sits between terminal monitoring and a full observability stack: more historical and shareable than nvitop, but much lighter to deploy than Prometheus/Grafana for a small lab.
Install the PyPI distribution and start the managed local stack:
pip install constella-gpu
constella service startconstella-gpu is the full installation with backend, Web UI, and TUI. Smaller
deployments can install constella-gpu-web, constella-gpu-tui, or
constella-gpu-backend; see Packaging for the feature
matrix.
On an Ascend host, use constella service start --device ascend.
From a source checkout, start the manager and local accelerator agent with:
cd Constella
./scripts/service/setup.sh
./scripts/service/start.shFor an Ascend host, select the hardware backend explicitly:
./scripts/service/start.sh --device ascendOpen:
http://127.0.0.1:8765/overview
Or stay in the terminal with the keyboard-first TUI:
constella tuiThe TUI connects to the same realtime cluster stream as the Web UI. Use
constella tui --url https://gpu.example.com for a remote manager, or run the
equivalent constella-tui entry point.
The fifth TUI view presents compact high-resolution NVIDIA GPM performance curves and uses the same node, GPU, and time-range controls as the other views.
If the service runs on a remote server, forward the port from your local machine:
ssh -N -L 8765:127.0.0.1:8765 <user>@<server>Enable SQLite history when workload history and analytics are needed:
DB_PATH=run/constella.db ./scripts/service/start.shStart the high-resolution sidecar when short-job curve cache should run outside the manager process:
DB_PATH=run/constella.db HIGHRES_SIDECAR=1 ./scripts/service/start.shThe sidecar listens on 127.0.0.1:8766 by default and subscribes to the manager stream at ws://127.0.0.1:8765/api/highres/stream. Simple deployments can skip the sidecar; the manager still exposes the built-in /api/highres/* endpoints.
NVML GPM is automatically probed on NVIDIA agents and remains isolated from the
base NVML path. Set CONSTELLA_NVML_GPM=off to disable collection, or
CONSTELLA_NVIDIA_GPM_ROLLUP=off to keep realtime performance data without
persisting its rollups. CONSTELLA_NVIDIA_GPM_HIGHRES=off disables only the
in-memory performance curves. Ascend agents never initialize the NVIDIA provider.
Prepare the remote node manifest:
cp docs/nodes.example.yaml nodes.yamlEdit manager_url, manager_hostname, and each node's device (nvidia or
ascend), then configure passwordless SSH from the manager host to each node.
flowchart LR
M["Manager<br/>FastAPI + Web UI"] -->|"SSH setup/control"| A["gpu-node-a<br/>agent"]
M -->|"SSH setup/control"| B["gpu-node-b<br/>agent"]
M -->|"SSH setup/control"| C["gpu-node-c<br/>agent"]
A -->|"WebSocket samples"| M
B -->|"WebSocket samples"| M
C -->|"WebSocket samples"| M
Start remote GPU agents:
./scripts/cluster/start.shscripts/service/start.shcreatesrun/agent-tokenon first local-agent startup, andscripts/cluster/start.shuses that token for remote agents.- If the manager host should not monitor local GPUs, start with
LOCAL_AGENT=0. - Remote nodes do not need
uv; the manager syncs a minimal agent runtime.
Constella uses independent, explicitly selected hardware chains:
nvidia: NVML, thennvidia-smifallback.ascend: DCMI (libdcmi.so), thennpu-smifallback.
The DCMI backend exposes AICore and HBM utilization, memory, temperature,
power, PCI identity, driver/DCMI versions, and running-process memory.
Multi-die cards remain visible as one device card per die. The API includes
card_id, die_id, card_count, and accelerator_count; rated power and live
power are counted once per physical card, while duplicate PIDs across dies are
counted as one active process.
flowchart LR
LA["Local agent<br/>selected device chain"] -->|"WS /api/agents/ws"| M["Manager<br/>FastAPI ingest"]
RA["Remote agents<br/>NVML→nvidia-smi or DCMI→npu-smi"] -->|"WS /api/agents/ws"| M
M --> S["ClusterState<br/>latest snapshots + 120-point history"]
S --> API["HTTP /api/cluster/snapshot"]
S --> WS["WebSocket /ws/cluster"]
S -.optional.-> DB["SQLite<br/>rollups + sessions"]
DB -.optional.-> AN["Analytics + job curves"]
S -.optional.-> HR["Highres cache / sidecar"]
API --> UI["Vite TypeScript UI"]
WS --> UI
AN --> UI
HR --> UI
The manager does not sample GPUs directly. Local and remote nodes both report current sample points through the same agent WebSocket path. SQLite, analytics, and high-resolution job curves are optional side paths and do not block realtime snapshots. See Design for the full data flow.
- Design: architecture, data path, low-overhead strategy, and data contracts.
- Operations: startup, access, cluster agent management, status, and verification commands.
- SQLite History: persistence, rollups, maintenance, and job curves.
- NVIDIA GPM Performance: requirements, metric interpretation, retention, and troubleshooting.
- Cloudflare Tunnel: domain access without opening an inbound server port.
- Lab user system: Access identity, roles, audit, and multi-node Linux account binding.
- Lab deployment: build, environment, rollout, backup, and member onboarding.
- Node manifest example:
nodes.yamltemplate for remote agents. - PyPI CLI: installed service, probe, agent, and cluster commands.
- Packaging: build and safely smoke-test wheel and source distributions.
- Scripts: service, cluster, tunnel, maintenance, and dev script entry points.
packages/backend/ Python backend, agents, cluster manager, samplers, API/WebSocket
packages/lab/ Optional Access-protected lab identity and account-binding edition
packages/web/ Installable production Web assets
packages/tui/ Textual terminal client, theme, and usage notes
src/constella_gpu/ Full-distribution metadata package
frontend/ Vite + TypeScript frontend
scripts/ categorized service, cluster, tunnel, maintenance, and dev scripts
docs/ design and operations notes
tests/ unit tests
uv sync
uv run pytest
cd frontend
npm install
npm run buildFrontend dev server:
cd frontend
npm run devFor a release build, scripts/package/build.sh builds the frontend into the Web
and Lab distributions and produces all five wheel/source-distribution pairs.
GET /api/healthGET /api/cluster/snapshotGET /api/settingsPATCH /api/settingsWS /ws/clusterWS /api/agents/wsGET /api/history/gpuGET /api/history/tasksGET /api/usersGET /api/analytics/overviewGET /api/analytics/node/{node_id}GET /api/highres/statusGET /api/highres/jobsGET /api/highres/jobs/{job_key}GET /api/highres/jobs/{job_key}/gpuGET /api/highres/performanceGET /api/highres/jobs/{job_key}/performanceGET /api/docs
When SQLite is not enabled, history, analytics, and job curve search APIs return enabled:false; realtime cluster monitoring continues through /api/cluster/snapshot and /ws/cluster.


