echolock measures what an edge GPU actually does rather than what it declares. Every probe sends a known stimulus, checks the result against an independent oracle, and times the response. It records the environment around each run: clocks, temperature and power every second, and kernel Xid/OOM/hang lines. Results are time series, and a fast wrong answer counts as a failure.
GB10 (DGX Spark class, sm_121a) is the first target. Probe definitions are backend-neutral;
CUDA is the first backend and Vulkan compute the second. Design and principles:
docs/brief.md.
| Milestone | State |
|---|---|
| 1. Skeleton, environment capture, Xid watcher, 1 s sampler | done, and run on GB10: env capture, 1 s sampler and Xid watcher all work on Asus GX10 hardware (10 runs, kernel log available, 0 events) (docs/gb10-first-run.md) |
| 2. Port memory, file_page, sync with JSONL output | done; run on GB10 sm_121a with all oracles passing (docs/gb10-first-run.md) |
| 3. Derived SM clock, pointer-chase latency map, per-format MMA throughput | done; run on GB10 sm_121a — fp4 measured 210 TFLOPS there, and the flat clock trace is flagged in finding 4 (docs/gb10-first-run.md) |
| 4. Load generators and grid runner | done; loads measured + tested on this box, GPU load compiles (needs GB10 to run) (load/) |
| 5. Vulkan backend | all six probes done and verified on AMD via RADV: memory, pointer_chase, mma (fp16/bf16/int8 bit-exact), sync, clock, file_page (backends/vulkan/) |
| AMD: environment, kernel log, docs, full probe suite | done on amdgpu hardware; all six probes run on AMD via RADV and through ./echolock run (docs/amd.md) |
| NVIDIA discrete (sm_120a, RTX 5090): first real-GPU run of the CUDA backend | all six probes pass on hardware including fp4 (kind::f8f6f4, ~514 TFLOPS measured); first run exposed and fixed four probe bugs: memory copies clobbered the device fill before the contention reads, the mma D-fragment store had d1/d2 swapped, the mma fp8/fp4 pack4 row/column addressing was wrong, and fp4's packing assumed 2 values/byte instead of the PTX kind::f8f6f4 central-4-bits byte container (plus a declare_compiled_arch stream-ordering race and a retry for transient calibration timeouts). fp4 needs the family build (sm_120a/sm_121a). Runs alongside a resident inference engine with small --gib. |
| NVIDIA edge (sm_121a, GB10 / Asus GX10): first GB10 run of the CUDA backend | all six probes pass on hardware, 110 checks 0 failed, fp4 210 TFLOPS, GPU read 243 GB/s median, contention ratio 0.7, and it cost a co-resident TP-3 inference engine nothing measurable at --gib 1. Two fixes came out of it (host.cpu_model was empty on Arm; --pin-cpus was accepted and silently ignored outside sync) and three open notes: the unified-memory util_mem_pct: 0, a zero-variance clock trace, and the order-blindness of the XOR oracle (docs/gb10-first-run.md) |
./echolock build # nvcc from PATH or /usr/local/cuda; ARCH defaults to sm_121a,
# or sm_120a automatically on a CC 12.0 GPU
./echolock run memory # -> runs/<utc>-memory-cuda-idle/ with report.md
./echolock run sync --repeat 3 # three runs, one combined report (flakiness becomes visible)
./echolock run sync --pin-cpus 0,10 # also blocking-schedule runs pinned to CPU 0 and CPU 10
./echolock run file_page --scratch-dir /path/on/nvme --drop-caches
./echolock run clock # derived SM clock, idle and under a memory load
./echolock run pointer_chase --scratch-dir /path/on/nvme
./echolock run mma # fp16/bf16/tf32/int8/fp8/fp4 tensor throughput
./echolock run memory --backend vulkan # bandwidth/copies/contention via RADV (AMD)
./echolock run pointer_chase --backend vulkan # GPU memory latency map
./echolock grid memory --conditions idle,cpu50,mem80 # one run per load condition
./echolock run memory -- --gib 2 --seconds 1 # after --: arguments for the probe
./echolock report runs/A runs/B # several runs side by side (e.g. idle vs loaded)Utilities: ./echolock env prints the environment snapshot, ./echolock sample streams 1 s
samples, and ./echolock kmsg streams kernel GPU/OOM/hang events. Reading /dev/kmsg needs root
or kernel.dmesg_restrict=0; without it the run still happens, and the manifest records that the
kernel log was not watched (and counts messages the ring lost). --drop-caches needs root or
passwordless sudo, and only runs before steps that don't manage page residency themselves:
file_page's resident and evict phases are never preceded by a drop, because that would evict the
pages they are about to measure. Licence: MIT (LICENSE).
Requirements: CUDA 13 toolkit (nvcc) and Python 3.11+ with the standard library only. The
probes link the CUDA runtime statically and reach the driver API through entry points, so they
need no -lcuda and no vendor libraries.
AMD GPUs are supported for the environment side today: clocks, temperature, power, utilisation,
VRAM and kernel-log resets come from the amdgpu driver's own sysfs files, with no ROCm, no AMD
SMI and no other library. What is open on AMD's side and what an AMD probe backend would use is
written up in docs/amd.md. All six probes also run on AMD through the Vulkan backend (RADV), so the full measurement
suite — bandwidth, copies, contention, the latency map, tensor throughput, device clocks,
submit/fence lag and host-page behaviour — is measured on AMD hardware today, through an open
driver. Where CUDA and Vulkan semantics differ on AMD (no real-time device clock in GLSL, no
file-backed host-memory import, GLSL cooperative-matrix accumulator semantics), the probes
document the difference instead of forcing the CUDA shape onto Vulkan; see
backends/vulkan/README.md for the findings and
docs/lessons.md for the debugging traps behind them.
probes/<name>/spec.md stimulus, oracle, parameters, metrics (backend-neutral)
backends/cuda/<name>.cu CUDA implementation; common/echolock.cuh holds the JSONL, oracle and clock helpers
backends/vulkan/ later
load/ load generators (cpu, mem, io, gpu) speaking one JSONL protocol
env/ capture.py (versions, firmware, limits), sampler.py (1 s), kmsg.py (Xid watcher), nvml.py, amdgpu.py
run/ plans.py (probe → process steps), runner.py (steps + env → run directory), grid.py (load conditions)
report/ markdown.py: declared vs measured, one column per run label
docs/ brief.md (design), record-format.md (JSONL schema), amd.md (AMD enablement),
lessons.md (debugging lessons), crosscheck.md (cross-platform plan),
review-request.md (for agent reviewers),
gb10-first-run.md (first GB10/GB10-OS run: outcomes, findings, caveats)
tests/ python3 -m unittest discover -s tests -t .
Each run writes runs/<run_id>/ containing manifest.json, env.json, probe.jsonl,
samples.jsonl, kernel.jsonl, stderr.log and report.md. Every stream uses CLOCK_MONOTONIC
nanoseconds, so a kernel Xid lines up with the probe sample and the 1 s sample around it. Ctrl-C or
kill stops the probe's process group (it runs in its own session, so nothing else would) and
keeps the partial run: the manifest is always written and the report reads the directory anyway.
Format: docs/record-format.md.
- Oracles. Every timed read or copy is checked against an XOR the host derives from the generator. The file_page pattern is now counter-based, so the host knows every row's XOR without reading the file.
- Per-launch series. Every launch is timed on its own, where the originals reported a mean or a percentile.
- Silent fallbacks. Compression granted but not faster, and eviction requested but pages
still resident (checked with
mincore), are recorded asfeaturerecords. - Eviction. file_page evicts with
posix_fadvise(DONTNEED)by default instead of reading a file larger than RAM; the original method is still available as--evict-file. - Graph check. In the sync probe each graph node increments a counter, which checks that all nodes ran.