Skip to content
baristahausPublic

About

Map accelerator capability

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

echolock

echolock measures what an edge GPU actually does rather than what it declares. Every probe sends a known stimulus, checks the result against an independent oracle, and times the response. It records the environment around each run: clocks, temperature and power every second, and kernel Xid/OOM/hang lines. Results are time series, and a fast wrong answer counts as a failure.

GB10 (DGX Spark class, sm_121a) is the first target. Probe definitions are backend-neutral; CUDA is the first backend and Vulkan compute the second. Design and principles: docs/brief.md.

Status

Milestone State
1. Skeleton, environment capture, Xid watcher, 1 s sampler done, and run on GB10: env capture, 1 s sampler and Xid watcher all work on Asus GX10 hardware (10 runs, kernel log available, 0 events) (docs/gb10-first-run.md)
2. Port memory, file_page, sync with JSONL output done; run on GB10 sm_121a with all oracles passing (docs/gb10-first-run.md)
3. Derived SM clock, pointer-chase latency map, per-format MMA throughput done; run on GB10 sm_121a — fp4 measured 210 TFLOPS there, and the flat clock trace is flagged in finding 4 (docs/gb10-first-run.md)
4. Load generators and grid runner done; loads measured + tested on this box, GPU load compiles (needs GB10 to run) (load/)
5. Vulkan backend all six probes done and verified on AMD via RADV: memory, pointer_chase, mma (fp16/bf16/int8 bit-exact), sync, clock, file_page (backends/vulkan/)
AMD: environment, kernel log, docs, full probe suite done on amdgpu hardware; all six probes run on AMD via RADV and through ./echolock run (docs/amd.md)
NVIDIA discrete (sm_120a, RTX 5090): first real-GPU run of the CUDA backend all six probes pass on hardware including fp4 (kind::f8f6f4, ~514 TFLOPS measured); first run exposed and fixed four probe bugs: memory copies clobbered the device fill before the contention reads, the mma D-fragment store had d1/d2 swapped, the mma fp8/fp4 pack4 row/column addressing was wrong, and fp4's packing assumed 2 values/byte instead of the PTX kind::f8f6f4 central-4-bits byte container (plus a declare_compiled_arch stream-ordering race and a retry for transient calibration timeouts). fp4 needs the family build (sm_120a/sm_121a). Runs alongside a resident inference engine with small --gib.

| NVIDIA edge (sm_121a, GB10 / Asus GX10): first GB10 run of the CUDA backend | all six probes pass on hardware, 110 checks 0 failed, fp4 210 TFLOPS, GPU read 243 GB/s median, contention ratio 0.7, and it cost a co-resident TP-3 inference engine nothing measurable at --gib 1. Two fixes came out of it (host.cpu_model was empty on Arm; --pin-cpus was accepted and silently ignored outside sync) and three open notes: the unified-memory util_mem_pct: 0, a zero-variance clock trace, and the order-blindness of the XOR oracle (docs/gb10-first-run.md) |

Quick start (on the GPU machine)

./echolock build                      # nvcc from PATH or /usr/local/cuda; ARCH defaults to sm_121a,
                                      # or sm_120a automatically on a CC 12.0 GPU
./echolock run memory                 # -> runs/<utc>-memory-cuda-idle/ with report.md
./echolock run sync --repeat 3        # three runs, one combined report (flakiness becomes visible)
./echolock run sync --pin-cpus 0,10   # also blocking-schedule runs pinned to CPU 0 and CPU 10
./echolock run file_page --scratch-dir /path/on/nvme --drop-caches
./echolock run clock                  # derived SM clock, idle and under a memory load
./echolock run pointer_chase --scratch-dir /path/on/nvme
./echolock run mma                    # fp16/bf16/tf32/int8/fp8/fp4 tensor throughput
./echolock run memory --backend vulkan    # bandwidth/copies/contention via RADV (AMD)
./echolock run pointer_chase --backend vulkan  # GPU memory latency map
./echolock grid memory --conditions idle,cpu50,mem80   # one run per load condition
./echolock run memory -- --gib 2 --seconds 1          # after --: arguments for the probe
./echolock report runs/A runs/B       # several runs side by side (e.g. idle vs loaded)

Utilities: ./echolock env prints the environment snapshot, ./echolock sample streams 1 s samples, and ./echolock kmsg streams kernel GPU/OOM/hang events. Reading /dev/kmsg needs root or kernel.dmesg_restrict=0; without it the run still happens, and the manifest records that the kernel log was not watched (and counts messages the ring lost). --drop-caches needs root or passwordless sudo, and only runs before steps that don't manage page residency themselves: file_page's resident and evict phases are never preceded by a drop, because that would evict the pages they are about to measure. Licence: MIT (LICENSE).

Requirements: CUDA 13 toolkit (nvcc) and Python 3.11+ with the standard library only. The probes link the CUDA runtime statically and reach the driver API through entry points, so they need no -lcuda and no vendor libraries.

AMD GPUs are supported for the environment side today: clocks, temperature, power, utilisation, VRAM and kernel-log resets come from the amdgpu driver's own sysfs files, with no ROCm, no AMD SMI and no other library. What is open on AMD's side and what an AMD probe backend would use is written up in docs/amd.md. All six probes also run on AMD through the Vulkan backend (RADV), so the full measurement suite — bandwidth, copies, contention, the latency map, tensor throughput, device clocks, submit/fence lag and host-page behaviour — is measured on AMD hardware today, through an open driver. Where CUDA and Vulkan semantics differ on AMD (no real-time device clock in GLSL, no file-backed host-memory import, GLSL cooperative-matrix accumulator semantics), the probes document the difference instead of forcing the CUDA shape onto Vulkan; see backends/vulkan/README.md for the findings and docs/lessons.md for the debugging traps behind them.

Layout

probes/<name>/spec.md     stimulus, oracle, parameters, metrics (backend-neutral)
backends/cuda/<name>.cu   CUDA implementation; common/echolock.cuh holds the JSONL, oracle and clock helpers
backends/vulkan/          later
load/                     load generators (cpu, mem, io, gpu) speaking one JSONL protocol
env/                      capture.py (versions, firmware, limits), sampler.py (1 s), kmsg.py (Xid watcher), nvml.py, amdgpu.py
run/                      plans.py (probe → process steps), runner.py (steps + env → run directory), grid.py (load conditions)
report/                   markdown.py: declared vs measured, one column per run label
docs/                     brief.md (design), record-format.md (JSONL schema), amd.md (AMD enablement),
                          lessons.md (debugging lessons), crosscheck.md (cross-platform plan),
                          review-request.md (for agent reviewers),
                          gb10-first-run.md (first GB10/GB10-OS run: outcomes, findings, caveats)
tests/                    python3 -m unittest discover -s tests -t .

Each run writes runs/<run_id>/ containing manifest.json, env.json, probe.jsonl, samples.jsonl, kernel.jsonl, stderr.log and report.md. Every stream uses CLOCK_MONOTONIC nanoseconds, so a kernel Xid lines up with the probe sample and the 1 s sample around it. Ctrl-C or kill stops the probe's process group (it runs in its own session, so nothing else would) and keeps the partial run: the manifest is always written and the report reads the directory anyway. Format: docs/record-format.md.

What changed from the ninfer-gb10 originals

  • Oracles. Every timed read or copy is checked against an XOR the host derives from the generator. The file_page pattern is now counter-based, so the host knows every row's XOR without reading the file.
  • Per-launch series. Every launch is timed on its own, where the originals reported a mean or a percentile.
  • Silent fallbacks. Compression granted but not faster, and eviction requested but pages still resident (checked with mincore), are recorded as feature records.
  • Eviction. file_page evicts with posix_fadvise(DONTNEED) by default instead of reading a file larger than RAM; the original method is still available as --evict-file.
  • Graph check. In the sync probe each graph node increments a counter, which checks that all nodes ran.

About

Map accelerator capability

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages