Skip to content

Latest commit

 

History

History
48 lines (38 loc) · 1.63 KB

File metadata and controls

48 lines (38 loc) · 1.63 KB

Nsight protocol for dKV tail evidence

本实验不修改 FLA kernel,因此不能直接写出每个 CTA 的开始/结束 timestamp。 Profiler 只用于解释 clean timing,不取代 clean timing。

先确认当前 Nsight Compute 可用的 metric:

ncu --version
ncu --query-metrics > profiles/available_metrics.txt

正式比较在同一个 profiler 进程内完成:只做一次 local autotune,然后在同一 outer NVTX range 内依次发射 local 与 sink-heavy:

ncu --set full \
  --nvtx \
  --nvtx-include "nsa_local-vs-sink-heavy_dkv_core/" \
  -o profiles/e1_local_vs_sink_dkv \
  python scripts/profile_stage.py \
    --config configs/e1_h100.json \
    --pattern local-vs-sink-heavy \
    --stage dkv_core \
    --iterations 1 \
    --benchmark-result results/raw/e1_h100.json

脚本先在 NVTX range 外用 local 完成 backward compile/autotune,再构造两个 目标,并断言两次目标 warmup 都没有改变 dQ/dKV best config。 --benchmark-result 还会自动核对 clean benchmark JSON 中的 calibration_best_config;不一致时脚本直接终止,该 profile 不会被采集。

重点从当前版本实际存在的 metrics 中选择:

  • kernel duration;
  • waves per multiprocessor;
  • achieved occupancy / active warps;
  • long-scoreboard 与 memory-dependency stalls;
  • DRAM/L2 traffic;
  • Tensor Core activity。

解释边界:

  • fan-in statistics 给出 work distribution;
  • clean CUDA-event latency 给出完整 dKV makespan;
  • Nsight 给出 stall/occupancy/wave 证据;
  • 三者联合可以讨论 kernel tail,但不能写成“测得最慢 CTA 为 X μs”。