本实验不修改 FLA kernel,因此不能直接写出每个 CTA 的开始/结束 timestamp。 Profiler 只用于解释 clean timing,不取代 clean timing。
先确认当前 Nsight Compute 可用的 metric:
ncu --version
ncu --query-metrics > profiles/available_metrics.txt正式比较在同一个 profiler 进程内完成:只做一次 local autotune,然后在同一
outer NVTX range 内依次发射 local 与 sink-heavy:
ncu --set full \
--nvtx \
--nvtx-include "nsa_local-vs-sink-heavy_dkv_core/" \
-o profiles/e1_local_vs_sink_dkv \
python scripts/profile_stage.py \
--config configs/e1_h100.json \
--pattern local-vs-sink-heavy \
--stage dkv_core \
--iterations 1 \
--benchmark-result results/raw/e1_h100.json脚本先在 NVTX range 外用 local 完成 backward compile/autotune,再构造两个
目标,并断言两次目标 warmup 都没有改变 dQ/dKV best config。
--benchmark-result 还会自动核对 clean benchmark JSON 中的
calibration_best_config;不一致时脚本直接终止,该 profile 不会被采集。
重点从当前版本实际存在的 metrics 中选择:
- kernel duration;
- waves per multiprocessor;
- achieved occupancy / active warps;
- long-scoreboard 与 memory-dependency stalls;
- DRAM/L2 traffic;
- Tensor Core activity。
解释边界:
- fan-in statistics 给出 work distribution;
- clean CUDA-event latency 给出完整 dKV makespan;
- Nsight 给出 stall/occupancy/wave 证据;
- 三者联合可以讨论 kernel tail,但不能写成“测得最慢 CTA 为 X μs”。