Skip to content

Phase 6: Add profiling-guided fast-path expansion telemetry #35

Description

@inaniloquentee

Background

Split from #6 Phase 6: Fast Path Expansion From Profiling Bottlenecks.

This issue owns the profiling and target-selection layer after the consistency/audit and operator-contract stack is available. The goal is to make fast-path work follow measured vime bottlenecks rather than convenient kernel boundaries, while keeping strict-consistency claims separated from opportunistic speed paths.

Scope

  • Add clean warmup and stable-window profiling entry points for representative vime workloads.
  • Add or formalize NVTX ranges for rollout wait, rollout execution, training forward, logprob, backward, optimizer, weight sync, and communication.
  • Record communication metrics: AllReduce/AllGather count and time, rank skew, overlap status, and topology notes.
  • Record compute/communication segment telemetry: compute time, exposed communication time, overlapped communication time, idle/wait time, stream/event plan, bucket assignment, and overlap fallback reason.
  • Add a profiling report path that separates operator-level, actor-window, and full-step measurements.
  • Use Phase 3-5 attribution outputs to choose fast-path targets, especially TP communication and high-count small elementwise/reduce/copy kernels.
  • Add backend descriptors or telemetry fields that state whether communication is in-op, post-op NCCL, custom all-reduce, or overlapped bucket, and whether that policy is strict-eligible.
  • Keep linear_logp operator timing/memory claims separate from full-step claims.

Non-Goals

Acceptance Criteria

  • Profiling output names warmup, measurement window, model/workload shape, world sizes, dtype, backend choices, and runtime/build fingerprints.
  • NVTX/telemetry separates rollout, train compute, logprob, backward, optimizer, weight sync, and communication time.
  • Communication telemetry records count/time/skew/overlap and enough topology metadata to compare runs.
  • Fast-path candidates are selected from measured bottlenecks and linked back to Phase 3-5 attribution or profiling evidence.
  • Reports clearly separate operator-level, actor-window, and full-step metrics.
  • Strict-consistency eligibility is never inferred from performance telemetry alone.

Hardware / Execution Notes

  • Unit tests for report builders and telemetry normalization should run on CPU or single-card developer machines.
  • NVTX emission and single-GPU operator/actor-window profiling need one CUDA GPU.
  • Real communication and rank-skew validation need distributed GPUs, typically at least 2 GPUs for TP/collectives and preferably 4-8 GPUs for representative TP/SP/CP sweeps.
  • Blackwell/B-card hardware is only required for Blackwell-specific profiling or claims, not for the generic vime profiling framework.

Dependencies

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    consistencyRollout-training consistencyphase-6Phase 6 workstreamrl-kernelRL-Kernel integrationtelemetryStructured telemetry and observability

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions