Skip to content

feat(observability): add NVTX tracing and Prometheus metrics - #242

Draft
BruceLoveDecimal wants to merge 1 commit into
RL-Align:mainfrom
BruceLoveDecimal:feat/observability
Draft

feat(observability): add NVTX tracing and Prometheus metrics#242
BruceLoveDecimal wants to merge 1 commit into
RL-Align:mainfrom
BruceLoveDecimal:feat/observability

Conversation

@BruceLoveDecimal

Copy link
Copy Markdown

Observability: NVTX tracing + Prometheus metrics
Adds a two-layer observability stack for RL-Kernel.

NVTX tracing (micro) — rlk:: ranges around every C++ operator binding plus Python stage ranges (rlk::score, rlk::train_step, …), visible on an Nsight Systems timeline. Opt-in via RL_KERNEL_NVTX=1; off by default (cached branch, no cost).

Prometheus metrics (macro) — a per-process /metrics endpoint exposing token throughput, rollout QPS, stage latencies, backend/fallback state, KV-cache fragmentation, peak memory, and training loss, plus a ready-made Grafana dashboard. prometheus-client is an optional extra; without it every call is a silent no-op.

Notes

Instrumentation is exception-transparent and per-stage (never per-token). Measured overhead +0.9% on an RTX 5090.
Existing StageResult.metrics dicts are unchanged.
Verified on GPU: CUDA extension builds, full suite green (572 passed), nsys shows the rlk::
ranges, live /metrics served.
Closes #72

Signed-off-by: BruceLoveDecimal <liuqihao970610@gmail.com>
@coderabbitai

coderabbitai Bot commented Jul 22, 2026

Copy link
Copy Markdown

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 0e7667e0-e3b4-44e9-a102-42eebb77b396

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@@ -111,6 +114,14 @@ def __init__(
self.reward_adapter = reward_adapter or default_reward_adapter

def score(self, inputs: StatelessForwardInputs) -> StatelessForwardResult:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A better way is to use decorator here. See many hardcoded modifications like this

@nodeeeeee nodeeeeee left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You may write a decorator as wrapper instead of using _train_impl and _score_impl.

Also, could you help enable counting op calls, in the metrics we can see how many times an op is called.

Tested on my h100 and everything seems good.

@zhangj1an

Copy link
Copy Markdown
Collaborator

Thanks for helping me review this @nodeeeeee T-T

I will resolve conflicts and approve this PR on Tuesday (I need to check whether we need follow up PRs as the codebase has evolved a bit). Also thank you bruce for your beautiful code!!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants