Skip to content

feat(d3d11_service): per-stage weave timing under DXR_WEAVE_GPU_TIMING (2D-under-the-lens P0) - #1824

Merged
dfattal merged 1 commit into
mainfrom
feat/weave-timing-p0
Oct 5, 2026
Merged

dfattal merged 1 commit into
mainfrom
feat/weave-timing-p0

Conversation

@dfattal

@dfattal dfattal commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

What

P0 of the "near-zero-cost 2D under the lens" plan: instrumentation only. It extends the existing DXR_WEAVE_GPU_TIMING=1 diagnostic on the D3D11 service present-owner path (comp_d3d11_service_weave_submit, the path that weaves the browser's whole-page 2D overlay) so every stage has its own cost. Later phases are judged against these numbers.

The knob stays off by default. The env var is read once (cached static). It writes one WARN line per client every 5 s, never per frame. With the knob off, the added cost is a handful of os_monotonic_get_ns() reads per submit and nothing else: no queries are created and nothing is logged.

Companion plug-in PR (same knob, DP-internal split): DisplayXR/displayxr-leia-plugin#302.

What is measured

GPU. Six timestamps per submit, on the service immediate context, inside one disjoint query. Read back non-blocking from an 8-slot ring, frames late. A slot still in flight when the ring wraps is dropped, never waited on.

stage brackets averaged over
clear1058 the #1058 woven-output clear frames that cleared the output
input_copy v6 crop CopySubresourceRegion (+ lift writes) v6 crop frames
sbs_clear_blit batch SBS scratch clear + per-rect L/R stretch blits batch frames
weave set_overlay_2d + process_atlas, i.e. the DP's weave including SR's in-weave 2D compose and the DP's alpha gate all frames
fallback_blit runtime post-weave premul overlay blit frames whose overlay took the fallback
epilogue wish publish (+ #73 diag dump when armed), up to the fence signal all frames
total BEGIN..END all frames

There is no woven-output copy inside the service. The caller copies the output back in its own process, so epilogue is the only service-side GPU work after the weave. v6 zero-copy and legacy single-rect frames have no ingest stage.

CPU wall time, per accepted submit:

stage measures
submit the whole synchronous weave_submit, from entry to return, including the fence signal + Flush
render_mutex the render_mutex_fair_lock wait
ctx_mutex the immediate_ctx_mutex wait
in_acquire the input IDXGIKeyedMutex::AcquireSync(0, 4)
ov_acquire the overlay AcquireSync(0, 4), averaged over submits that carried an overlay

Counts over the window:

  • accepted submits;
  • refused: submits that returned false after render_mutex was taken;
  • submits by path: v6zc / v6crop / batch / legacy;
  • submits by 2D-overlay route: dp (composited inside the weave), fallback (runtime post-weave blit), none;
  • unchanged: submits whose overlay the caller declared unchanged (v14).

How to read the log line

The old weave GPU time [...] line is replaced by:

weave timing [<adapter>] <win_w>x<win_h> <window>s: submits=N refused=N paths(v6zc=N v6crop=N batch=N legacy=N) overlay(dp=N fallback=N none=N unchanged=N) | CPU ms avg/max: submit A/M render_mutex A/M ctx_mutex A/M in_acquire A/M ov_acquire A/M | GPU ms avg/max n=N disjoint=N: clear1058 A/M (n=N) input_copy A/M (n=N) sbs_clear_blit A/M (n=N) weave A/M fallback_blit A/M (n=N) epilogue A/M total A/M
  • n= after GPU is the number of query sets read back in the window. It can be lower than submits when sets were dropped in flight or discarded as disjoint. Each GPU stage's (n=) is its own denominator.
  • weave is the number the later phases attack. Read it next to the plug-in's Leia D3D11 DP timing line, which splits it into SR weave vs alpha gate.
  • submit minus (render_mutex + ctx_mutex + in_acquire + ov_acquire) is the CPU time the service spends issuing work. A large render_mutex points at contention with the render thread, not at the weave.
  • overlay(dp=…) vs fallback= says which route the page's 2D layer actually took. When fallback>0, fallback_blit is the cost the DP route avoids.

The existing unconditional #625 weave timing / #625 weave split lines are left as they are.

Docs

Adds the missing DXR_WEAVE_GPU_TIMING row to the census in docs/roadmap/control-panel-performance-settings.md (Diagnostics / observers). The knob was not listed there before.

Testing

  • scripts\build_windows.bat build: green (=== ALL DONE ===; incremental rebuild after formatting also green, no new warnings).
  • Not yet panel-tested. The hub session runs the panel A/B. Set DXR_WEAVE_GPU_TIMING=1 in the service's environment: stop the service and relaunch it from a .bat that sets the env, through explorer.exe so it stays at medium integrity. A client-side env var does not reach it.

🤖 Generated with Claude Code

…G (2D-under-the-lens P0)

Extends the DXR_WEAVE_GPU_TIMING=1 diagnostic on the present-owner path
(comp_d3d11_service_weave_submit) so every stage is attributed:

- GPU (6 non-blocking timestamps per submit, 8-slot ring): #1058 output
  clear, v6 crop copy (+lift) / batch SBS scratch clear+blit, the DP
  weave (set_overlay_2d + process_atlas), the post-weave overlay
  fallback blit, and the wish-publish epilogue, plus the total;
- CPU wall time: render_mutex wait, immediate_ctx_mutex wait, input and
  overlay keyed-mutex acquires, and the whole synchronous submit;
- counts: accepted / refused submits, by path (v6 zero-copy / v6 crop /
  batch / legacy) and by 2D-overlay route (DP / fallback / none, and
  declared-unchanged).

One WARN line per client every 5 s ("weave timing [...]"), replacing the
old "weave GPU time" line. Off by default. Adds the knob to the env-var
census in docs/roadmap/control-panel-performance-settings.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dfattal
dfattal marked this pull request as ready for review October 5, 2026 08:26
@dfattal
dfattal requested a review from a team as a code owner October 5, 2026 08:26
@dfattal
dfattal merged commit f8a3b0f into main Oct 5, 2026
40 checks passed
@dfattal
dfattal deleted the feat/weave-timing-p0 branch October 5, 2026 08:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant