Skip to content

perf(d3d11_service): batch weave clears only the stale SBS scratch regions (2D-under-the-lens P1) - #1826

Closed
dfattal wants to merge 1 commit into
mainfrom
perf/batch-sbs-scratch
Closed

dfattal wants to merge 1 commit into
mainfrom
perf/batch-sbs-scratch

Conversation

@dfattal

@dfattal dfattal commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

P1 of the "2D under the lens" perf work. The P0 timing (#1824) measured sbs_clear_blit at 0.7-1.2 ms/frame on the browser present-owner (spec-v3 batch) weave: RTX 3080 laptop, 3840x2040, whole-page overlay, static gallery. This PR cuts the clear half of that stage. The output stays bit-identical.

Why the batch path copies the browser's views into an SBS scratch

The browser submits one window-sized input. Each woven element's L|R pair sits squeezed side by side at that element's own window position, so a rect (x,y,w,h) holds L in [x, x+w/2) and R in [x+w/2, x+w). That is not a tiled atlas. The SR weaver's contract is one weave() per frame: per-rect calls at N times the frame rate degrade its predictor (#96). So the runtime builds one window-sized 2x1 SBS pair, (2*win_w) x win_h, by stretch-blitting each rect's halves to the rect's position in the L and R tiles. It then weaves the whole window in a single process_atlas. With v5 firstChunk the scratch is cleared to (0,0,0,0) first, so the gaps between tiles weave transparent and the browser can draw the woven output back whole-window.

Why zero-copy does not apply

ADR-030's zero-copy is about handing the DP an app atlas that already is the active mode's tiled atlas (u_tiling_can_zero_copy()). The batch input never is: its views are interleaved per element at window positions, not contiguous tiles. So no "crop", and therefore no zero-copy, is possible here, and this PR adds no zero-copy proxy of any kind. The path that does have a zero-copy case is spec-v6, where the caller packs a real N-view atlas. Moving the browser's whole-page-overlay mode onto v6 would remove the scratch entirely, but that is a browser-side change (follow-up).

What changes

The scratch used to be cleared whole on every frame-first submit: 7680x2040 RGBA8, 62.7 MB. Only the regions blitted since the previous frame-first can hold non-zero pixels; everything else was cleared last frame and nothing has written it since. So:

  • The scratch keeps a dirty-region list: two regions per blitted rect (L and R tile), clipped to the scratch rather than the tile, so a rect that crosses the window edge and bleeds into the other tile is tracked too. Lift rects are tracked as their whole rect, which is conservative.
  • A frame-first submit clears (ClearView with rects) only the dirty regions that this submit will not overwrite with an identical opaque blit. "Identical" means the same clipped region from a non-lifted rect. The test uses exactly the blit loop's eligibility rules.
  • It falls back to a whole clear when the scratch is new (contents undefined), the list overflowed (256 regions), the stale area is more than half the scratch, or DXR_WEAVE_SBS_FULL_CLEAR=1.
  • Static page (same rects every frame): nothing is cleared. Scrolling: last frame's tile footprints are cleared, not 62.7 MB.

Invariant: while weave_sbs_clean_valid holds, every pixel outside the dirty list is (0,0,0,0). The texture handed to the DP is bit-identical to the whole-clear version. Gaps are zero either way, and every rect is still blitted after the clear, in the same order. The weave blit is an opaque overwrite with no discard (convert_srgb=0, no corner/feather), so a skipped region really is fully rewritten. A caller that never sets firstChunk is unchanged (it never cleared and still doesn't).

What this PR does not reduce is the blit half of the stage (2 stretched draws per rect). The split below says how much of the 0.7-1.2 ms each half is. A follow-up could halve the scratch to win_w x win_h: un-stretched views, with the weaver doing the 2x upscale it already does on the legacy path. That would halve clear, blit and the weave's input read. It changes the weaver's input dims and the sampling phase on odd rect x, so it needs its own A/B.

Timing (P0 line kept)

weave timing [...] keeps every existing field (sbs_clear_blit is unchanged) and appends:

| sbs split: clear A/B blit C/D (n=N) frame-first clears whole=W partial=P none=Z cleared X.XX Mpx/clear

That comes from a seventh GPU timestamp (WEAVE_TS_SBS_CLEARED). It is stamped as a placeholder on every path so query sets always complete, and the batch path re-stamps it after its clear.

Expected savings

  • sbs_clear_blit drops by the clear's share. On a static page none should dominate and cleared should be ~0 Mpx/clear, against 15.7 Mpx before.
  • If the driver turned the old whole clear into a metadata fast-clear, the saving is small. The sbs split line will show this directly (clear before ~ after). I am stating that up front rather than promising a number.
  • weave is unchanged by this PR. The plug-in PR (displayxr-leia-plugin, alpha gate) is the other half of the weave-stage saving.

Test plan

Build: scripts\build_windows.bat build in a worktree (done; compiles clean).

Deploy (same as P0): scripts/push-runtime-pf.sh from the worktree. Relaunch the service via an explorer.exe .bat that sets DXR_WEAVE_GPU_TIMING=1 and SR_SIMULATED_VIEWER=0,0,600,63 in the service env.

  1. bperf A/B on the same build. Use the DisplayXR Browser whole-page overlay on the static gallery page at 100% overlay reuse. Take ~60 s per arm:

    • A: service env + DXR_WEAVE_SBS_FULL_CLEAR=1 (old behaviour).
    • B: without it.

    Compare sbs_clear_blit, sbs split: clear / blit, weave, total and CPU submit. Expect B to show none dominating, ~0 Mpx/clear and a lower clear. blit and weave should be flat. Also run a scrolling pass and expect partial with small Mpx.

  2. Visual, Add macOS runtime installer script #73 woven dump. touch %TEMP%\dxr_weave_dump_trigger on a static frame and again mid-scroll, in both arms. dxr73_weave_sbs and dxr73_weave_output should match A vs B. Expect a pixel diff of 0 on the SBS scratch for an identical page state. Look for no stale tiles in the gaps after a scroll or navigation.

  3. Eyeball. Scroll the gallery fast, navigate away and back, and resize the window (a resize recreates the scratch, so the first frame is a whole clear). No trails, no ghost tiles between elements.

  4. Lift rect (if a NeurD box is available): a lift-flagged element still weaves, and the area outside it stays transparent.

Risks

  • A writer to the scratch that the dirty list doesn't see would leave stale pixels in the gaps. Today only this function writes it (blits + lift, both tracked). Any future writer must mark the scratch dirty or set weave_sbs_clean_valid = false.
  • ClearView with rects is D3D11.1 (ID3D11DeviceContext4 is already required here).
  • Rollback without a rebuild: DXR_WEAVE_SBS_FULL_CLEAR=1.

🤖 Generated with Claude Code

…gions

The spec-v3 batch weave (the browser present-owner path) assembles a
(2*win_w) x win_h side-by-side scratch from the browser's window-sized input
and used to ClearRenderTargetView the WHOLE scratch on every frame-first
submit before blitting the rects into it - 62.7 MB of RGBA8 at 3840x2040.
Only the regions blitted since the previous frame-first can hold anything but
(0,0,0,0), so the scratch now carries a dirty-region list and the frame-first
clear covers just the regions this submit will not overwrite with an
identical opaque blit (ClearView with rects; a whole clear when the scratch
is new, the list overflowed, or the stale area exceeds half the scratch).
A static page clears nothing; a scrolling one clears last frame's tile
footprints. The scratch handed to the DP is bit-identical: gaps are
(0,0,0,0) either way and every rect is still blitted after the clear, in the
same order. Nothing about the DP handoff changes (same texture, same dims) -
not an ADR-030 zero-copy question; the batch input is per-rect squeezed SBS
at window positions, not a tiled atlas, so no zero-copy predicate applies.

DXR_WEAVE_SBS_FULL_CLEAR=1 restores the whole clear (A/B, rollback).

DXR_WEAVE_GPU_TIMING: the weave timing line keeps every existing field
(sbs_clear_blit unchanged) and appends a split - sbs clear vs blit GPU ms,
plus frame-first clears by kind (whole/partial/none) and Mpx cleared - via a
seventh timestamp stamped on every path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dfattal

dfattal commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator Author

Closing per David: measured on the panel (bperf qw_*, 4 alternating passes), the whole SBS clear was already a ~0.015 ms fast clear, and this saves ~0.01 ms. The real cost of the SBS step is the per-rect blit (0.7-1.2 ms). The follow-up is a half-width scratch or the browser handing over a packed (v6) atlas so the copy disappears.

@dfattal dfattal closed this Oct 6, 2026
@dfattal
dfattal deleted the perf/batch-sbs-scratch branch October 6, 2026 08:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant