Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **Background-aware foreground clipping — `XR_DXR_depth_budget` (#318).** A transparent overlay used to throw away *every* pixel behind the display plane: content behind the plane carries positive disparity while being drawn *over* desktop pixels at zero disparity, and the eyes can't reconcile "it covers the icon" with "it is behind the icon". But that contradiction needs a background with a **horizontal** depth cue right there — text, icons, window edges. Over a plain wall of colour, a vertical gradient or horizontal stripes there is nothing to disagree with, and the rear of the model looks fine. The runtime now measures the desktop behind the app (off the capture the display's plug-in already takes for transparency) and publishes a ramped, advisory **rear depth budget**; the provider chains it on `XrViewState` at every `xrLocateViews` and adds it to the per-eye foreground far. So with an empty desktop behind it the avatar's back **grows in** after ~0.4 s, and when a text window moves behind it the clip **slides shut** in ~0.15 s. `foregroundOnlyClip` is unchanged as the app's opt-in — the budget only decides how much rear volume that clip currently allows.
- **Honours both clip paths.** BiRP takes it through the native per-eye far (`dxr_prov_get_eye_clip`); URP's `DisplayXR/ForegroundClipURP` gets it as the new `_DXRRearOffset` global, because its shader derives each eye's display-plane distance itself (per eye *and* per zone) rather than taking ours. HDRP has no foreground-clip pass and is unaffected.
- **`DisplayXRContentBounds`** — a new component for the content root (the avatar, the model). It unions the world bounds of its renderers each frame and hands the box to the provider, which projects it through every eye and reports the union as `XrContentBoundsDXR` on `xrEndFrame`, so the runtime measures only the desktop **behind the content**. Without it the runtime judges the whole window — an empty Notepad's own menu and status bars measure a cue of 0.93 and keep the budget shut even when the avatar is in the far corner. Strongly recommended for any transparent overlay; purely a hint, and every failure path (a corner behind an eye, a box that clamps off-canvas) falls back to the whole canvas rather than to "neutral".
- **The silhouette, not a box around it (`XrContentMaskDXR`, spec v3).** A rectangle around a character is roughly three times its area, so most of what the runtime measured was background the model never covers — and any horizontal structure in that surplus closed the clip. Against a runtime advertising spec version 3 the provider now chains the app's **content occupancy mask**: `DisplayXRTransparentOverlay` already renders the silhouette per eye, unions L+R, and reads it back for `SetWindowRgn`, and that mask is already normalised to the window client rect — exactly the grid the extension wants — so the whole producer is one call in the existing readback callback and **no app needs any change**. Native copies it, reduces it to at most 256 per side with an **any-coverage** filter (never an average: a thin limb that only partly covers a cell still occupies it), and chains it *beside* `XrContentBoundsDXR` rather than instead of it, so a v2 runtime still gets the rect it understands. Opt out with `DisplayXRTransparentOverlay.reportContentMask`. Runtime precedence: mask → bounds → 3D zones → whole canvas, so every fallback is the conservative direction.
- **The silhouette, not a box around it (`XrContentMaskDXR`, spec v3).** A rectangle around a character is roughly three times its area, so most of what the runtime measured was background the model never covers — and any horizontal structure in that surplus closed the clip. Against a runtime advertising spec version 3 the provider now chains the app's **content occupancy mask**: `DisplayXRTransparentOverlay` already renders the silhouette per eye, unions L+R, and reads it back for `SetWindowRgn`, and that mask is already normalised to the window client rect — exactly the grid the extension wants — so the whole producer is one call in the existing readback callback and **no app needs any change**. **The two artefacts are not the same thing in general** — the click-through region is about which pixels were actually *painted* (post-clip alpha is correct for it), while the depth-budget mask must describe the content as it would render at an *unrestricted* budget, or the clip state feeds back into the runtime's measurement and the rear clip oscillates (`displayxr-runtime#1470`, spec v4). They coincide here only because the shared producer rasterises **pre-clip geometry** with the far plane defeated (`DisplayXRSilhouette.shader`), never a post-clip alpha readback — which is why this plugin satisfies v4 by construction. Native copies it, reduces it to at most 256 per side with an **any-coverage** filter (never an average: a thin limb that only partly covers a cell still occupies it), and chains it *beside* `XrContentBoundsDXR` rather than instead of it, so a v2 runtime still gets the rect it understands. Opt out with `DisplayXRTransparentOverlay.reportContentMask`. Runtime precedence: mask → bounds → 3D zones → whole canvas, so every fallback is the conservative direction.
- **`DisplayXRDepthBudget`** — read-only `State` / `FarOffsetVH` / `CueEnergy` / `RearOffsetWorld` for HUDs, debugging and tests, plus one log line per state change (never per frame). Nothing to wire.
- **Apply as delivered; do not smooth.** The runtime already time-ramps the value (~300 ms opening, ~150 ms closing), so the plane glides rather than pops; an app-side filter on top fights that ramp for a slower, less predictable plane.
- **Nothing changes without it.** The extension is enabled only when the runtime enumerates it, and the offset is 0 — clip exactly on the plane, byte-for-byte today's behaviour — on an older runtime, on a runtime whose display processor supplies no background source, or for an **opaque** session. That last case is deliberate: the runtime's transparent flag is fixed at `xrCreateSession`, so a busy-background budget of 0 can arrive for a session that never composites over the desktop, and the app's own state has to win.
Expand Down
9 changes: 9 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -106,6 +106,15 @@ build will tell you. One container run catches them all:
unions and reads back exactly that mask for `SetWindowRgn`, already window-normalised -
so **no app needs a change**; the mask is chained BESIDE the bounds rect so a v2 runtime
still gets the rect. Precedence: mask -> bounds -> 3D zones -> whole canvas.
**The mask and the click-through region are not necessarily one artefact.** The window
region is about which pixels were *painted*, so post-clip alpha is right for it; the
depth-budget mask must be the silhouette at an **unrestricted** budget (spec v4), or the
clip state feeds back into what the runtime measures and the rear clip oscillates ~1 Hz
on a static desktop (`displayxr-runtime#1470`). They coincide here only because the
shared producer rasterises **pre-clip geometry** with the far defeated (the z-pin in
`Runtime/Resources/DisplayXRSilhouette.shader`) rather than reading back swapchain alpha
— so this plugin is v4-correct by construction. Do not "optimise" that pass into an
alpha readback.
`DisplayXRDepthBudget` exposes state/value for HUDs and one log line per state change. Absent extension, older runtime, no background
source, or an **opaque** session → offset 0 = today's clip-at-the-plane, exactly. The
opaque case is deliberate: the runtime's transparent flag is fixed at `xrCreateSession`,
Expand Down
31 changes: 31 additions & 0 deletions Runtime/Resources/DisplayXRSilhouette.shader
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,11 @@
// SetWindowRgn — outside the silhouette the OS treats our window as if
// it didn't exist (cross-process click-through with full fidelity).
//
// The SAME readback has a second consumer: it is chained to the runtime
// as XrContentMaskDXR (XR_DXR_depth_budget v3+), telling the runtime
// which patch of desktop to measure when deciding this overlay's rear
// depth budget. See the mask invariant on the z override below.
//
// Renderer-agnostic: ZWrite/ZTest off + Cull off so it rasterizes
// regardless of depth, backface culling, or material setup. We only
// care "did anything draw here?" — the answer drives whether the OS
Expand Down Expand Up @@ -52,6 +57,9 @@ Shader "Hidden/DisplayXR/Silhouette"
v2f o;
float4 worldPos = mul(unity_ObjectToWorld, v.vertex);
o.pos = mul(_DXRViewProj, worldPos);
// LOAD-BEARING TWICE OVER — DO NOT REMOVE, AND DO NOT
// REPLACE THIS PASS WITH A POST-CLIP ALPHA READBACK.
//
// Pin clip-space z to mid-frustum (z = 0.5w → z_ndc = 0.5,
// inside both D3D [0,1] and GL [-1,1] ranges) so the
// projection's near/far planes never clip foreground/back
Expand All @@ -62,6 +70,29 @@ Shader "Hidden/DisplayXR/Silhouette"
// click-through holes that the woven render doesn't have. The
// mask only needs x/y coverage; depth is irrelevant here
// (ZTest Always, no depth buffer), so overriding z is safe.
//
// That was reason one (the click-through region). Reason
// two: this z override is ALSO the v4 invariant of
// XR_DXR_depth_budget, whose mask this readback feeds. The
// spec requires the content mask to be the silhouette the
// content WOULD have at an unrestricted budget — rasterised
// ignoring the far clip. Because this pass draws pre-clip
// GEOMETRY with the far defeated (here) and the depth-budget
// foreground clip is a separate fragment-discard pass over the
// camera's colour target (_DXRForegroundFar / _DXRRearOffset),
// this plugin already satisfies v4 by construction and cannot
// enter the oscillation of displayxr-runtime#1470, where the
// mask is derived from the RENDERED (post-far-clip) content and
// so feeds the clip state back into the runtime's measurement:
// clipped → small silhouette over quiet desktop → budget opens
// → rear half appears → silhouette grows over busy desktop →
// budget closes → round again, ~1 Hz on a static desktop.
//
// The click-through region may keep post-clip alpha (it is about
// which pixels were painted); the depth-budget mask may not. The
// two consumers share one artefact here ONLY because that
// artefact is pre-clip geometry. Sourcing this mask from the
// swapchain alpha to "save a pass" would reintroduce #1470.
o.pos.z = o.pos.w * 0.5;
return o;
}
Expand Down
31 changes: 31 additions & 0 deletions docs~/architecture/click-through-mask.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,35 @@ DisplayXRTransparentOverlay.LateUpdate (C#) native (win32.c)
displays a column-by-column union of the eyes; a cyclopean-only mask is narrower than the
visible silhouette on high-disparity foreground geometry (hands, hat) and would clip.

## The mask's second consumer: the rear depth budget

The very same readback is chained to the runtime as **`XrContentMaskDXR`**
(`XR_DXR_depth_budget` v3+, `DisplayXRTransparentOverlay.OnHitMaskReadback` →
`dxr_prov_set_content_mask`, opt out with `reportContentMask`). It tells the runtime
*which patch of desktop to measure* when deciding how much rear depth this overlay may
draw. Sharing one artefact is what made v3 free here — but the two consumers have
**different contracts**, and only one of them is satisfied by a rendered image:

| Consumer | What it must describe |
|---|---|
| `SetWindowRgn` click-through region | which pixels were **actually painted** — post-clip alpha is correct |
| `XrContentMaskDXR` (spec **v4**) | the silhouette **as it would render at an unrestricted budget** — the far clip must be **ignored** |

If the depth-budget mask is derived from the *rendered* (post-far-clip) content, the clip
state feeds back into the runtime's measurement and the rear clip oscillates roughly once
a second on a completely static desktop — `DisplayXR/displayxr-runtime#1470`, which is
what drove spec v4.

**This plugin satisfies v4 by construction and must keep doing so.** The silhouette pass
rasterises **pre-clip geometry** (`CommandBuffer.DrawRenderer` of `clickableRenderers`)
with `ZTest Always` and the clip-space z pinned to mid-frustum, so no near/far plane can
touch it (`Runtime/Resources/DisplayXRSilhouette.shader`); the depth-budget foreground
clip is an entirely separate fragment-discard pass over the *camera's colour target*
(`_DXRForegroundFar` / `_DXRRearOffset`) and never runs over the mask RT. **Do not
replace this pass with a swapchain-alpha readback** to save a draw — that is precisely
the change that reintroduces #1470. Content bounds (`DisplayXRContentBounds`) are
likewise geometry-derived (`Renderer.bounds`) and clip-independent.

## The alignment invariant (the one thing that must hold)

The mask is rasterized **full-window** (in overlay client pixels), but the woven 3D content
Expand Down Expand Up @@ -131,6 +160,8 @@ is woven.
| File | Role |
|---|---|
| `Runtime/DisplayXRTransparentOverlay.cs` | C# per-eye silhouette render + async readback (`RenderHitMaskAndRequestReadback`, `OnHitMaskReadback`); `DumpHitMaskPng` (`DXR_DUMP_HIT_MASK=1`). |
| `Runtime/Resources/DisplayXRSilhouette.shader` | The silhouette rasteriser. Its clip-space z pin is what keeps the mask independent of the near/far planes — and therefore of the rear depth budget (spec v4 / runtime#1470). |
| `Runtime/DisplayXRDepthBudget.cs`, `Runtime/DisplayXRContentBounds.cs` | The other half of the second consumer: the published budget, and the (geometry-derived, clip-independent) bounds rect chained beside the mask. |
| `native~/displayxr_win32.c` | `displayxr_set_overlay_hit_mask` — stamp mask into the target rect → `ExtCreateRegion` → `SetWindowRgn`; `region_target_hwnd`/`region_target_ready`. |
| `native~/displayxr_native_shared.cpp` | `displayxr_get_canvas_rect_px` (reader) + `displayxr_set_canvas_rect` (writer, re-homed in #166) over the `s_canvas_rect` statics. |
| `native~/displayxr_xrprovider/displayxr_provider_session.cpp` | `dxr_prov_get_zone_count`/`dxr_prov_get_zone_rect_px` (multi-zone rects); `dxr_prov_set_3d_zone_rect` (zone weave); publishes the stereo matrices the mask reads. |
Expand Down
35 changes: 29 additions & 6 deletions native~/displayxr_extensions.h
Original file line number Diff line number Diff line change
Expand Up @@ -484,9 +484,12 @@ typedef XrResult(XRAPI_PTR *PFN_xrSetWorkspaceViewRigDXR)(XrSession session, con
// there); the provider derives its own vH in that case.
//
// Source of truth: displayxr-runtime/src/external/openxr_includes/openxr/
// XR_DXR_depth_budget.h (SPEC_VERSION 2). These declarations are byte-identical to
// it. Enable ONLY when enumerated — the extension is the app's opt-in AND the
// runtime's gate (a session that never enables it costs the runtime nothing).
// XR_DXR_depth_budget.h (SPEC_VERSION 4). These declarations are byte-identical to
// it - v3 added XrContentMaskDXR and v4 changed only the WORDING of what that mask
// must contain (see the struct below), so no layout has moved since v2 and no
// consumer needs a recompile. Enable ONLY when enumerated — the extension is the
// app's opt-in AND the runtime's gate (a session that never enables it costs the
// runtime nothing).
#define XR_DXR_DEPTH_BUDGET_EXTENSION_NAME "XR_DXR_depth_budget"
// This is our FLOOR - the minimum runtime spec version this plugin is designed
// against - NOT the latest version transcribed below. Do NOT bump it to match the
Expand Down Expand Up @@ -550,9 +553,29 @@ typedef struct XrContentBoundsDXR {
float marginNormalized;
} XrContentBoundsDXR;

// INPUT (v3): the app's content OCCUPANCY MASK for this frame - the union over ALL
// views of its rendered silhouette. Chain on XrFrameEndInfo::next, beside or instead
// of XrContentBoundsDXR.
// INPUT (v3, semantics clarified in v4): the app's content OCCUPANCY MASK for this
// frame - the union over ALL views of the silhouette of the content subject to the
// rear budget, rasterised IGNORING the far clip, i.e. AS IT WOULD RENDER AT AN
// UNRESTRICTED BUDGET. Chain on XrFrameEndInfo::next, beside or instead of
// XrContentBoundsDXR.
//
// THE MASK MUST NOT BE A FUNCTION OF THE BUDGET THE RUNTIME PUBLISHED. v3's original
// wording ("its rendered silhouette") made it one, and that closes a feedback loop:
// clipped -> a small silhouette over a quiet patch of desktop -> the budget opens ->
// the rear half appears -> the silhouette grows over a busy patch -> the budget
// closes -> round again, roughly once a second on a STATIC desktop
// (displayxr-runtime#1470; v4 says it explicitly, and runtimes ratchet the measured
// region as a floor for apps still shipping the clipped one). This mirrors the
// bounds rule exactly: geometry the CURRENT budget happens to be clipping away still
// belongs in the mask.
//
// So the mask and the app's CLICK-THROUGH window region are NOT the same artefact in
// general - the window region keeps the clipped alpha, because that one is about
// which pixels were actually painted. They coincide in THIS plugin only because the
// shared producer (DisplayXRTransparentOverlay's silhouette pass) rasterises PRE-CLIP
// GEOMETRY rather than reading back post-clip alpha; see the z-pin in
// Runtime/Resources/DisplayXRSilhouette.shader. Deriving either one from the swapchain
// alpha would reintroduce #1470.
//
// Why a mask at all: XrContentBoundsDXR is a RECTANGLE, and a rectangle around a
// character-shaped silhouette is roughly three times its area, so most of what the
Expand Down