Skip to content

Shadow performance programme: current shadows within 2 ms and a fluid native-resolution frame #525

Description

@pasquelin

Why

The maintainer reports stuttering and the loss of 120 FPS on ordinary examples. His 27 September 2026 screenshot of spin-an-astrolabe shows 35 FPS, 21.43 ms GPU, 38,432 triangles, 1,184 draw calls including 1,172 shadow draws, 586 shadow pages drawn, 580 requested, 0 cached and 0 refetched. This is an observation, not a controlled benchmark: screenshot commit, canvas size, browser and sampling interval are unknown. It strongly motivates measuring shadow submission/raster work; it does not prove that all invalidation is unnecessary or that the cache is thrashing.

This is the coordination and acceptance issue for the shadow programme, not one implementation PR. Earlier improvements (#763, #766, #811, #818, #815) do not establish the final performance target. #868 was rejected; its single-depth-draw design is not delivered. The latest #850 proof is KO. The body below consolidates the 27 September maintainer request and supersedes outdated implementation assumptions in earlier comments; measured results remain evidence.

To do

Order: baseline → measure the closed #867 and repair #850 → #831 capacity/cache steps → measured fine selection → #344 scheduling/submission → conditional #22. The owner splits each programme step into a single-PR child at pick time (about 500 handwritten lines maximum); no coder takes this parent as one rewrite. Existing image defects are not postponed behind speculative optimisations.

Examples and design rationale

Case Required behaviour What it distinguishes
Astrolabe rotating, fixed lamp/camera Update moving rings and affected shadow work; retain unchanged static contributions A small triangle count can still cause excessive per-page submission
Same astrolabe stopped through its public control Work settles after real changes finish A moving scene is not expected to have the cache ratio of a still one
One box moving over a static floor Static depth remains reusable; old and new dynamic coverage are correct Static/dynamic separation, including removal of the previous shadow
Wall outside camera, shadow on visible floor Preserve the wall's contribution Receiver-driven selection is not camera-frustum culling of casters
Camera turn/resize and a lamp moving every frame Newly needed shadows are current, without holes or delayed silhouettes Feedback latency, pool lifetime, page validity

Code context

Anchors checked against local develop 454fc23b3; re-resolve lines at implementation time.

  • packages/sdk-browser/src/lighting/direct/shadowWgsl.ts:59 — existing deduplicated receiver requests and filter-aware page lookup.
  • packages/sdk-browser/src/webgpu/shadow/pageRequests.ts:28 — asynchronous feedback; :102 maps readback after submission. It is not a synchronous CPU wait today, but CPU consumption still affects scheduling.
  • packages/sdk-core/src/scene/light-shadow/plan.ts:30 — existing planner, allocation/admission and validity rules; preserve their invariants.
  • packages/sdk-browser/src/webgpu/pages/render/encodeShadows.ts:61 — host region planning; encodeShadowPass.ts:70 — raster and per-region viewports/draws.
  • packages/sdk-browser/src/gpu/shadow/cullShader.ts:59, gpu/shadow/staticLayer.ts:25, webgpu/shadow/mobility.ts:16 — existing selection and static/dynamic infrastructure, reused.
  • packages/sdk-browser/src/webgpu/pages/render/cpuSteps.ts:1, packages/sdk-core/src/contracts/shadowMetrics.ts:1 — existing observation contracts, not a second profiler.
  • packages/sdk-browser/src/webgpu/core/materialPasses.ts — material pass encoding, for the material-surface item moved from Material performance: redundant pixel work removed and every GPU pass of the frame attributed #685.

Proof

  • Measurer only, one browser on the machine. Pin before/after commits, machine/GPU, browser/backend, display Hz, CSS canvas size, actual drawing-buffer size, DPR, camera, animation input, quality, memory grant and warm-up. Use native resolution; the requested reference case is 1728×1117 CSS, DPR 2, and the astrolabe's actual canvas must also be recorded rather than inferred from the screenshot.
  • Three comparable runs per side, with run-to-run spread. Publish CPU and GPU frame envelopes separately, rAF intervals p50/p95/p99/max, missed refreshes (interval > 1.5 × display period), long tasks, cold-start/resize events separately from steady animation. A 60 Hz headless run cannot prove 120 Hz presentation. Do not use the known unlit uncapped path as a performance proof.
  • Programme targets: GPU frame p95 ≤ 8 ms, shadow contribution ≤ 2 ms with a documented attribution method, main-thread engine work ≤ 2 ms as in Streaming without holes: rules and objectives for geometry, memory and shadows #483. At an actual 120 Hz display the frame budget is 8.33 ms; publish missed refreshes and p99, not just average FPS. No engine long task >16 ms during the reference navigation. Targets are requirements, not achieved results.
  • Targeted cases: astrolabe rotation/stop/drag, falling-boxes, moving point light (Moving-light shadows stay current: GPU scheduling and submission within the frame budget #344), plus an open world; sponza and facade-7 for the whole frame moved from Material performance: redundant pixel work removed and every GPU pass of the frame attributed #685. Walker/car, opaque/masked/blended shadows and resize are regression cases selected for the changed mechanism; full campaign belongs to the release.
  • Image: unchanged resolution, lighting, shadow distance and filtering; deterministic animation checkpoints against the same pose, A/A baseline, no disappearing/stale shadow or temporal trail. Apply CONTRIBUTING's image rule, never accept an unlit frame as a fast frame. No quality reduction to meet the timing.
  • Every implementation child has behaviour tests and check:changed, test:changed, validate; the named browser proof is run in the measurer queue after merge. This parent closes only when its children and the end-to-end targets are measured as met, not merely when implementation issues close.

Links

Parent #483. Implementation/work ownership: #867, #850, #831, #344, #22. Related #685 (narrowed on 29 September; its frame target lives here), #26, #234, #816. Delivered historical work: #811, #818, #815, #902, #1128; rejected experiment: #868. Latest #850 KO.

Published techniques: comparison and applicability

  • Epic: Virtual Shadow Maps describes fine page selection and separate static caching. These justify removing redundant work; they do not establish a speedup for our implementation.
  • Epic support, UE 5.7 receiver masks explains why partial dynamic pages may be uncached while static pages remain cached. Its directional-light result is not proof for the astrolabe's point light.
  • Unity HDRP mixed cached shadows combines cached static geometry and dynamic rendering, with extra atlas memory and copy cost. Measure both sides of that trade-off here.
  • Blender EEVEE shadow module documents GPU allocation/update without host synchronization. Its shadow pipeline describes direct atlas writes using atomic minima. These are architectural references, not evidence that Blender's backend facilities or exact depth behaviour are portable to WebGPU.
  • Apple: load/store actions explains attachment traffic between passes. Measure whole-frame cost and pass layout on the maintainer's Apple GPU, not draw count alone.
  • WebGPU specification defines the available commands and limits. Record actual browser/device features; no mandatory hardware ray tracing, multi-viewport extension or unverified native-only facility.

Reimplement from public descriptions only; never copy engine source. Dynamic ray-traced lighting is not a prerequisite for this one-lamp regression. The techniques above are candidates constrained by correctness, measured cost and the web backend.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    lightingDynamic lighting, shadows, bounce, reflectionsto measureMerged; waiting for the recette's timing🔴 criticalCritical priority: rendering correctness and FPS

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions