You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The maintainer reports stuttering and the loss of 120 FPS on ordinary examples. His 27 September 2026 screenshot of spin-an-astrolabe shows 35 FPS, 21.43 ms GPU, 38,432 triangles, 1,184 draw calls including 1,172 shadow draws, 586 shadow pages drawn, 580 requested, 0 cached and 0 refetched. This is an observation, not a controlled benchmark: screenshot commit, canvas size, browser and sampling interval are unknown. It strongly motivates measuring shadow submission/raster work; it does not prove that all invalidation is unnecessary or that the cache is thrashing.
This is the coordination and acceptance issue for the shadow programme, not one implementation PR. Earlier improvements (#763, #766, #811, #818, #815) do not establish the final performance target. #868 was rejected; its single-depth-draw design is not delivered. The latest #850 proof is KO. The body below consolidates the 27 September maintainer request and supersedes outdated implementation assumptions in earlier comments; measured results remain evidence.
To do
Publish an attributable baseline. The measurer records the current full frame first, then CPU shadow planning/encoding, GPU cull/raster/resolve, passes, draws, static/dynamic work and memory. Check timestamp boundaries: the broad Shadows stage has included unrelated work in earlier reports. Never sum overlapping GPU timers or CPU and GPU durations to manufacture a frame cost. Unsupported values remain null.
Select only required shadow work — owned by this programme. The next measured selection step must use the existing receiver requests, cluster hierarchy and per-mover bounds, not introduce another BVH. Cover the complete sampling footprint (including filter taps, transparent receivers and other actual shadow consumers). A caster outside the camera can still shadow a visible receiver and must remain. Fine receiver masks are a candidate only when the baseline shows excess caster/texel work. Partial dynamic coverage must never be marked as a fully valid cached page. Split this bounded step into a child when it is picked, before coding.
Remove feedback-dependent CPU scheduling and repeated submission — Moving-light shadows stay current: GPU scheduling and submission within the frame budget #344. Separate GPU page scheduling from raster submission; stage the migration behind the same page identity, residency and validity contracts. Same-frame shadow demand must be available before raster and final sampling. Post-shading readback cannot schedule that same frame. Resource creation and the global memory grant stay host responsibilities; per-page hot-path work moves only when the measured gain justifies it.
Material surfaces, a steady 11.4–12.0 ms, with texture fetch included: remove repeated shading and fetches, and values that can be prepared once or cooked. Reuse the visibility/material pipeline; do not add another material renderer. perf(geometry): skip material classification for prepared single-class scenes #902 already skips material depth on one-class frames.
Spikes:shadow atlas and light tiles reach 10–11 ms at p95 and drive the envelope p95 to 46–53 ms.
Pass traffic: on the maintainer's Apple GPU, compare attachment traffic and total cost when compatible work is grouped. Count both the intermediates added and the passes removed. Shader micro-optimisations and fewer draw calls are not evidence of a full-frame win.
Order: baseline → measure the closed #867 and repair #850 → #831 capacity/cache steps → measured fine selection → #344 scheduling/submission → conditional #22. The owner splits each programme step into a single-PR child at pick time (about 500 handwritten lines maximum); no coder takes this parent as one rewrite. Existing image defects are not postponed behind speculative optimisations.
A small triangle count can still cause excessive per-page submission
Same astrolabe stopped through its public control
Work settles after real changes finish
A moving scene is not expected to have the cache ratio of a still one
One box moving over a static floor
Static depth remains reusable; old and new dynamic coverage are correct
Static/dynamic separation, including removal of the previous shadow
Wall outside camera, shadow on visible floor
Preserve the wall's contribution
Receiver-driven selection is not camera-frustum culling of casters
Camera turn/resize and a lamp moving every frame
Newly needed shadows are current, without holes or delayed silhouettes
Feedback latency, pool lifetime, page validity
Code context
Anchors checked against local develop454fc23b3; re-resolve lines at implementation time.
packages/sdk-browser/src/lighting/direct/shadowWgsl.ts:59 — existing deduplicated receiver requests and filter-aware page lookup.
packages/sdk-browser/src/webgpu/shadow/pageRequests.ts:28 — asynchronous feedback; :102 maps readback after submission. It is not a synchronous CPU wait today, but CPU consumption still affects scheduling.
packages/sdk-core/src/scene/light-shadow/plan.ts:30 — existing planner, allocation/admission and validity rules; preserve their invariants.
packages/sdk-browser/src/webgpu/pages/render/encodeShadows.ts:61 — host region planning; encodeShadowPass.ts:70 — raster and per-region viewports/draws.
packages/sdk-browser/src/gpu/shadow/cullShader.ts:59, gpu/shadow/staticLayer.ts:25, webgpu/shadow/mobility.ts:16 — existing selection and static/dynamic infrastructure, reused.
packages/sdk-browser/src/webgpu/pages/render/cpuSteps.ts:1, packages/sdk-core/src/contracts/shadowMetrics.ts:1 — existing observation contracts, not a second profiler.
Measurer only, one browser on the machine. Pin before/after commits, machine/GPU, browser/backend, display Hz, CSS canvas size, actual drawing-buffer size, DPR, camera, animation input, quality, memory grant and warm-up. Use native resolution; the requested reference case is 1728×1117 CSS, DPR 2, and the astrolabe's actual canvas must also be recorded rather than inferred from the screenshot.
Three comparable runs per side, with run-to-run spread. Publish CPU and GPU frame envelopes separately, rAF intervals p50/p95/p99/max, missed refreshes (interval > 1.5 × display period), long tasks, cold-start/resize events separately from steady animation. A 60 Hz headless run cannot prove 120 Hz presentation. Do not use the known unlit uncapped path as a performance proof.
Programme targets: GPU frame p95 ≤ 8 ms, shadow contribution ≤ 2 ms with a documented attribution method, main-thread engine work ≤ 2 ms as in Streaming without holes: rules and objectives for geometry, memory and shadows #483. At an actual 120 Hz display the frame budget is 8.33 ms; publish missed refreshes and p99, not just average FPS. No engine long task >16 ms during the reference navigation. Targets are requirements, not achieved results.
Image: unchanged resolution, lighting, shadow distance and filtering; deterministic animation checkpoints against the same pose, A/A baseline, no disappearing/stale shadow or temporal trail. Apply CONTRIBUTING's image rule, never accept an unlit frame as a fast frame. No quality reduction to meet the timing.
Every implementation child has behaviour tests and check:changed, test:changed, validate; the named browser proof is run in the measurer queue after merge. This parent closes only when its children and the end-to-end targets are measured as met, not merely when implementation issues close.
Published techniques: comparison and applicability
Epic: Virtual Shadow Maps describes fine page selection and separate static caching. These justify removing redundant work; they do not establish a speedup for our implementation.
Epic support, UE 5.7 receiver masks explains why partial dynamic pages may be uncached while static pages remain cached. Its directional-light result is not proof for the astrolabe's point light.
Unity HDRP mixed cached shadows combines cached static geometry and dynamic rendering, with extra atlas memory and copy cost. Measure both sides of that trade-off here.
Blender EEVEE shadow module documents GPU allocation/update without host synchronization. Its shadow pipeline describes direct atlas writes using atomic minima. These are architectural references, not evidence that Blender's backend facilities or exact depth behaviour are portable to WebGPU.
Apple: load/store actions explains attachment traffic between passes. Measure whole-frame cost and pass layout on the maintainer's Apple GPU, not draw count alone.
WebGPU specification defines the available commands and limits. Record actual browser/device features; no mandatory hardware ray tracing, multi-viewport extension or unverified native-only facility.
Reimplement from public descriptions only; never copy engine source. Dynamic ray-traced lighting is not a prerequisite for this one-lamp regression. The techniques above are candidates constrained by correctness, measured cost and the web backend.
Why
The maintainer reports stuttering and the loss of 120 FPS on ordinary examples. His 27 September 2026 screenshot of
spin-an-astrolabeshows 35 FPS, 21.43 ms GPU, 38,432 triangles, 1,184 draw calls including 1,172 shadow draws, 586 shadow pages drawn, 580 requested, 0 cached and 0 refetched. This is an observation, not a controlled benchmark: screenshot commit, canvas size, browser and sampling interval are unknown. It strongly motivates measuring shadow submission/raster work; it does not prove that all invalidation is unnecessary or that the cache is thrashing.This is the coordination and acceptance issue for the shadow programme, not one implementation PR. Earlier improvements (#763, #766, #811, #818, #815) do not establish the final performance target. #868 was rejected; its single-depth-draw design is not delivered. The latest #850 proof is KO. The body below consolidates the 27 September maintainer request and supersedes outdated implementation assumptions in earlier comments; measured results remain evidence.
To do
Shadowsstage has included unrelated work in earlier reports. Never sum overlapping GPU timers or CPU and GPU durations to manufacture a frame cost. Unsupported values remainnull.to measure) owns only the exact transmittance clear; its measured gain is still pending. Shadow pool resizing preserves the image, frame cost and peak memory budget #850 owns runtime pool resizing, including its new image, pass-count and transient-memory failures. Neither ticket may claim that its local change delivers the whole 120 FPS target.shadow atlasandlight tilesreach 10–11 ms at p95 and drive the envelope p95 to 46–53 ms.Order: baseline → measure the closed #867 and repair #850 → #831 capacity/cache steps → measured fine selection → #344 scheduling/submission → conditional #22. The owner splits each programme step into a single-PR child at pick time (about 500 handwritten lines maximum); no coder takes this parent as one rewrite. Existing image defects are not postponed behind speculative optimisations.
Examples and design rationale
Code context
Anchors checked against local
develop454fc23b3; re-resolve lines at implementation time.packages/sdk-browser/src/lighting/direct/shadowWgsl.ts:59— existing deduplicated receiver requests and filter-aware page lookup.packages/sdk-browser/src/webgpu/shadow/pageRequests.ts:28— asynchronous feedback;:102maps readback after submission. It is not a synchronous CPU wait today, but CPU consumption still affects scheduling.packages/sdk-core/src/scene/light-shadow/plan.ts:30— existing planner, allocation/admission and validity rules; preserve their invariants.packages/sdk-browser/src/webgpu/pages/render/encodeShadows.ts:61— host region planning;encodeShadowPass.ts:70— raster and per-region viewports/draws.packages/sdk-browser/src/gpu/shadow/cullShader.ts:59,gpu/shadow/staticLayer.ts:25,webgpu/shadow/mobility.ts:16— existing selection and static/dynamic infrastructure, reused.packages/sdk-browser/src/webgpu/pages/render/cpuSteps.ts:1,packages/sdk-core/src/contracts/shadowMetrics.ts:1— existing observation contracts, not a second profiler.packages/sdk-browser/src/webgpu/core/materialPasses.ts— material pass encoding, for the material-surface item moved from Material performance: redundant pixel work removed and every GPU pass of the frame attributed #685.Proof
check:changed,test:changed,validate; the named browser proof is run in the measurer queue after merge. This parent closes only when its children and the end-to-end targets are measured as met, not merely when implementation issues close.Links
Parent #483. Implementation/work ownership: #867, #850, #831, #344, #22. Related #685 (narrowed on 29 September; its frame target lives here), #26, #234, #816. Delivered historical work: #811, #818, #815, #902, #1128; rejected experiment: #868. Latest #850 KO.
Published techniques: comparison and applicability
Reimplement from public descriptions only; never copy engine source. Dynamic ray-traced lighting is not a prerequisite for this one-lamp regression. The techniques above are candidates constrained by correctness, measured cost and the web backend.