Skip to content

Add WebGPU wedge watchdog with one-way WASM fallback - #35

Merged
felirami merged 1 commit into
mainfrom
fable/webgpu-watchdog
Aug 16, 2026
Merged

felirami merged 1 commit into
mainfrom
fable/webgpu-watchdog

Conversation

@felirami

Copy link
Copy Markdown
Collaborator

Defect

Live smoke of v1.3.1 on an M3 Pro under GPU contention: the offscreen document's WebGPU inference stalled mid-page (Wikipedia article, Amazon search), the offscreen init health check failed after 30 seconds, and 104 badges stayed pending with no recovery. AGENTS.md covers init ("probe the adapter; do not latch a WebGPU error") but there was no runtime recovery once a WebGPU session.run hung or the device was lost: the exclusive inference lock stayed occupied by the wedged run and the whole queue starved silently.

Fix

New pure module src/webgpu-watchdog.js, wired into src/offscreen.js:

  1. Per-inference watchdog. Every session.run on a WebGPU session (CF and DINO, via wrapSessionWithWatchdog) races MODEL_RUN_WEDGE_TIMEOUT_MS (20 s, JSDoc-documented: generous vs sub-5 s normal runs, well under the 45 s INFERENCE_TIMEOUT_MS so the wedge verdict plus WASM rebuild plus retry fit in one inference budget). A run past the timer is treated as a wedged backend.
  2. Device loss. GPUDevice.lost (via ort.env.webgpu.device) marks a one-way health latch; device-level errors thrown by a run (isWebGpuDeviceError) count the same way. Fetch, decode, and size-cap errors deliberately never match.
  3. Automatic fallback. On wedge or device loss: both sessions are disposed (fire-and-forget, since release on a wedged device can itself hang), both are rebuilt on the WASM execution provider, the switch is logged once (console.warn "Clueside: WebGPU backend unhealthy (...)"), the in-flight image is rescored on WASM, and everything waiting behind the inference lock or in the page queue runs on the rebuilt WASM sessions. The latch is one-way for the browser session: no flap back to WebGPU until the offscreen document restarts.
  4. No silent starvation. An error after the WASM retry propagates to the analyze response, so the content script paints its normal error badge instead of leaving the badge pending. The DINO head also stops swallowing wedge/device errors (it previously returned null on any error), so a wedged DINO run triggers the same fallback instead of silently degrading.

Not touched

Fusion policy, thresholds, bands, models, probe, benchmark numbers.

Evidence

  • npm test: 248 passing, 0 failing (was 228; 20 new tests in tests/webgpu-watchdog.test.mjs covering the wedge timer, error classification, fallback decision, one-way latch, and offscreen wiring).
  • esbuild src/offscreen.js --bundle compiles clean (695 kb, same entry point the build uses).

Live smoke of v1.3.1 (M3 Pro under GPU contention): a WebGPU session.run
stalled mid-page on Wikipedia and Amazon search, the offscreen document
failed its 30 second init health check, and 104 badges stayed pending
with no recovery path.

src/webgpu-watchdog.js (new, pure, node:test-covered):
- MODEL_RUN_WEDGE_TIMEOUT_MS 20s per single session.run, well under the
  45s inference budget so the wedge verdict plus WASM retry fit inside it
- isWedgeError / isWebGpuDeviceError / shouldFallBackToWasm decision
  logic, watchModelRun race, wrapSessionWithWatchdog session wrapper,
  createWebGpuHealth one-way latch

src/offscreen.js:
- WebGPU CF and DINO sessions are wrapped with the watchdog
- GPUDevice.lost marks the health latch
- On wedge or device loss: dispose both sessions, rebuild both on WASM,
  log the switch once, rescore the in-flight image, and let everything
  behind the inference lock and in the page queue run on WASM
- One-way for the browser session; no flap back until the offscreen
  document restarts
- An error after the WASM retry propagates so the badge gets its normal
  error state instead of staying pending

tests/webgpu-watchdog.test.mjs: 20 new tests. Full suite 248 passing
(was 228). No fusion, threshold, model, or probe changes.
@felirami
felirami merged commit 7347f79 into main Aug 16, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant