Skip to content

vk_native: forget a destroyed window-space image's cached HUD view (#1782) - #1783

Merged
dfattal merged 1 commit into
DisplayXR:mainfrom
byungjul:fix/vk-hud-blend-stale-cache
Oct 2, 2026
Merged

dfattal merged 1 commit into
DisplayXR:mainfrom
byungjul:fix/vk-hud-blend-stale-cache

Conversation

@byungjul

@byungjul byungjul commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Fixes #1782.

Problem

On the Vulkan native compositor, resizing a window-space layer (the app destroys its swapchain and creates a new one) hung the GPU within 2–3 resizes: VK_ERROR_DEVICE_LOST in draw_zones_pass, NVIDIA Xid 109, and every frame after that failed.

vk_hud_blend caches an image view + descriptor set per HUD source image, keyed by the raw VkImage handle, and never removes an entry. The driver reuses a destroyed image's handle for the next image it creates, so the new swapchain image got the old entry: a view of the destroyed (and differently sized) image. Details and handle logs are in #1782.

Fix

  • vk_hud_blend_forget_image() removes one cache entry: it destroys the view and frees the descriptor set. The descriptor pool now has FREE_DESCRIPTOR_SET_BIT.
  • vk_native swapchain destroy tells the compositor which images are going away (comp_vk_native_compositor_swapchain_images_destroyed). The window-space pass forgets them at its start, before it looks anything up. That is safe for the same reason the existing atlas-framebuffer eviction there is: the previous frame's submit is waited on before this recording begins.
  • The queue has its own leaf lock. Destroy runs on the app thread, the pass can run on the weave thread, and c->mutex may be held by a wedged weave thread (Android in-process: after an Adreno GSL timestamp violation on a target recreate, the next weave hangs forever in the vendor interlacer's vkWaitForFences (mini-window freeze) #1394). If more images die than the queue holds, the pass clears the whole cache instead.

comp_multi_system.c (hud_blend, chrome_blend, shared_chrome_blend) uses the same cache and is not changed here. Its source images may need the same treatment; I haven't looked at whether they are ever recreated.

Testing

Ubuntu 26.04, RTX 4090 Laptop, Leia DS1, Leia SR plug-in 2.7.6, transparent session. App: the lenovo-avatar Unity player with displayxr-unity window-space UI on Vulkan. A dev-only harness disables and re-enables each HUD with a different resolution, which destroys and recreates its swapchain.

Runtime Result
v2.21.11 (installed .deb and a local build) GPU hang within 2–3 resizes, 6 of 6 runs
v2.21.11 + this change 2 runs × 40 resizes (two HUDs, every 2 s): no xrEndFrame error, no Xid, both HUDs drawn

The local build has to use Vulkan 1.3.204 headers (Ubuntu 22.04's) to match the Leia plug-in's vk_bundle ABI. With newer headers the plug-in is refused and the session runs unwoven, and then the hang doesn't reproduce.

The branch is on current main and builds there (build_linux.sh, Release, no new warnings). The hardware test ran on v2.21.11 with the same patch, which applies cleanly.

Not fixed here: the SIGSEGV on quit (#1779) still happens with and without this change.

🤖 Generated with Claude Code

@dfattal dfattal left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approve with nits. Correct for the targeted bug. The dead image is queued before vkDestroyImage, so there's no reused-handle (ABA) window, and every lookup goes through the one drain. ws_dead.mutex is a true leaf lock and never waits on c->mutex. Teardown order is fine (vkDeviceWaitIdle → vk_hud_blend_fini). No new logs. It merges cleanly with main after #1784; I'm updating the branch so CI runs on the combined tree.

1 (low): descriptors can be freed while in flight on the split (weave-on-scanout) path. The new comment says the previous frame's submit has always been waited on before the window-space pass records. That holds for the windowed path (fence), shared-texture (vkQueueWaitIdle) and DP self-submit. It does not hold for vk_split_composite_window_space: its 3-deep VK_SPLIT_PLANE_CMD_RING only waits on the slot it reuses, so up to 2 earlier frames may still be running. If more than 64 images die between passes, the overflow branch frees the descriptor sets of every cached entry, including live HUD images bound by an in-flight command buffer (VUID-vkFreeDescriptorSets-pDescriptorSets-00309). Fix: wait for the outstanding ws_cmd_value[] before forgetting on the split path, or at least correct the comment. The neighbouring atlas-framebuffer eviction comment makes the same claim.

2 (nit): the overflow branch can spin forever. while (image_count > 0) forget(cached_images[0].image) never ends if entry 0's image is VK_NULL_HANDLE, because forget returns early on null. It's safe today (the only caller null-checks src_image), but a vk_hud_blend_forget_all() would be sturdier and would keep the compositor out of cached_images.

3 (nit): doc comment in the wrong place. The new function's doxygen block sits between vk_compositor_render_window_space_into_atlas's doc block and that function, so the big block now documents the wrong function. Move the new function above it.

Older issue, not introduced here: vk_swapchain_destroy calls vkDestroyImage with no GPU wait. comp_multi_system's blends still never evict, as the PR body notes.

Items 1–3 can be a follow-up. Note that #1787 (#1786) adds a second pipeline inside the same vk_hud_blend so your cache eviction still covers it; it rebases cleanly on top of this.

…isplayXR#1782)

vk_hud_blend caches an image view + descriptor set per HUD source image,
keyed by the raw VkImage handle, and never evicted. The driver reuses a
destroyed image's handle for the next image it creates, so an app that
resizes a window-space layer (destroy + create its swapchain) got the old
entry back: the window-space pass sampled a view of freed memory and the
GPU hung within 2-3 resizes (VK_ERROR_DEVICE_LOST, NVIDIA Xid 109).

- vk_hud_blend_forget_image() drops one entry (the pool can now free sets).
- vk_native swapchain destroy queues its images on the compositor; the
  window-space pass forgets them before it looks anything up. The queue has
  its own leaf lock: destroy runs on the app thread, the pass can run on the
  weave thread, and c->mutex may be held by a wedged weave thread (DisplayXR#1394).

comp_multi_system's hud/chrome blends keep the old behaviour.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dfattal
dfattal force-pushed the fix/vk-hud-blend-stale-cache branch from 8da6cca to 869b5f3 Compare October 2, 2026 07:15
@dfattal
dfattal merged commit de3f621 into DisplayXR:main Oct 2, 2026
34 checks passed
@dfattal

dfattal commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

@byungjul thanks. #1784 and this PR are both merged. This one was rebased onto #1784 first so CI ran on the combined tree. Here's what followed:

Your Linux hardware setup is the only place the combined #1784 + #1783 + #1787 behaviour has been run with transparency on. A quick re-check of the HUD panels on current main would be welcome.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

vk_native: resizing a window-space layer hangs the GPU (vk_hud_blend caches views by VkImage handle; handles get reused)

2 participants