vk_native: forget a destroyed window-space image's cached HUD view (#1782) - #1783
Conversation
dfattal
left a comment
There was a problem hiding this comment.
Approve with nits. Correct for the targeted bug. The dead image is queued before vkDestroyImage, so there's no reused-handle (ABA) window, and every lookup goes through the one drain. ws_dead.mutex is a true leaf lock and never waits on c->mutex. Teardown order is fine (vkDeviceWaitIdle → vk_hud_blend_fini). No new logs. It merges cleanly with main after #1784; I'm updating the branch so CI runs on the combined tree.
1 (low): descriptors can be freed while in flight on the split (weave-on-scanout) path. The new comment says the previous frame's submit has always been waited on before the window-space pass records. That holds for the windowed path (fence), shared-texture (vkQueueWaitIdle) and DP self-submit. It does not hold for vk_split_composite_window_space: its 3-deep VK_SPLIT_PLANE_CMD_RING only waits on the slot it reuses, so up to 2 earlier frames may still be running. If more than 64 images die between passes, the overflow branch frees the descriptor sets of every cached entry, including live HUD images bound by an in-flight command buffer (VUID-vkFreeDescriptorSets-pDescriptorSets-00309). Fix: wait for the outstanding ws_cmd_value[] before forgetting on the split path, or at least correct the comment. The neighbouring atlas-framebuffer eviction comment makes the same claim.
2 (nit): the overflow branch can spin forever. while (image_count > 0) forget(cached_images[0].image) never ends if entry 0's image is VK_NULL_HANDLE, because forget returns early on null. It's safe today (the only caller null-checks src_image), but a vk_hud_blend_forget_all() would be sturdier and would keep the compositor out of cached_images.
3 (nit): doc comment in the wrong place. The new function's doxygen block sits between vk_compositor_render_window_space_into_atlas's doc block and that function, so the big block now documents the wrong function. Move the new function above it.
Older issue, not introduced here: vk_swapchain_destroy calls vkDestroyImage with no GPU wait. comp_multi_system's blends still never evict, as the PR body notes.
Items 1–3 can be a follow-up. Note that #1787 (#1786) adds a second pipeline inside the same vk_hud_blend so your cache eviction still covers it; it rebases cleanly on top of this.
…isplayXR#1782) vk_hud_blend caches an image view + descriptor set per HUD source image, keyed by the raw VkImage handle, and never evicted. The driver reuses a destroyed image's handle for the next image it creates, so an app that resizes a window-space layer (destroy + create its swapchain) got the old entry back: the window-space pass sampled a view of freed memory and the GPU hung within 2-3 resizes (VK_ERROR_DEVICE_LOST, NVIDIA Xid 109). - vk_hud_blend_forget_image() drops one entry (the pool can now free sets). - vk_native swapchain destroy queues its images on the compositor; the window-space pass forgets them before it looks anything up. The queue has its own leaf lock: destroy runs on the app thread, the pass can run on the weave thread, and c->mutex may be held by a wedged weave thread (DisplayXR#1394). comp_multi_system's hud/chrome blends keep the old behaviour. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
8da6cca to
869b5f3
Compare
|
@byungjul thanks. #1784 and this PR are both merged. This one was rebased onto #1784 first so CI ran on the combined tree. Here's what followed:
Your Linux hardware setup is the only place the combined #1784 + #1783 + #1787 behaviour has been run with transparency on. A quick re-check of the HUD panels on current |
Fixes #1782.
Problem
On the Vulkan native compositor, resizing a window-space layer (the app destroys its swapchain and creates a new one) hung the GPU within 2–3 resizes:
VK_ERROR_DEVICE_LOSTindraw_zones_pass, NVIDIA Xid 109, and every frame after that failed.vk_hud_blendcaches an image view + descriptor set per HUD source image, keyed by the rawVkImagehandle, and never removes an entry. The driver reuses a destroyed image's handle for the next image it creates, so the new swapchain image got the old entry: a view of the destroyed (and differently sized) image. Details and handle logs are in #1782.Fix
vk_hud_blend_forget_image()removes one cache entry: it destroys the view and frees the descriptor set. The descriptor pool now hasFREE_DESCRIPTOR_SET_BIT.comp_vk_native_compositor_swapchain_images_destroyed). The window-space pass forgets them at its start, before it looks anything up. That is safe for the same reason the existing atlas-framebuffer eviction there is: the previous frame's submit is waited on before this recording begins.c->mutexmay be held by a wedged weave thread (Android in-process: after an Adreno GSL timestamp violation on a target recreate, the next weave hangs forever in the vendor interlacer's vkWaitForFences (mini-window freeze) #1394). If more images die than the queue holds, the pass clears the whole cache instead.comp_multi_system.c(hud_blend,chrome_blend,shared_chrome_blend) uses the same cache and is not changed here. Its source images may need the same treatment; I haven't looked at whether they are ever recreated.Testing
Ubuntu 26.04, RTX 4090 Laptop, Leia DS1, Leia SR plug-in 2.7.6, transparent session. App: the lenovo-avatar Unity player with displayxr-unity window-space UI on Vulkan. A dev-only harness disables and re-enables each HUD with a different resolution, which destroys and recreates its swapchain.
xrEndFrameerror, no Xid, both HUDs drawnThe local build has to use Vulkan 1.3.204 headers (Ubuntu 22.04's) to match the Leia plug-in's
vk_bundleABI. With newer headers the plug-in is refused and the session runs unwoven, and then the hang doesn't reproduce.The branch is on current
mainand builds there (build_linux.sh, Release, no new warnings). The hardware test ran on v2.21.11 with the same patch, which applies cleanly.Not fixed here: the SIGSEGV on quit (#1779) still happens with and without this change.
🤖 Generated with Claude Code