Skip to content

q4k mmq optimizations - #80

Draft
liangliangchang wants to merge 3 commits into
gfx11from
lichang.q4k-opt
Draft

q4k mmq optimizations#80
liangliangchang wants to merge 3 commits into
gfx11from
lichang.q4k-opt

Conversation

@liangliangchang

Copy link
Copy Markdown

Overview

Additional information

Requirements

liangliangchang and others added 3 commits July 31, 2026 13:18
Pipeline Q4_K tile loads with split WMMA dequantization and select the measured tile width by matrix shape.

Assisted-by: GPT-5.6 Sol
Co-authored-by: Cursor <cursoragent@cursor.com>
Match the refactored 32-row warp layout and limit the specialized J64 path to the small shape where it provides a stable gain.

Assisted-by: GPT-5.6 Sol
Co-authored-by: Cursor <cursoragent@cursor.com>
Use the generic vec-dot with J64 two-row waves and J128 one-row waves, retaining the J64 prefetch loop for the best measured shape coverage.

Assisted-by: GPT-5.6 Sol
Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant