Skip to content

Add support for candle-flash-attn and candle-flash-attn-v3 in CUDA - #14

Merged
alvarobartt merged 6 commits into
add-missing-laya-variantsfrom
add-flash-attention
Sep 25, 2026
Merged

alvarobartt merged 6 commits into
add-missing-laya-variantsfrom
add-flash-attention

Conversation

@alvarobartt

Copy link
Copy Markdown
Owner

Description

This PR adds support for both candle-flash-attn and candle-flash-attn-v3 in CUDA, but only for Ampere, Ada Lovelace and Hopper, meaning that both Turing and Blackwell as left aside. On Turing because the kernels require features such as cp.async or bfloat16 tensor cores, which are only in Ampere onward; and then for Blackwell the limitation is on candle kernels which are missing coverage for sm100 and sm120 for the time being, so not discarding that support might come soon!

@alvarobartt
alvarobartt added this pull request to stack #15 September 25, 2026 07:48
@alvarobartt
alvarobartt merged commit 8921718 into main Sep 25, 2026
16 checks passed
@alvarobartt
alvarobartt deleted the add-flash-attention branch September 25, 2026 13:14

This branch was successfully deployed

1 active deployment
pr-build — 2e821f38 Deployed Sep 25, 2026 by alvarobartt via Approve PR build #7
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant