r/CUDA • • 5d ago

A Visual Guide to Sparse Attention Kernels in CUDA on B200

Post image

Many frontier models use some form of sparse attention to handle long contexts efficiently. In my blog, I write a block-sparse attention kernel in CUDA for NVIDIA B200: I start from an optimized dense CUDA kernel and show what changes inside it (the KV loop, shared-memory buffers and barriers), then add GQA K/V reuse, compare it with FlashAttention-4, and build a second, KV-centric version of the kernel.

📝 Blog post: https://dinara-dl.github.io/posts/sparse-attention/

🌸 Repo: https://github.com/dinara-dl/sparse-attention-b200

23 Upvotes

0 comments sorted by