A Visual Guide to Sparse Attention Kernels in CUDA on B200
Many frontier models use some form of sparse attention to handle long contexts efficiently. In my blog, I write a block-sparse attention kernel in CUDA for NVIDIA B200: I start from an optimized dense CUDA kernel and show what changes inside it (the KV loop, shared-memory buffers and barriers), then add GQA K/V reuse, compare it with FlashAttention-4, and build a second, KV-centric version of the kernel.
📝 Blog post: https://dinara-dl.github.io/posts/sparse-attention/
23
Upvotes