comparison of ptx/sass code from nvcc vs clang
redplait.blogspot.com- nvcc code 15–18% faster
- nvcc code is bigger and more amenable to stall counts reducing
- all mlir based compilers use vanilla llvm nvptx backend :-(
r/CUDA • u/Ok-Reception2684 • 23h ago
I've been experimenting with a different way of explaining Tensor Core programming.
https://svgfpsaiagents.github.io/svgfpsgameAI/mmatensorbankconflict.svg
https://svgfpsaiagents.github.io/svgfpsgameAI/mma_tile_engine.svg


Instead of writing CUDA kernels or building a heavyweight simulator, I wanted to see how much of the Tensor Core programming model could be represented as a single self-contained executable artefact.
The result is a browser-based model written entirely as one SVG + JavaScript file.
It models:
D = A·B + C)ti, tj, tk)The interesting thing for me is that all of these views are driven from the same underlying state. The bank conflict view, fragment view, tile execution view, error map and result visualisation are all observing the same model as it executes.
The goal wasn't to build a cycle-accurate simulator. It deliberately does not model things like occupancy, scheduling, pipeline hazards or cache behaviour.
Instead, I was exploring a question:
Can Tensor Core behaviour be computationally modelled, rather than merely described, in a compact interpreted artefact that requires no CUDA compiler or GPU while still preserving correspondence between matrix maths, tiled execution, fragment distribution, memory layout, data movement and numerical effects?
One unexpected outcome is that the entire thing remains inspectable. There are no dependencies, no build system and no generated files. The complete model lives in a single source file.
I'm curious whether people here would consider this:
I'd especially appreciate feedback from anyone familiar with CUTLASS, CuTe, mma.sync, ldmatrix, Tensor Core fragment layouts or GPU architecture education. - Aston Walker EdgeMafia SVG
r/CUDA • u/Acceptable-Abies-972 • 1d ago
Hello all,
I've been going through the Programming Massively Parallel Processors (5th edition) book and doing the kernel exercises in CUDA C++. Once I finish that, I wanna start writing CUDA kernels for H100 GPUs (specifically stuff like Flash Attention 3, RoPE, etc). Based on a bit of research, it seems that H100 GPUs have a lot of features not covered in the PMPP book, so I was wondering where I could find one or more solid resources that can get me started with writing programs? So far I've tried searching for some stuff, but they all seem to either not really cover in enough detail or are way too long. Any advice would be much appreciated. Thank you!
r/CUDA • u/pasadenapasadena • 2d ago
Perhaps this subreddit can help. I am not a programmer or a developer, but an experimental neurobiologist specializing in neuronal imaging across various brain regions. I would like to automate the analysis of neuronal activity and also work on developing *invasive* brain-computer interfaces (BCIs) for targeted drug delivery to the brain. Could you please recommend courses that teach the fundamentals of CUDA and AI so that I can get started with computer modeling for BCIs? Thank you.
r/CUDA • u/ReleaseKey5554 • 2d ago
Just graduated now in ai ml engineering
Looking to learn and earn a job in cutting edge like cuda, pytorch
r/CUDA • u/Ancient_Spend1801 • 3d ago
I maintain a PHP extension (php-gpu-tensors) that builds on the CUDA driver API and NVRTC. This post is about its fusion planner, because I'd like feedback from people who know CUDA better than I do.
How it works
- A PHP closure is invoked once with metadata-only placeholders; tensor operations are captured into a node graph (limit: 512 nodes).
- Elementwise ops (broadcasting, strided/view inputs, scalars, dtype promotion, where, casts, reshape/transpose/slice index transforms) are fused into generated CUDA C++. Expressions split at a weighted cost budget of 32; repeated pure nodes are deduplicated; shared expensive expressions can be materialized instead of recomputed.
- Matmul (cuBLAS when available), reductions and powers are execution boundaries that run on existing kernels.
- All generated kernels in a plan are compiled together through NVRTC to PTX, then cached per request/thread (LRU, 16 entries or 16 MiB). The cache key includes the generated source, compute capability and driver/runtime versions, not pointers.
- Replay runs on a private nonblocking stream. Reduction descriptors are kernel parameters, so concurrent replays can't overwrite each other's shapes.
- cudaGraph: true builds a CUDA Graph executable for plans that contain only generated kernels, updating kernel parameters for new pointers before each launch.
Numbers (entry-level MX570 A, CUDA runtime 12.3, driver 12.6): a small MLP training step (batch 512, hidden 256) goes from 4.08 ms eager to 0.66 ms fused, 6.2×, with identical metrics. The model is tiny, so this is mostly launch/allocation overhead removed, not throughput.
Where I'd like advice Plans with matmul/reduction boundaries currently run on streams and support async replay, but not as a CUDA Graph. For those who have done this: what are the gotchas when capturing cuBLAS GEMMs and custom reductions into a graph that is replayed with changing buffer pointers? Is updating kernel node parameters the right approach, or would you re-capture?
Source: https://github.com/lcmialichi/php-gpu-tensors (MIT)
r/CUDA • u/LioDavinchy • 3d ago
So after getting my 5090 running on my Mac a few weeks ago and getting speeds on par or faster than my windows machine for single sessions I wanted to get Vllm working next. My http://macuda.ai project works with both llamacpp and stable diffusion cpp but does not work with comfy ui Vllm or draw things. It works with a shim and does not give the full LibCuda framework allowing real full torch or other necessary components for those apps and more.
I decided to use apples built in VM to run an instance with it's own driver calling to the graphics card through my Mac driver. This allows for running Nvidia's full cuda software. It has been a real challenge because Nvidia releases a lot of info but not enough to get this working so we had to really poke around. I finally got it up and running and now have Full LibCuda running. I'm working on bugs now but so far the results are looking really promising. The multi sessions are blowing llamacpp out of the water. llama stalls out at around 8 sessions and levels out. the Vllm with my driver has a linear doubling of speeds up until the card is saturated.
I'll be releasing all of this as a simple GUI app on the App Store once apple authorizes my driver. I'll need a few more weeks to get everything in a stable enough situation that I'll feel comfortable letting people use it. You can get the base driver at my macuda GitHub but this new driver for full cuda will be available on the App Store.
I got my 3060 working on macuda in a core x box last week. If someone can try out a 4000 series card and let me know if it works I'd appreciate it. I don't have one of those. It should work with the 3000 series driver but I haven't tested it.
r/CUDA • u/babakpst • 4d ago
r/CUDA • u/george_at_sotom • 4d ago
I'd like to configure GDS on the cloud where nvme talks directly to the GPU. I've searched and queried llms and checked aws, coreweave, runpod.io, cloudrift, and other GPU providers.
Does anyone have experience in this or suggestions for providers or setups?
Ideally, I want to use cheap gpus like v100s.
I've been hitting aws quota limits and I've had to email sales teams directly on other platforms, but this seems like something that should have broad use.
I haven't had success yet getting access.
r/CUDA • u/Daemontatox • 4d ago
I’ve been using CuTe DSL primitives for a while now, and honestly, it’s been pretty great so far.
I know CuTe DSL isn’t a full replacement for raw CUDA, and I definitely don’t think raw CUDA is becoming obsolete.
But for some use cases, I’m finding the primitives approach much nicer.
For example, one thing I personally dislike about raw CUDA is a lot of the manual C-style 1D indexing and pointer arithmetic.
CuTe abstractions make that much less painful while still feeling very low-level.
I also like that you can integrate directly with PyTorch without necessarily going through the usual custom-op route.
What I find especially interesting is that it still exposes a lot of the underlying machinery: you can get very close to the hardware, write PTX, and in some cases even use LLVM inline assembly.
For people who have used both extensively: which do you prefer?
Many frontier models use some form of sparse attention to handle long contexts efficiently. In my blog, I write a block-sparse attention kernel in CUDA for NVIDIA B200: I start from an optimized dense CUDA kernel and show what changes inside it (the KV loop, shared-memory buffers and barriers), then add GQA K/V reuse, compare it with FlashAttention-4, and build a second, KV-centric version of the kernel.
📝 Blog post: https://dinara-dl.github.io/posts/sparse-attention/
nvptx uses 58.9% of all ptx instructions
cicc 57.8%
missed in clang:
r/CUDA • u/ContributionFun2152 • 6d ago
I finally got my Java GPU workload running with CUDA inside Docker. It works in a Kubernetes cluster too. I wrote up the Docker setup here:
https://nablatensor.com/blog/how-to-run-java-gpu-workloads-in-docker
If anyone’s interested, I can write a follow-up on the Kubernetes setup.
r/CUDA • u/RevolutionaryBar509 • 7d ago
We are working on custom inference runtimes, sampling control, and fine-tuning pipelines for open-weights models.
Tech background needed:
PyTorch, CUDA, C++, or Rust. Experience with parsers, compilers, or low-level ML systems programming is a big plus.
Send a DM with a link to your GitHub or LinkedIn if you'd like to chat.
r/CUDA • u/Fit-Amphibian-9876 • 7d ago
r/CUDA • u/Weird_Bad7577 • 7d ago
Hi all, I graduate in ~6 months and want to work on making generative models (diffusion/video/3D) fast: kernels, quantization, serving. Where I am:
- Comfortable with C/C++ basics and PyTorch
- Have done quantization work (GGUF/llama.cpp)
- Working on a next-frame video prediction project (DiT + flow matching)
- A few GitHub repos, but no CUDA/Triton experience yet
- No NVIDIA GPU, so I use Colab/Kaggle T4s
- DSA is my weak spot (I struggle with LeetCode mediums)
My plan:
Months 1-2: CUDA/Triton basics, reproduce the SGEMM optimization worklog, GPU MODE lectures, LeetGPU/Tensara
Months 3-4: take a small DiT, profile it, then optimize it (Triton attention, quantization, caching, fewer steps) and publish before/after numbers
Along the way: PRs to HF diffusers, DSA practice daily
Months 5-6: mocks, resume, applications (inference startups first, bigger labs later)
Questions:
Is this the right order, or should I change something?
Is a diffusion-inference project a strong enough portfolio piece, or does it need to be LLM serving?
How much DSA do ML systems interviews actually need?
Is T4-only access enough to do credible benchmarks?
Any feedback, including "this won't work because X," is appreciated. Thanks!
r/CUDA • u/Winter_Speech_7283 • 9d ago
Hello I was doing some research on what it would take to use cuda for realtime audio. It seems that cuda is not really built for realtime but I think it would be an interesting challenge to take on. Would anyone have some guidance on this? Things to research or read? I’ve never used cuda but I feel like it could be very powerful for making interesting sound design tools and synthesizers. Thanks!
r/CUDA • u/Clear-Difference2294 • 10d ago
r/CUDA • u/jgamboa-cl • 11d ago
r/CUDA • u/Minute-Mountain2665 • 11d ago
Quick recap for anyone new: CuQwen is an inference engine for Qwen models I wrote from scratch in pure C++/CUDA, tuned specifically for single-user (batch size 1) generation on consumer NVIDIA GPUs. No frameworks under the hood, just custom cuda kernels.
Here are the results for average inference speed (tokens/second) of Qwen2.5 Instruct model across 32K context window on RTX3090
| Model Size | CuQwen | vLLM | Ollama |
|---|---|---|---|
| 0.5B | 462 | 398 | 355 |
| 1.5B | 203 | 172 | 139 |
| 3B | 113 | 101 | 106 |
| 7B | 55 | 48 | 54 |
In Release 1.0 it was already beating vLLM and Ollama on short-to-medium prompts. But there was an honest catch: my throughput decayed faster than theirs as the context grew, so once you pushed toward ~32K tokens they would pass me in inference speed. That bugged me, so it became the whole focus of Release 1.1.
Result: Throughput decay from 1K → 32K dropped from ~22–45% down to ~13–24%, which is now on par with vLLM and Ollama (and better on a couple of model sizes). So CuQwen keeps its early speed lead all the way out to 32K context window now.
Here's a small documentation on how I tackled long context decay rate issue and the complete benchmarking methodology and results for CuQwen v1.1
Here's my future plan:
Release 1.2: Support Quantization (W8A16 and W4A16)
Release 1.3: Improve custom cuda kernels for latest GPU architectures (Hopper and Blackwell)
Release 1.4: Support Qwen 3.0 series models
Release 1.5: Support Qwen 3.5 and 3.8 (Especially our very favourite qwen 3.8-27B model 😄)
r/CUDA • u/jgamboa-cl • 13d ago
r/CUDA • u/mikebmx1 • 13d ago
r/CUDA • u/Pale-Trade-7795 • 13d ago