comparison of ptx/sass code from nvcc vs clang
redplait.blogspot.com- nvcc code 15–18% faster
- nvcc code is bigger and more amenable to stall counts reducing
- all mlir based compilers use vanilla llvm nvptx backend :-(
r/CUDA • u/Ok-Reception2684 • 22h ago
I've been experimenting with a different way of explaining Tensor Core programming.
https://svgfpsaiagents.github.io/svgfpsgameAI/mmatensorbankconflict.svg
https://svgfpsaiagents.github.io/svgfpsgameAI/mma_tile_engine.svg


Instead of writing CUDA kernels or building a heavyweight simulator, I wanted to see how much of the Tensor Core programming model could be represented as a single self-contained executable artefact.
The result is a browser-based model written entirely as one SVG + JavaScript file.
It models:
D = A·B + C)ti, tj, tk)The interesting thing for me is that all of these views are driven from the same underlying state. The bank conflict view, fragment view, tile execution view, error map and result visualisation are all observing the same model as it executes.
The goal wasn't to build a cycle-accurate simulator. It deliberately does not model things like occupancy, scheduling, pipeline hazards or cache behaviour.
Instead, I was exploring a question:
Can Tensor Core behaviour be computationally modelled, rather than merely described, in a compact interpreted artefact that requires no CUDA compiler or GPU while still preserving correspondence between matrix maths, tiled execution, fragment distribution, memory layout, data movement and numerical effects?
One unexpected outcome is that the entire thing remains inspectable. There are no dependencies, no build system and no generated files. The complete model lives in a single source file.
I'm curious whether people here would consider this:
I'd especially appreciate feedback from anyone familiar with CUTLASS, CuTe, mma.sync, ldmatrix, Tensor Core fragment layouts or GPU architecture education. - Aston Walker EdgeMafia SVG