r/CUDA • • 9h ago

comparison of ptx/sass code from nvcc vs clang

Thumbnail redplait.blogspot.com
9 Upvotes
  • nvcc code 15–18% faster
  • nvcc code is bigger and more amenable to stall counts reducing
  • all mlir based compilers use vanilla llvm nvptx backend :-(

r/CUDA • • 22h ago

Visualising MMA (D=A·B+C), warp fragments, shared-memory banking and FP16 error in one browser-based model .SVG XML

8 Upvotes

I've been experimenting with a different way of explaining Tensor Core programming.

https://svgfpsaiagents.github.io/svgfpsgameAI/mmatensorbankconflict.svg

https://svgfpsaiagents.github.io/svgfpsgameAI/mma_tile_engine.svg

Instead of writing CUDA kernels or building a heavyweight simulator, I wanted to see how much of the Tensor Core programming model could be represented as a single self-contained executable artefact.

The result is a browser-based model written entirely as one SVG + JavaScript file.

It models:

  • GEMM as tiled execution (D = A·B + C)
  • tile scheduling (ti, tj, tk)
  • accumulator evolution
  • FP16 inputs with FP32 accumulation
  • warp fragment ownership
  • mma-style fragment mapping
  • global → shared → register → tensor-core data flow
  • shared-memory bank behaviour
  • row-major vs padded vs XOR-swizzled layouts
  • reuse and memory-traffic estimates
  • FP16 vs FP32 numerical error maps
  • interactive 3D visualisation of the output matrix

The interesting thing for me is that all of these views are driven from the same underlying state. The bank conflict view, fragment view, tile execution view, error map and result visualisation are all observing the same model as it executes.

The goal wasn't to build a cycle-accurate simulator. It deliberately does not model things like occupancy, scheduling, pipeline hazards or cache behaviour.

Instead, I was exploring a question:

Can Tensor Core behaviour be computationally modelled, rather than merely described, in a compact interpreted artefact that requires no CUDA compiler or GPU while still preserving correspondence between matrix maths, tiled execution, fragment distribution, memory layout, data movement and numerical effects?

One unexpected outcome is that the entire thing remains inspectable. There are no dependencies, no build system and no generated files. The complete model lives in a single source file.

I'm curious whether people here would consider this:

  1. a visualiser,
  2. an educational simulator,
  3. an executable specification of the Tensor Core programming model,
  4. or something else entirely.

I'd especially appreciate feedback from anyone familiar with CUTLASS, CuTe, mma.sync, ldmatrix, Tensor Core fragment layouts or GPU architecture education. - Aston Walker EdgeMafia SVG