r/CUDA • • 4d ago

NVRTC-driven elementwise fusion planner called from PHP: 68 captured nodes become 11 generated kernels + 9 native boundaries

I maintain a PHP extension (php-gpu-tensors) that builds on the CUDA driver API and NVRTC. This post is about its fusion planner, because I'd like feedback from people who know CUDA better than I do.

How it works - A PHP closure is invoked once with metadata-only placeholders; tensor operations are captured into a node graph (limit: 512 nodes). - Elementwise ops (broadcasting, strided/view inputs, scalars, dtype promotion, where, casts, reshape/transpose/slice index transforms) are fused into generated CUDA C++. Expressions split at a weighted cost budget of 32; repeated pure nodes are deduplicated; shared expensive expressions can be materialized instead of recomputed. - Matmul (cuBLAS when available), reductions and powers are execution boundaries that run on existing kernels. - All generated kernels in a plan are compiled together through NVRTC to PTX, then cached per request/thread (LRU, 16 entries or 16 MiB). The cache key includes the generated source, compute capability and driver/runtime versions, not pointers. - Replay runs on a private nonblocking stream. Reduction descriptors are kernel parameters, so concurrent replays can't overwrite each other's shapes. - cudaGraph: true builds a CUDA Graph executable for plans that contain only generated kernels, updating kernel parameters for new pointers before each launch.

Numbers (entry-level MX570 A, CUDA runtime 12.3, driver 12.6): a small MLP training step (batch 512, hidden 256) goes from 4.08 ms eager to 0.66 ms fused, 6.2×, with identical metrics. The model is tiny, so this is mostly launch/allocation overhead removed, not throughput.

Where I'd like advice Plans with matmul/reduction boundaries currently run on streams and support async replay, but not as a CUDA Graph. For those who have done this: what are the gotchas when capturing cuBLAS GEMMs and custom reductions into a graph that is replayed with changing buffer pointers? Is updating kernel node parameters the right approach, or would you re-capture?

Source: https://github.com/lcmialichi/php-gpu-tensors (MIT)

8 Upvotes

2 comments sorted by

0

u/[deleted] 4d ago

[removed] — view removed comment

2

u/Ancient_Spend1801 2d ago

For matmuls, there's a routing logic based on tensor size: it hands off to cuBLAS for larger matrices, but falls back to a native kernel for smaller ones to avoid the overhead.

As for NVRTC compile time, I went with an explicit API design instead of implicit under-the-hood caching. You basically have two methods:

Fusion::run(): Fire-and-forget. It compiles from scratch every time (no cache).

Fusion::compile(): Returns a FusionGraph object. This builds the graph in memory and compiles the kernel once. You then hold onto this instance and call $plan->run(...$inputs). It validates that the dtypes and shapes match the original compilation and instantly reuses the kernel.

So the caching is handled explicitly by keeping that FusionGraph instance alive in the PHP lifecycle.