latent state // log

CUDA Compute // 1 post

2026
JUL 28 // CUDA · ML
Matrix Multiplication, From Definition to Cache Lines
One multiply, three implementations, three orders of magnitude — measured on a Skylake desktop, then explained: cache lines, arithmetic intensity, and the tiling idea every BLAS and every GPU matmul kernel is built on.