1 paper · 1 filter
Gal Bloch, Ariel Gera, Matan Orbach +2
We present \textbf{Flash-GMM}, a fused Triton kernel for efficient computation of Gaussian Mixture Models (GMMs) over large-scale data in a single GPU pass. By eliminating the need…