15 citations · 16 across the 2 of their papers we have counts for
2 papers
cs.PF2025★ 1 cited
cuTeSpMM: Accelerating Sparse-Dense Matrix Multiplication using GPU Tensor Cores
Lizhi Xiang, Omid Asudeh, Gerald Sabin +2
Many recent GPUs feature matrix multiplication engines (aka Tensor Core Units or TCUs) that perform small fixed-size matrix-matrix products at very high throughput. They have been…
cs.LG2023★ 15 cited
HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks
Jinqi Xiao, Chengming Zhang, Yu Gong +5
Low-rank compression is an important model compression strategy for obtaining compact neural network models. In general, because the rank values directly determine the model comple…