2 papers
cs.LG2024
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
Seungmin Yu, Xiaodie Yi, Hayun Lee +1
N:M sparsity pruning is a powerful technique for compressing deep neural networks, utilizing NVIDIA's Sparse Tensor Core technology. This method benefits from hardware support for…
cs.LG2024
Realizing Unaligned Block-wise Pruning for DNN Acceleration on Mobile Devices
Hayun Lee, Dongkun Shin
With the recent proliferation of on-device AI, there is an increasing need to run computationally intensive DNNs directly on mobile devices. However, the limited computing and memo…