1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CL2024★ 1 cited
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
Hyun-rae Jo, Dongkun Shin
Recently, large language models (LLM) based on transformers are facing memory bottleneck issues due to KV cache, especially in long sequence handling. Previous researches proposed…
cs.LG2024
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
Seungmin Yu, Xiaodie Yi, Hayun Lee +1
N:M sparsity pruning is a powerful technique for compressing deep neural networks, utilizing NVIDIA's Sparse Tensor Core technology. This method benefits from hardware support for…
cs.LG2024
Realizing Unaligned Block-wise Pruning for DNN Acceleration on Mobile Devices
Hayun Lee, Dongkun Shin
With the recent proliferation of on-device AI, there is an increasing need to run computationally intensive DNNs directly on mobile devices. However, the limited computing and memo…