1 citations · 2 across the 4 of their papers we have counts for
4 papers
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
Hyun-rae Jo, Dongkun Shin
Recently, large language models (LLM) based on transformers are facing memory bottleneck issues due to KV cache, especially in long sequence handling. Previous researches proposed…
Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
Seungmin Yu, Xiaodie Yi, Hayun Lee +1
N:M sparsity pruning is a powerful technique for compressing deep neural networks, utilizing NVIDIA's Sparse Tensor Core technology. This method benefits from hardware support for…
Octave-YOLO: Cross frequency detection network with octave convolution
Sangjune Shin, Dongkun Shin
Despite the rapid advancement of object detection algorithms, processing high-resolution images on embedded devices remains a significant challenge. Theoretically, the fully convol…
Realizing Unaligned Block-wise Pruning for DNN Acceleration on Mobile Devices
Hayun Lee, Dongkun Shin
With the recent proliferation of on-device AI, there is an increasing need to run computationally intensive DNNs directly on mobile devices. However, the limited computing and memo…