1 paper
Kwanhee Lee, Hyeondo Jang, Dongyeop Lee +2
Neural network pruning is a promising technique to mitigate the excessive computational and memory requirements of large language models (LLMs). Despite its promise, however, progr…