1 paper
Bo-Kyeong Kim, Geonmin Kim, Tae-Ho Kim +4
Structured pruning of modern large language models (LLMs) has emerged as a way of decreasing their high computational needs. Width pruning reduces the size of projection weight mat…