1 paper
Jungmin Lee, Gwangeun Byeon, Yulhwa Kim +1
Pruning has emerged as a promising direction for accelerating large language model (LLM) inference, yet existing approaches often suffer from instability because they rely on offli…