1 paper · 1 filter
Pingjie Wang, Ziqing Fan, Shengchao Hu +3
Structured pruning is a promising hardware-friendly compression technique for large language models (LLMs), which is expected to be retraining-free to avoid the enormous retraining…