1 paper · 1 filter
Yuanhe Tian, Junjie Liu, Xican Yang +2
Pruning provides a practical solution to reduce the resources required to run large language models (LLMs) to benefit from their effective capabilities as well as control their cos…