1 paper · 1 filter
Yaya Sy, Christophe Cerisara, Irina Illina
Current LLM structured pruning methods typically involve two steps: (1) compression with calibration data and (2) costly continued pretraining on billions of tokens to recover lost…