1 paper · 1 filter
Juyun Wee, Minjae Park, Jaeho Lee
Depth pruning aims to reduce the inference cost of a large language model without any hardware-specific complications, by simply removing several less important transformer blocks.…