1 paper · 1 filter
He Xiao, Qingyao Yang, Dirui Xie +7
Large language models with billions of parameters are often over-provisioned: many layers contribute little unique information yet dominate the memory and energy footprint during i…