1 paper · 1 filter
Sangmin Yoo, Srikanth Malla, Chiho Choi +2
The inference of large language models imposes significant computational workloads, often requiring the processing of billions of parameters. Although early-exit strategies have pr…