1 paper · 1 filter
Mingyu Yang, Mehdi Rezagholizadeh, Guihong Li +2
With the growing demand for deploying large language models (LLMs) across diverse applications, improving their inference efficiency is crucial for sustainable and democratized acc…