1 paper · 1 filter
Yangyijian Liu, Jun Li, Wu-Jun Li
The high memory and computation demand of large language models (LLMs) makes them challenging to be deployed on consumer devices due to limited GPU memory. Offloading can mitigate…