1 paper · 1 filter
Yangyijian Liu, Hongyi Ye, Mingyang Li +1
Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference ne…