3 papers
cs.DC2026
Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices
Yangyijian Liu, Hongyi Ye, Mingyang Li +1
Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference ne…
cs.DC2025
PIPO: Pipelined Offloading for Efficient Inference on Consumer Devices
Yangyijian Liu, Jun Li, Wu-Jun Li
The high memory and computation demand of large language models (LLMs) makes them challenging to be deployed on consumer devices due to limited GPU memory. Offloading can mitigate…
cs.LG2024
LCQ: Low-Rank Codebook based Quantization for Large Language Models
Wen-Pu Cai, Ming-Yang Li, Wu-Jun Li
Large language models~(LLMs) have recently demonstrated promising performance in many tasks. However, the high storage and computational cost of LLMs has become a challenge for dep…