1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2025
When Compression Meets Model Compression: Memory-Efficient Double Compression for Large Language Models
Weilan Wang, Yu Mao, Dongdong Tang +3
Large language models (LLMs) exhibit excellent performance in various tasks. However, the memory requirements of LLMs present a great challenge when deploying on memory-limited dev…
cs.LG2024★ 1 cited
On the Compressibility of Quantized Large Language Models
Yu Mao, Weilan Wang, Hongchao Du +2
Deploying Large Language Models (LLMs) on edge or mobile devices offers significant benefits, such as enhanced data privacy and real-time processing capabilities. However, it also…