Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Towards Extreme Pruning of LLMs with Plug-and-Play Mixed Sparsity
Chi Xu, Gefei Zhang, Yantong Zhu +4
N:M structured pruning is essential for large language models (LLMs) because it can remove less important network weights and reduce the memory and computation requirements. Existi…
cs.CL2024
SCALM: Towards Semantic Caching for Automated Chat Services with Large Language Models
Jiaxing Li, Chi Xu, Feng Wang +3
Large Language Models (LLMs) have become increasingly popular, transforming a wide range of applications across various domains. However, the real-world effectiveness of their quer…