attention mechanisms 1continuous monitoring 1eeg analysis 1streaming inference 1transformer models 1
From the 1 of 28 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
Hang Guo, Yawei Li, Luca Benini
Recent advances in Large Language Model (LLM) compression, such as quantization and pruning, have achieved notable success. However, as these techniques gradually approach their re…
cs.CL2025
Towards Extreme Pruning of LLMs with Plug-and-Play Mixed Sparsity
Chi Xu, Gefei Zhang, Yantong Zhu +4
N:M structured pruning is essential for large language models (LLMs) because it can remove less important network weights and reduce the memory and computation requirements. Existi…