2 papers
cs.CV2025
QwT-v2: Practical, Effective and Efficient Post-Training Quantization
Ningyuan Tang, Minghao Fu, Hao Yu +1
Network quantization is arguably one of the most practical network compression approaches for reducing the enormous resource consumption of modern deep neural networks. They usuall…
cs.CL2024
Effectively Compress KV Heads for LLM
Hao Yu, Zelan Yang, Shen Li +2
The advent of pre-trained large language models (LLMs) has revolutionized various natural language processing tasks. These models predominantly employ an auto-regressive decoding m…