1 citations · 1 across the 5 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2024★ 1 cited
Art and Science of Quantizing Large-Scale Models: A Comprehensive Overview
Yanshu Wang, Tong Yang, Xiyan Liang +5
This paper provides a comprehensive overview of the principles, challenges, and methodologies associated with quantizing large-scale neural network models. As neural networks have…
cs.LG2024
QET: Enhancing Quantized LLM Parameters and KV cache Compression through Element Substitution and Residual Clustering
Yanshu Wang, Wang Li, Zhaoqian Yao +1
The matrix quantization entails representing matrix elements in a more space-efficient form to reduce storage usage, with dequantization restoring the original matrix for use. We f…
cs.LG2024
Athena: Efficient Block-Wise Post-Training Quantization for Large Language Models Using Second-Order Matrix Derivative Information
Yanshu Wang, Wenyang He, Tong Yang
Large Language Models (LLMs) have significantly advanced natural language processing tasks such as machine translation, text generation, and sentiment analysis. However, their larg…