150 citations · 153 across the 2 of their papers we have counts for
2 papers
cs.AR2023★ 150 cited
OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
Cong Guo, Jiaming Tang, Weiming Hu +6
Transformer-based large language models (LLMs) have achieved great success with the growing model size. LLMs' size grows by every two years, which outpaces the hardware…
cs.LG2022★ 3 cited
ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network Quantization
Cong Guo, Chen Zhang, Jingwen Leng +5
Quantization is a technique to reduce the computation and memory cost of DNN models, which are getting increasingly large. Existing quantization solutions use fixed-point integer o…