150 citations · 218 across the 17 of their papers we have counts for
Showing cs.ARShow all
2 papers · 1 filter
cs.AR2023
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
Jianyi Cheng, Cheng Zhang, Zhewen Yu +3
Model quantization represents both parameters (weights) and intermediate values (activations) in a more compact format, thereby directly reducing both computational and memory cost…
cs.AR2023★ 150 cited
OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
Cong Guo, Jiaming Tang, Weiming Hu +6
Transformer-based large language models (LLMs) have achieved great success with the growing model size. LLMs' size grows by every two years, which outpaces the hardware…