1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CV2024
Towards Accurate Post-training Quantization for Reparameterized Models
Luoming Zhang, Yefei He, Wen Fei +4
Model reparameterization is a widely accepted technique for improving inference speed without compromising performance. However, current Post-training Quantization (PTQ) methods of…
cs.AI2023★ 1 cited
Dual Grained Quantization: Efficient Fine-Grained Quantization for LLM
Luoming Zhang, Wen Fei, Weijia Wu +3
Large Language Models (LLMs) pose significant hardware challenges related to memory requirements and computational ability. There are two mainstream quantization schemes for LLMs:…