most citedRPTQ: Reorder-based Post-training Quantization for Large Language Models

20 citations · 34 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG20241 cited

I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models

Xing Hu, Yuan Cheng, Dawei Yang +4

Post-training quantization (PTQ) serves as a potent technique to accelerate the inference of large language models (LLMs). Nonetheless, existing works still necessitate a considera…

cs.AR2024

Algorithm-hardware co-design for Energy-Efficient A/D conversion in ReRAM-based accelerators

Chenguang Zhang, Zhihang Yuan, Xingchen Li +1

Deep neural networks are widely deployed in many fields. Due to the in-situ computation (known as processing in memory) capacity of the Resistive Random Access Memory (ReRAM) cross…

cs.LG20237 cited

PB-LLM: Partially Binarized Large Language Models

Yuzhang Shang, Zhihang Yuan, Qiang Wu +1

This paper explores network binarization, a radical form of quantization, compressing model weights to a single bit, specifically for Large Language Models (LLMs) compression. Due…

cs.CL202320 cited

RPTQ: Reorder-based Post-training Quantization for Large Language Models

Zhihang Yuan, Lin Niu, Jiawei Liu +7

Large-scale language models (LLMs) have demonstrated impressive performance, but their deployment presents challenges due to their significant memory usage. This issue can be allev…

cs.CV20232 cited

Improving Post-Training Quantization on Object Detection with Task Loss-Guided Lp Metric

Lin Niu, Jiawei Liu, Zhihang Yuan +3

Efficient inference for object detection networks is a major challenge on edge devices. Post-Training Quantization (PTQ), which transforms a full-precision model into low bit-width…

cs.LG20234 cited

Benchmarking the Reliability of Post-training Quantization: a Particular Focus on Worst-case Performance

Zhihang Yuan, Jiawei Liu, Jiaxiang Wu +6

Post-training quantization (PTQ) is a popular method for compressing deep neural networks (DNNs) without modifying their original architecture or training procedures. Despite its e…