activity
20192025
most citedRPTQ: Reorder-based Post-training Quantization for Large Language Models

20 citations · 49 across the 9 of their papers we have counts for

collaborators

10 papers

cs.AR20251 cited

AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM

Yuanpeng Zhang, Xing Hu, Xi Chen +10

SRAM Processing-in-Memory (PIM) has emerged as the most promising implementation for high-performance PIM, delivering superior computing density, energy efficiency, and computation…

cs.AR2024

Algorithm-hardware co-design for Energy-Efficient A/D conversion in ReRAM-based accelerators

Chenguang Zhang, Zhihang Yuan, Xingchen Li +1

Deep neural networks are widely deployed in many fields. Due to the in-situ computation (known as processing in memory) capacity of the Resistive Random Access Memory (ReRAM) cross…

cs.CV2023

Latency-aware Unified Dynamic Networks for Efficient Image Recognition

Yizeng Han, Zeyu Liu, Zhihang Yuan +4

Dynamic computation has emerged as a promising avenue to enhance the inference efficiency of deep networks. It allows selective activation of computational units, leading to a redu…

cs.CL202320 cited

RPTQ: Reorder-based Post-training Quantization for Large Language Models

Zhihang Yuan, Lin Niu, Jiawei Liu +7

Large-scale language models (LLMs) have demonstrated impressive performance, but their deployment presents challenges due to their significant memory usage. This issue can be allev…

cs.CV20232 cited

Improving Post-Training Quantization on Object Detection with Task Loss-Guided Lp Metric

Lin Niu, Jiawei Liu, Zhihang Yuan +3

Efficient inference for object detection networks is a major challenge on edge devices. Post-Training Quantization (PTQ), which transforms a full-precision model into low bit-width…

cs.LG20234 cited

Benchmarking the Reliability of Post-training Quantization: a Particular Focus on Worst-case Performance

Zhihang Yuan, Jiawei Liu, Jiaxiang Wu +6

Post-training quantization (PTQ) is a popular method for compressing deep neural networks (DNNs) without modifying their original architecture or training procedures. Despite its e…