most citedInteger or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models

5 citations · 6 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL20241 cited

BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Dayou Du, Yijia Zhang, Shijie Cao +4

The upscaling of Large Language Models (LLMs) has yielded impressive advances in natural language processing, yet it also poses significant deployment challenges. Weight quantizati…

cs.CL2023

AFPQ: Asymmetric Floating Point Quantization for LLMs

Yijia Zhang, Sicheng Zhang, Shijie Cao +4

Large language models (LLMs) show great performance in various tasks, but face deployment challenges from limited memory capacity and bandwidth. Low-bit weight quantization can sav…

cs.LG2023

Adam Accumulation to Reduce Memory Footprints of both Activations and Gradients for Large-scale DNN Training

Yijia Zhang, Yibo Han, Shijie Cao +5

Running out of GPU memory has become a main bottleneck for large-scale DNN training. How to reduce the memory footprint during training has received intensive research attention. W…

cs.CL2023

Accurate and Structured Pruning for Efficient Automatic Speech Recognition

Huiqiang Jiang, Li Lyna Zhang, Yuang Li +7

Automatic Speech Recognition (ASR) has seen remarkable advancements with deep neural networks, such as Transformer and Conformer. However, these models typically have large model s…

cs.LG20235 cited

Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models

Yijia Zhang, Lingran Zhao, Shijie Cao +6

Efficient deployment of large language models (LLMs) necessitates low-bit quantization to minimize model size and inference cost. While low-bit integer formats (e.g., INT8/INT4) ha…