5 citations · 5 across the 5 of their papers we have counts for
8 papers
DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video Generation
Shuaiting Li, Zelin Gao, Haibin Shen +3
Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high memory and computational costs hinder practical deployment. Quantization…
Rethinking Residual Errors in Compensation-based LLM Quantization
Shuaiting Li, Juncan Deng, Kedong Xu +5
Methods based on weight compensation, which iteratively apply quantization and weight compensation to minimize the output error, have recently demonstrated remarkable success in qu…
ViM-VQ: Efficient Post-Training Vector Quantization for Visual Mamba
Juncan Deng, Shuaiting Li, Zeyu Wang +3
Visual Mamba networks (ViMs) extend the selective state space model (Mamba) to various vision tasks and demonstrate significant potential. As a promising compression technique, vec…
SSVQ: Unleashing the Potential of Vector Quantization with Sign-Splitting
Shuaiting Li, Juncan Deng, Chenxuan Wang +5
Vector Quantization (VQ) has emerged as a prominent weight compression technique, showcasing substantially lower quantization errors than uniform quantization across diverse models…
MVQ:Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
Shuaiting Li, Chengxuan Wang, Juncan Deng +5
Vector quantization(VQ) is a hardware-friendly DNN compression method that can reduce the storage cost and weight-loading datawidth of hardware accelerators. However, conventional…
VQ4ALL: Efficient Neural Network Representation via a Universal Codebook
Juncan Deng, Shuaiting Li, Zeyu Wang +3
The rapid growth of the big neural network models puts forward new requirements for lightweight network representation methods. The traditional methods based on model compression h…