activity
20242026
collaborators

6 papers

cs.LG2026

Rethinking Residual Errors in Compensation-based LLM Quantization

Shuaiting Li, Juncan Deng, Kedong Xu +5

Methods based on weight compensation, which iteratively apply quantization and weight compensation to minimize the output error, have recently demonstrated remarkable success in qu…

cs.CV2025

SSVQ: Unleashing the Potential of Vector Quantization with Sign-Splitting

Shuaiting Li, Juncan Deng, Chenxuan Wang +5

Vector Quantization (VQ) has emerged as a prominent weight compression technique, showcasing substantially lower quantization errors than uniform quantization across diverse models…

cs.CV2025

ViM-VQ: Efficient Post-Training Vector Quantization for Visual Mamba

Juncan Deng, Shuaiting Li, Zeyu Wang +3

Visual Mamba networks (ViMs) extend the selective state space model (Mamba) to various vision tasks and demonstrate significant potential. As a promising compression technique, vec…

cs.CV2025

Efficiency Meets Fidelity: A Novel Quantization Framework for Stable Diffusion

Shuaiting Li, Juncan Deng, Zeyu Wang +5

Text-to-image generation via Stable Diffusion models (SDM) have demonstrated remarkable capabilities. However, their computational intensity, particularly in the iterative denoisin…

cs.CV2024

MVQ:Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization

Shuaiting Li, Chengxuan Wang, Juncan Deng +5

Vector quantization(VQ) is a hardware-friendly DNN compression method that can reduce the storage cost and weight-loading datawidth of hardware accelerators. However, conventional…

cs.LG2024

VQ4ALL: Efficient Neural Network Representation via a Universal Codebook

Juncan Deng, Shuaiting Li, Zeyu Wang +3

The rapid growth of the big neural network models puts forward new requirements for lightweight network representation methods. The traditional methods based on model compression h…