collaborators

22 papers

cs.CV2026

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization

Zhong Wang, Zukang Xu, Xing Hu +1

Vision-Language Models (VLMs) achieve outstanding performance, yet their huge model size severely hinders deployment on edge devices with limited resources. As an efficient model c…

cs.LG2026

TORQ: Two-Level Orthogonal Rotation for MXFP4 Quantization

Zukang Xu, Xing Hu, Dawei Yang

As Large Language Models (LLMs) advance toward practical deployment, the Microscaling FP4 (MXFP4) format has emerged as a cornerstone for next-generation low-bit inference, owing t…

cs.CV2026

CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model

Houji Wen, Jiangyong Yu, Jun Li +1

Segment Anything Models (SAMs) are extensively used in computer vision for universal image segmentation, but deploying them on resource-constrained devices is challenging due to th…

cs.IR2026

R3-VAE: Reference Vector-Guided Rating Residual Quantization VAE for Generative Recommendation

Qiang Wan, Ze Yang, Dawei Yang +8

Generative Recommendation (GR) has gained traction for its merits of superior performance and cold-start capability. As the vital role in GR, Semantic Identifiers (SIDs) represent…

cs.LG2026

BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs

Zhixiong Zhao, Zukang Xu, Dawei Yang

Large language models (LLMs) have driven major progress in NLP, yet their substantial memory and compute demands still hinder practical deployment. Binarization can compress weight…

cs.LG2026

MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization

Zhixiong Zhao, Zukang Xu, Zhixuan Chen +1

Mixture-of-Experts (MoE) based large language models (LLMs) offer strong performance but suffer from high memory and computation costs. Weight binarization provides extreme efficie…