collaborators

6 papers

cs.LG2026

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation

Shihao Zhang, Rayan Saab

Post-training quantization is widely used for compressing large neural networks, but aggressive low-bit quantization can significantly degrade model quality. A common remedy is to…

cs.LG2026

Efficient Matrix Implementation for Rotary Position Embedding

Chen Minqi, Zhongqi Yue, Shihao Zhang +5

Rotary Position Embedding (RoPE) has become a core component of modern Transformer architectures across language, vision, and 3D domains. However, existing implementations rely on…

cs.LG2026

Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos

Haoyu Zhang, Shihao Zhang, Ian Colbert +1

Post-training quantization (PTQ) has become a crucial tool for reducing the memory and compute costs of modern deep neural networks, including large language models (LLMs). Among P…

cs.LG2026

Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization

Shihao Zhang, Haoyu Zhang, Ian Colbert +1

We introduce Qronos -- a new state-of-the-art post-training quantization algorithm that sequentially rounds and updates neural network weights. Qronos not only explicitly corrects…

cs.LG2025

Beacon: Post-Training Quantization with Integrated Grid Selection

Shihao Zhang, Rayan Saab

Quantization is a widely used compression technique for reducing the memory and computation costs of large pre-trained models. A key challenge in per-channel post-training quantiza…

cs.LG2025

Theoretical Guarantees for Low-Rank Compression of Deep Neural Networks

Shihao Zhang, Rayan Saab

Deep neural networks have achieved state-of-the-art performance across numerous applications, but their high memory and computational demands present significant challenges, partic…