3 papers
cs.LG2026
Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
Dung Anh Hoang, Cuong Pham, Cuong Nguyen +3
Large Language Models (LLMs) deliver strong performance across a wide range of NLP tasks, but their massive sizes hinder deployment on resource-constrained devices. To reduce their…
cs.LG2026
Gradient-Aligned Calibration for Post-Training Quantization of Diffusion Models
Dung Anh Hoang, Cuong Pham anh Trung Le, Jianfei Cai +1
Diffusion models have shown remarkable performance in image synthesis by progressively estimating a smooth transition from a Gaussian distribution of noise to a real image. Unfortu…
cs.LG2025
Adaptive Layer-Wise Transformations for Post-Training Quantization of Large Language Models
Cuong Pham, Hoang Anh Dung, Cuong C. Nguyen +4
Large language models require significant computational resources for deployment, making quantization essential for practical applications. However, the main obstacle to effective…