3 papers
cs.LG2026
Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
Dung Anh Hoang, Cuong Pham, Cuong Nguyen +3
Large Language Models (LLMs) deliver strong performance across a wide range of NLP tasks, but their massive sizes hinder deployment on resource-constrained devices. To reduce their…
cs.LG2025
Adaptive Layer-Wise Transformations for Post-Training Quantization of Large Language Models
Cuong Pham, Hoang Anh Dung, Cuong C. Nguyen +4
Large language models require significant computational resources for deployment, making quantization essential for practical applications. However, the main obstacle to effective…
cs.LG2025
Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language Models
Cuong Pham, Hoang Anh Dung, Cuong C. Nguyen +3
Large language models (LLMs) have significantly advanced natural language processing, but their massive parameter counts create substantial computational and memory challenges duri…