2 papers
cs.LG2026
SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
Yeonsik Park, Hyeonseong Kim, Seungkyu Choi
Post-training quantization (PTQ) has emerged as a prevailing technique for deploying large language models (LLMs) efficiently in terms of both memory and computation, across edge d…
cs.LG2025
Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring
Dongyoung Lee, Seungkyu Choi, Ik Joon Chang
Large-scale language models (LLMs) excel in language processing tasks but face deployment challenges due to high memory and computational demands. While low-bit quantization, such…