2 papers
cs.AI2026
Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs
Mohanad Odema, Gabrielle De Micheli, Dayin Gou +3
Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-le…
cs.LG2025
CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression
Dayin Gou, Sanghyun Byun, Nilesh Malpeddi +4
Large Language Models (LLMs) typically rely on a large number of parameters for token embedding, leading to substantial storage requirements and memory footprints. In particular, L…