2 papers
cs.CL2025
ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models
Junho Yoon, Geom Lee, Donghyeon Jeon +2
Quantization has been widely studied as an effective technique for reducing the memory requirement of large language models (LLMs), potentially improving the latency time as well.…
cs.CL2025
CacheFocus: Dynamic Cache Re-Positioning for Efficient Retrieval-Augmented Generation
Kun-Hui Lee, Eunhwan Park, Donghoon Han +1
Large Language Models (LLMs) excel across a variety of language tasks yet are constrained by limited input lengths and high computational costs. Existing approaches\textemdash such…