Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models
Rayyan Abdalla, Amir Hussein, Min Wu +1
Post-training quantization (PTQ) is critical for the efficient deployment of large language models (LLMs). Recent ultra-low-bit PTQ methods rely on rigid weight-saliency assumption…
cs.AI2026
Compress Then Adapt? No, Do It Together via Task-aware Union of Subspaces
Jingze Ge, Yun Liu, Xue Geng +4
Adapting large pretrained models to diverse tasks is now routine, yet the two dominant strategies of parameter-efficient fine-tuning (PEFT) and low-rank compression are typically c…