2 papers
cs.LG2026
Fine-Tuning Low-Bit Models with Gradient in Quantized Code Space
Shiguang Wu, Zhouchen Lin, Quanming Yao
Fine-tuning Low-bit models aims to adapt a quantized model while keeping the final deployed checkpoint in the same low-bit form. This setting is practically important as it reduces…
cs.LG2026
Attention Sinks Induce Gradient Sinks: Massive Activations as Gradient Regulators in Transformers
Yihong Chen, Zhouchen Lin, Quanming Yao
Attention sinks and massive activations are recurring and closely related phenomena in Transformer models. Existing explanations have largely focused on the forward pass, yet in pr…