2 papers
cs.CL2026
ART: Attention Run-time Termination for Efficient Large Language Model Decoding
Chen Qiu, Guozhong Li, Cristian McGee +2
Long-context decoding in Large Language Models (LLMs) is constrained by the cost of accessing and processing the Key-Value (KV) cache. Despite evidence that attention outputs depen…
cs.LG2026
Beyond LoRA: Is Sparsity-Induced Adaptation Better?
Elijah Cadenhead, Cristian McGee, Xin Li +2
Low-rank adaptation (LoRA) and its variants provide a memory- and compute-efficient alternative to full fine-tuning of pre-trained models. However, questions remain about the compa…