Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024★ 2 cited
LASER: Attention with Exponential Transformation
Sai Surya Duvvuri, Inderjit S. Dhillon
Transformers have had tremendous impact for several sequence related tasks, largely due to their ability to retrieve from any part of the sequence via softmax based dot-product att…
cs.LG2024
LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization
Jui-Nan Yen, Si Si, Zhao Meng +5
Low-rank adaption (LoRA) is a widely used parameter-efficient finetuning method for LLM that reduces memory requirements. However, current LoRA optimizers lack transformation invar…