1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.LG2024
LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization
Jui-Nan Yen, Si Si, Zhao Meng +5
Low-rank adaption (LoRA) is a widely used parameter-efficient finetuning method for LLM that reduces memory requirements. However, current LoRA optimizers lack transformation invar…
cs.CL2024
Accelerating Large Language Model Pretraining via LFR Pedagogy: Learn, Focus, and Review
Neha Prakriya, Jui-Nan Yen, Cho-Jui Hsieh +1
Traditional Large Language Model (LLM) pretraining relies on autoregressive language modeling with randomly sampled data from web-scale datasets. Inspired by human learning techniq…
cs.LG2023★ 1 cited
A Computationally Efficient Sparsified Online Newton Method
Fnu Devvrit, Sai Surya Duvvuri, Rohan Anil +3
Second-order methods hold significant promise for enhancing the convergence of deep neural network training; however, their large memory and computational demands have limited thei…