1 paper · 1 filter
Yuyang Wang, Haoyu Yao, Pengcheng Xie
Full-parameter fine-tuning of large language models has substantial memory costs because backpropagation stores activations and gradients. Zeroth-order optimization avoids this by…