1 paper · 1 filter
Rivaan Patil, Simon Dennis, Hao Guo +1
Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.7--3.2x the throughput of standard fine-tuning at ~40%…