3 papers
cs.LG2024
Beyond Backpropagation: Optimization with Multi-Tangent Forward Gradients
Katharina Flügel, Daniel Coquelin, Marie Weiel +3
The gradients used to train neural networks are typically computed using backpropagation. While an efficient way to obtain exact gradients, backpropagation is computationally expen…
cs.LG2024
AB-Training: A Communication-Efficient Approach for Distributed Low-Rank Learning
Daniel Coquelin, Katherina Flügel, Marie Weiel +5
Communication bottlenecks severely hinder the scalability of distributed neural network training, particularly in high-performance computing (HPC) environments. We introduce AB-tra…
cs.LG2024
Harnessing Orthogonality to Train Low-Rank Neural Networks
Daniel Coquelin, Katharina Flügel, Marie Weiel +4
This study explores the learning dynamics of neural networks by analyzing the singular value decomposition (SVD) of their weights throughout training. Our investigation reveals tha…