1 citations · 1 across the 10 of their papers we have counts for
10 papers · 1 filter
Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness
Siqiao Mu, Diego Klabjan
We establish convergence guarantees of gradient descent for general feedforward neural networks of arbitrary width or depth, with no special requirements on the initialization or d…
Class-Grouped Normalized Momentum and Faster Hyperparameter Exploration to Tackle Class Imbalance in Federated Learning
Haemin Park, Diego Klabjan, Martin W. Braun +2
Class imbalance poses a critical challenge in federated learning (FL), where underrepresented classes suffer from poor predictive performance yet cannot be addressed by standard ce…
Rank-Accuracy Trade-off for LoRA: A Gradient-Flow Analysis
Michael Rushka, Diego Klabjan
Previous empirical studies have shown that LoRA achieves accuracy comparable to full-parameter methods on downstream fine-tuning tasks, even for rank-1 updates. By contrast, the th…
On the Convergence Rate of LoRA Gradient Descent
Siqiao Mu, Diego Klabjan
The low-rank adaptation (LoRA) algorithm for fine-tuning large models has grown popular in recent years due to its remarkable performance and low computational requirements. LoRA t…
Descend or Rewind? Stochastic Gradient Descent Unlearning
Siqiao Mu, Diego Klabjan
Machine unlearning algorithms aim to remove the impact of selected training data from a model without the computational expenses of retraining from scratch. Two such algorithms are…
A Mirror Descent Perspective of Smoothed Sign Descent
Shuyang Wang, Diego Klabjan
Recent work by Woodworth et al. (2020) shows that the optimization dynamics of gradient descent for overparameterized problems can be viewed as low-dimensional dual dynamics induce…