7 papers
CacheMuon: Using Temporal Preconditioning To Approximate Polar Factor
Bishnu Dev, Sushil Bohara, Martin TakÃ¡Ä +1
Muon is an optimizer that computes updates using the polar factor of the momentum matrix and has shown strong empirical performance across a range of training settings. A key compo…
PreLort: Prefix-Nested LoRA for Federated Fine-Tuning under Rank Heterogeneity
Muhammad Waseem, Nurbek Tastan, Andrej Jovanovic +4
Federated fine-tuning of large language models using parameter-efficient methods such as LoRA enables privacy-preserving adaptation of foundation models. Heterogeneous hardware res…
FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment
Riccardo Zaccone, Stefanos Laskaridis, Marco Ciccone +1
The growing scale of deep neural networks, encompassing large language models (LLMs) and vision transformers (ViTs), has made training from scratch prohibitively expensive and depl…
Can Muon Fine-tune Adam-Pretrained Models?
Xingyu Qu, Peigeng Huang, Samuel Horvath
Muon has emerged as an efficient alternative to Adam for pretraining, yet remains underused for fine-tuning. A key obstacle is that most open models are pretrained with Adam, and n…
Byzantine-Robust Optimization under -Smoothness
Arman Bolatov, Samuel Horváth, Martin TakÃ¡Ä +1
We consider distributed optimization under Byzantine attacks in the presence of -smoothness, a generalization of standard -smoothness that captures functions with sta…
Learning in the Null Space: Small Singular Values for Continual Learning
Cuong Anh Pham, Praneeth Vepakomma, Samuel Horváth
Alleviating catastrophic forgetting while enabling further learning is a primary challenge in continual learning (CL). Orthogonal-based training methods have gained attention for t…