collaborators

7 papers

cs.LG2026

CacheMuon: Using Temporal Preconditioning To Approximate Polar Factor

Bishnu Dev, Sushil Bohara, Martin Takáč +1

Muon is an optimizer that computes updates using the polar factor of the momentum matrix and has shown strong empirical performance across a range of training settings. A key compo…

cs.DC2026

PreLort: Prefix-Nested LoRA for Federated Fine-Tuning under Rank Heterogeneity

Muhammad Waseem, Nurbek Tastan, Andrej Jovanovic +4

Federated fine-tuning of large language models using parameter-efficient methods such as LoRA enables privacy-preserving adaptation of foundation models. Heterogeneous hardware res…

cs.LG2026

FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model Deployment

Riccardo Zaccone, Stefanos Laskaridis, Marco Ciccone +1

The growing scale of deep neural networks, encompassing large language models (LLMs) and vision transformers (ViTs), has made training from scratch prohibitively expensive and depl…

cs.LG2026

Can Muon Fine-tune Adam-Pretrained Models?

Xingyu Qu, Peigeng Huang, Samuel Horvath

Muon has emerged as an efficient alternative to Adam for pretraining, yet remains underused for fine-tuning. A key obstacle is that most open models are pretrained with Adam, and n…

cs.LG2026

Byzantine-Robust Optimization under -Smoothness

Arman Bolatov, Samuel Horváth, Martin Takáč +1

We consider distributed optimization under Byzantine attacks in the presence of -smoothness, a generalization of standard -smoothness that captures functions with sta…

cs.LG2026

Learning in the Null Space: Small Singular Values for Continual Learning

Cuong Anh Pham, Praneeth Vepakomma, Samuel Horváth

Alleviating catastrophic forgetting while enabling further learning is a primary challenge in continual learning (CL). Orthogonal-based training methods have gained attention for t…