3 papers
cs.LG2024
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training
Philip Zmushko, Aleksandr Beznosikov, Martin Takáč +1
With the increase in the number of parameters in large language models, the process of pre-training and fine-tuning increasingly demands larger volumes of GPU memory. A significant…
cs.LG2024
Collaborative and Efficient Personalization with Mixtures of Adaptors
Abdulla Jasem Almansoori, Samuel Horváth, Martin Takáč
Heterogenous data is prevalent in real-world federated learning. We propose a parameter-efficient framework, Federated Low-Rank Adaptive Learning (FLoRAL), that allows clients to p…
math.OC2024
Methods for Convex -Smooth Optimization: Clipping, Acceleration, and Adaptivity
Eduard Gorbunov, Nazarii Tupitsa, Sayantan Choudhury +4
Due to the non-smoothness of optimization problems in Machine Learning, generalized smoothness assumptions have been gaining a lot of attention in recent years. One of the most pop…