5 papers
Decoupled Relative Learning Rate Schedules
Jan Ludziejewski, Jan MaÅaÅnicki, Maciej Pióro +8
In this work, we introduce a novel approach for optimizing LLM training by adjusting learning rates across weights of different components in Transformer models. Traditional method…
Faster Semi-streaming Matchings via Alternating Trees
Slobodan MitroviÄ, Anish Mukherjee, Piotr Sankowski +1
We design a deterministic algorithm for the -approximate maximum matching problem. Our primary result demonstrates that this problem can be solved in semi-stre…
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
Jan Ludziejewski, Maciej Pióro, Jakub Krajewski +8
Mixture of Experts (MoE) architectures have significantly increased computational efficiency in both research and real-world applications of large-scale machine learning models. Ho…
LLM generated responses to mitigate the impact of hate speech
Jakub Podolak, Szymon Åukasik, PaweÅ Balawender +4
In this study, we explore the use of Large Language Models (LLMs) to counteract hate speech. We conducted the first real-life A/B test assessing the effectiveness of LLM-generated…
Online Multi-level Aggregation with Delays and Stochastic Arrivals
Mathieu Mari, MichaÅ PawÅowski, Runtian Ren +1
This paper presents a new research direction for online Multi-Level Aggregation (MLA) with delays. In this problem, we are given an edge-weighted rooted tree , and we have to se…