3 papers
cs.LG2026
Deriving Hyperparameter Scaling Laws via Modern Optimization Theory
Egor Shulgin, Dimitri von Rütte, Tianyue H. Zhang +3
Hyperparameter transfer has become an important component of modern large-scale training recipes. Existing methods, such as muP, primarily focus on transfer between model sizes, wi…
cs.LG2026
Scaling Behavior of Discrete Diffusion Language Models
Dimitri von Rütte, Janis Fluri, Omead Pooladzandi +3
Modern LLM pre-training consumes vast amounts of compute and training data, making the scaling behavior, or scaling laws, of different models a key distinguishing factor. Discrete…
cs.CL2025
Generalized Interpolating Discrete Diffusion
Dimitri von Rütte, Janis Fluri, Yuhui Ding +3
While state-of-the-art language models achieve impressive results through next-token prediction, they have inherent limitations such as the inability to revise already generated to…