3 papers
cs.LG2026
How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization
Leena Chennuru Vankadara, Moritz Haas, Luke Hayward +2
Recent frontier large language models predominantly rely on Mixture-of-Experts (MoE) architectures. Despite empirical progress, there is still no principled understanding of how hy…
cs.AI2025
Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees
Alessandro Breccia, Federica Gerace, Marco Lippi +2
We study whether a transformer network can learn the deterministic sequence of trees generated by the iterated prime factorization of the natural numbers. Each integer is mapped in…
cs.LG2025
The Importance of Being Lazy: Scaling Limits of Continual Learning
Jacopo Graldi, Alessandro Breccia, Giulia Lanzillotta +2
Despite recent efforts, neural networks still struggle to learn in non-stationary environments, and our understanding of catastrophic forgetting (CF) is far from complete. In this…