2 papers
cs.LG2026
The Energy Consumption of Transformer Fine-Tuning: A Roofline-Inspired Scaling Model
Mansour Zoubeirou a Mayaki
Transformer-based models underpin modern natural language processing but incur rapidly growing computational and energy costs. As training scales in both model size and parallelism…
cs.LG2026
Generalization and Scaling Laws for Mixture-of-Experts Transformers
Mansour Zoubeirou a Mayaki
We develop a theory of generalization and scaling for Mixture-of-Experts (MoE) Transformers that cleanly separates \emph{active} per-input capacity from routing combinatorics. By c…