2 papers
cs.LG2026
Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts
Soumajyoti Sarkar, Yuxin Tang, Sheng Zha
In large-scale pretraining, the algorithm, architecture, and systems decisions are conventionally made in disconnected stages. A scaling law stage selects an architecture and train…
cs.LG2026
Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation
Joshua Shay Kricheli, Alexander Lawrence Reid, Soumajyoti Sarkar +2
Neural scaling laws approximate a language model's loss as a power-law function of parameter count and token count . Following Chinchilla-style compute-optimal training, man…