Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
Shaowen Wang, Ge Zhang, Kairong Luo +6
Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra…
cs.LG2026
On the Residual Scaling of Looped Transformers: Stability and Transferability
Shaowen Wang, Bingrui Li, Ge Zhang +3
Looped (weight-tied) Transformers apply a shared residual block times (, same at each step), increasing effective depth without adding p…