2 papers
cs.LG2026
Toward a First-Principles Update Geometry for the Language-Model Head
Aditya Somasundaram, Charles Guille-Escuret, Alexander Moreno +2
Muon motivates designing optimizer geometry around the function of each parameter block and uses the spectral norm for hidden linear layers. For the language-model head, the spectr…
cs.LG2026
Unlocking Lossless Speedups in LLMs via Discrete Diffusion
Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham +14
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overco…