From the 1 of 22 linked papers with an AI index.
22 papers
How the Hessian-Spectrum of Neural Networks Depends on Data
Jasraj Singh, Enea Monzio Compagnoni, Antonio Orvieto
The paper analytically derives the Hessian eigenvalues for linear neural networks of any size and data distribution, showing that solution sharpness under MSE loss is tied to the l…
Selective Rotary Position Embedding
Sajad Movahedi, Timur Carstensen, Arshia Afzal +3
Position information is essential for language modeling. In softmax transformers, Rotary Position Embeddings (\textit{RoPE}) encode positions through \textit{fixed-angle} rotations…
Beyond a Single Explanation of the Adam--SGD Gap
Chenxiang Zhang, Rustem Islamov, Enea Monzio Compagnoni +3
Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties.…
On the Interaction of Batch Noise, Adaptivity, and Compression, under -Smoothness: An SDE Approach
Enea Monzio Compagnoni, Rustem Islamov, Frank Norbert Proske +3
Distributed stochastic optimization intertwines (i) stochastic gradient noise, (ii) communication compression, and (iii) adaptive/normalized updates. While each factor has been stu…
GRASP: Deterministic argument ranking in interaction graphs
Diganta Misra, Antonio Orvieto, Rediet Abebe +1
Large language models are increasingly deployed as automated judges to evaluate the strength of arguments. As this role expands, their legitimacy depends on consistency, transparen…
Muown: Row-Norm Control for Muon Optimization
Kai Lion, Florian Hübler, Bingcong Li +2
Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work has observed that, for Muon wi…