8 citations · 15 across the 4 of their papers we have counts for
4 papers
What Algorithms can Transformers Learn? A Study in Length Generalization
Hattie Zhou, Arwen Bradley, Etai Littwin +5
Large language models exhibit surprising emergent generalization properties, yet also struggle on many simple reasoning tasks such as arithmetic and parity. This raises the questio…
Adaptivity and Modularity for Efficient Generalization Over Task Complexity
Samira Abnar, Omid Saremi, Laurent Dinh +8
Can transformers generalize efficiently on problems that require dealing with examples with different levels of difficulty? We introduce a new task tailored to assess generalizatio…
Tensor Programs IVb: Adaptive Optimization in the Infinite-Width Limit
Greg Yang, Etai Littwin
Going beyond stochastic gradient descent (SGD), what new phenomena emerge in wide neural networks trained by adaptive optimizers like Adam? Here we show: The same dichotomy between…
The Loss Surface of Residual Networks: Ensembles and the Role of Batch Normalization
Etai Littwin, Lior Wolf
Deep Residual Networks present a premium in performance in comparison to conventional networks of the same depth and are trainable at extreme depths. It has recently been shown tha…