281 citations · 565 across the 6 of their papers we have counts for
6 papers
Long Range Language Modeling via Gated State Spaces
Harsh Mehta, Ankit Gupta, Ashok Cutkosky +1
State space models have shown to be effective at modeling long range dependencies, specially on sequence classification tasks. In this work we focus on autoregressive sequence mode…
Exploring Length Generalization in Large Language Models
Cem Anil, Yuhuai Wu, Anders Andreassen +7
The ability to extrapolate from short problem instances to longer ones is an important form of out-of-distribution generalization in reasoning tasks, and is crucial when learning f…
Solving Quantitative Reasoning Problems with Language Models
Aitor Lewkowycz, Anders Andreassen, David Dohan +11
Language models have achieved remarkable performance on a wide range of tasks that require natural language understanding. Nevertheless, state-of-the-art models have generally stru…
In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning
Behnam Neyshabur, Ryota Tomioka, Nathan Srebro
We present experiments demonstrating that some other form of capacity control, different from network size, plays a central role in learning multilayer feed-forward networks. We ar…
On Symmetric and Asymmetric LSHs for Inner Product Search
Behnam Neyshabur, Nathan Srebro
We consider the problem of designing locality sensitive hashes (LSH) for inner product similarity, and of the power of asymmetric hashes in this context. Shrivastava and Li argue t…
Clustering, Hamming Embedding, Generalized LSH and the Max Norm
Behnam Neyshabur, Yury Makarychev, Nathan Srebro
We study the convex relaxation of clustering and hamming embedding, focusing on the asymmetric case (co-clustering and asymmetric hamming embedding), understanding their relationsh…