27 citations · 29 across the 9 of their papers we have counts for
7 papers · 1 filter
Pretraining Recurrent Networks without Recurrence
Akarsh Kumar, Phillip Isola
Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations. Standard backpropagation through time (BPTT) addresses this problem poorl…
The Truth Lies Somewhere in the Middle (of the Generated Tokens)
Sophie L. Wang, Phillip Isola, Brian Cheung
How should hidden states generated autoregressively be collapsed into a representation that reflects a language model's internal state? Despite tokens being generated under causal…
Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
Yulu Gan, Phillip Isola
Pretraining produces a learned parameter vector that is typically treated as a starting point for further iterative adaptation. In this work, we instead view the outcome of pretrai…
Training Transformers with Enforced Lipschitz Constants
Laker Newhouse, R. Preston Hess, Franz Cesista +3
Neural networks are often highly sensitive to input and weight perturbations. This sensitivity has been linked to pathologies such as vulnerability to adversarial examples, diverge…
Scalable Optimization in the Modular Norm
Tim Large, Yang Liu, Minyoung Huh +3
To improve performance in contemporary deep learning, one is interested in scaling up the neural network in terms of both the number and the size of the layers. When ramping up the…
The Platonic Representation Hypothesis
Minyoung Huh, Brian Cheung, Tongzhou Wang +1
We argue that representations in AI models, particularly deep networks, are converging. First, we survey many examples of convergence in the literature: over time and across multip…