activity
20242026
most citedThe Platonic Representation Hypothesis

27 citations · 29 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Pretraining Recurrent Networks without Recurrence

Akarsh Kumar, Phillip Isola

Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations. Standard backpropagation through time (BPTT) addresses this problem poorl…

cs.LG2026

The Truth Lies Somewhere in the Middle (of the Generated Tokens)

Sophie L. Wang, Phillip Isola, Brian Cheung

How should hidden states generated autoregressively be collapsed into a representation that reflects a language model's internal state? Despite tokens being generated under causal…

cs.LG2026

Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights

Yulu Gan, Phillip Isola

Pretraining produces a learned parameter vector that is typically treated as a starting point for further iterative adaptation. In this work, we instead view the outcome of pretrai…

cs.LG2025★ 1 cited

Training Transformers with Enforced Lipschitz Constants

Laker Newhouse, R. Preston Hess, Franz Cesista +3

Neural networks are often highly sensitive to input and weight perturbations. This sensitivity has been linked to pathologies such as vulnerability to adversarial examples, diverge…

cs.LG2024

Scalable Optimization in the Modular Norm

Tim Large, Yang Liu, Minyoung Huh +3

To improve performance in contemporary deep learning, one is interested in scaling up the neural network in terms of both the number and the size of the layers. When ramping up the…

cs.LG2024★ 27 cited

The Platonic Representation Hypothesis

Minyoung Huh, Brian Cheung, Tongzhou Wang +1

We argue that representations in AI models, particularly deep networks, are converging. First, we survey many examples of convergence in the literature: over time and across multip…