6 papers
Krause Synchronization Transformers
Jingkun Liu, Yisong Yue, Max Welling +1
Self-attention in Transformers relies on globally normalized softmax weights, causing all tokens to compete for influence at every layer. When composed across depth, this interacti…
Spontaneous symmetry breaking and Goldstone modes for deep information propagation
Nabil Iqbal, T. Anderson Keller, Yue Song +2
In physical systems, whenever a continuous symmetry is spontaneously broken, the system possesses excitations called Goldstone modes, which allow coherent information propagation o…
Kernel-Gradient Drifting Models
Maria Esteban-Casadevall, Jorge Carrasco-Pollo, Max Welling +3
We propose kernel-gradient drifting, a one-step generative modeling framework that replaces the fixed Euclidean displacement direction in drifting models with directions induced by…
Langevin Flows for Modeling Neural Latent Dynamics
Yue Song, T. Anderson Keller, Yisong Yue +2
Neural populations exhibit latent dynamical structures that drive time-evolving spiking activities, motivating the search for models that capture both intrinsic network dynamics an…
Unsupervised Representation Learning from Sparse Transformation Analysis
Yue Song, Thomas Anderson Keller, Yisong Yue +2
There is a vast literature on representation learning based on principles such as coding efficiency, statistical independence, causality, controllability, or symmetry. In this pape…
A Spatiotemporal Perspective on Dynamical Computation in Neural Information Processing Systems
T. Anderson Keller, Lyle Muller, Terrence J. Sejnowski +1
Spatiotemporal flows of neural activity, such as traveling waves, have been observed throughout the brain since the earliest recordings; yet there is still little consensus on thei…