4 papers
Is CLIP ideal? No. Can we fix it? Yes!
Raphi Kang, Yue Song, Georgia Gkioxari +1
Contrastive Language-Image Pre-Training (CLIP) is a popular method for learning multimodal latent spaces with well-organized semantics. Despite its wide range of applications, CLIP…
Langevin Flows for Modeling Neural Latent Dynamics
Yue Song, T. Anderson Keller, Yisong Yue +2
Neural populations exhibit latent dynamical structures that drive time-evolving spiking activities, motivating the search for models that capture both intrinsic network dynamics an…
Unsupervised Representation Learning from Sparse Transformation Analysis
Yue Song, Thomas Anderson Keller, Yisong Yue +2
There is a vast literature on representation learning based on principles such as coding efficiency, statistical independence, causality, controllability, or symmetry. In this pape…
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
Raphi Kang, Hongqiao Chen, Georgia Gkioxari +1
Spatio-temporal reasoning is a remarkable capability of Vision Language Models (VLMs), but the underlying mechanisms of such abilities remain largely opaque. We postulate that visu…