7 papers
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
Chen Huang, Xianhang Li, Vimal Thilak +2
Image-based Joint-Embedding Predictive Architecture (I-JEPA) offers a promising approach to visual self-supervised learning through masked feature prediction. However with the inhe…
Path-Constrained Mixture-of-Experts
Zijin Gu, Tatiana Likhomanenko, Vimal Thilak +2
Sparse Mixture-of-Experts (MoE) architectures route each token through a subset of experts at each layer independently. We propose viewing MoE computation through the lens of \emph…
How PARTs assemble into wholes: Learning the relative composition of images
Melika Ayoughi, Samira Abnar, Chen Huang +10
The composition of objects and their parts, along with object-object positional relationships, provides a rich source of information for representation learning. Hence, spatial-awa…
Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
Xianhang Li, Chen Huang, Chun-Liang Li +4
Video Joint Embedding Predictive Architectures (V-JEPA) learn generalizable off-the-shelf video representation by predicting masked regions in latent space with an exponential movi…
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
Samira Abnar, Harshay Shah, Dan Busbridge +3
Scaling the capacity of language models has consistently proven to be a reliable approach for improving performance and unlocking new capabilities. Capacity can be primarily define…
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
Zakaria Aldeneh, Vimal Thilak, Takuya Higuchi +2
This study explores using embedding rank as an unsupervised evaluation metric for general-purpose speech encoders trained via self-supervised learning (SSL). Traditionally, assessi…