7 papers
When to Align, When to Predict: A Phase Diagram for Multimodal Learning
Ilay Kamai, Hugues Van Assel, Aviv Regev +2
Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each…
Generate in Reconstruction Space, Match in Semantic Space: Transport Geometry for One-Step Generation
Hugues Van Assel, Edward De Brouwer, Saeed Saremi +2
Generative modeling and self-supervised representation learning (SSL) optimize structurally different objectives: generative training rewards distributional fidelity, while SSL rew…
Group Contrastive Learning for Weakly Paired Multimodal Data
Aditya Gorla, Hugues Van Assel, Jan-Christian Huetter +4
We present GROOVE, a semi-supervised multi-modal representation learning approach for high-content perturbation data where samples across modalities are weakly paired through share…
stable-pretraining-v1: Foundation Model Research Made Simple
Randall Balestriero, Hugues Van Assel, Sami BuGhanem +1
Foundation models and self-supervised learning (SSL) have become central to modern AI, yet research in this area remains hindered by complex codebases, redundant re-implementations…
Ditch the Denoiser: Emergence of Noise Robustness in Self-Supervised Learning from Data Curriculum
Wenquan Lu, Jiaqi Zhang, Hugues Van Assel +1
Self-Supervised Learning (SSL) has become a powerful solution to extract rich representations from unlabeled data. Yet, SSL research is mostly focused on clean, curated and high-qu…
Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning
Hugues Van Assel, Mark Ibrahim, Tommaso Biancalani +2
Reconstruction and joint embedding have emerged as two leading paradigms in Self Supervised Learning (SSL). Reconstruction methods focus on recovering the original sample from a di…