8 papers
When to Align, When to Predict: A Phase Diagram for Multimodal Learning
Ilay Kamai, Hugues Van Assel, Aviv Regev +2
Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each…
Generate in Reconstruction Space, Match in Semantic Space: Transport Geometry for One-Step Generation
Hugues Van Assel, Edward De Brouwer, Saeed Saremi +2
Generative modeling and self-supervised representation learning (SSL) optimize structurally different objectives: generative training rewards distributional fidelity, while SSL rew…
E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
Shuvom Sadhuka, Drew Prinster, Clara Fannjiang +4
Agentic AI systems execute a sequence of actions, such as reasoning steps or tool calls, in response to a user prompt. To evaluate the success of their trajectories, researchers ha…
Group Contrastive Learning for Weakly Paired Multimodal Data
Aditya Gorla, Hugues Van Assel, Jan-Christian Huetter +4
We present GROOVE, a semi-supervised multi-modal representation learning approach for high-content perturbation data where samples across modalities are weakly paired through share…
Sparse Mixture-of-Experts for Multi-Channel Imaging: Are All Channel Interactions Required?
Sukwon Yun, Heming Yao, Burkhard Hoeckendorf +3
Vision Transformers () have become the backbone of vision foundation models, yet their optimization for multi-channel domains - such as cell painting or satellite imag…
Supervised Contrastive Block Disentanglement
Taro Makino, Ji Won Park, Natasa Tagasovska +11
Real-world datasets often combine data collected under different experimental conditions. This yields larger datasets, but also introduces spurious correlations that make it diffic…