collaborators

5 papers

cs.CV2026

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning

Revant Teotia, Adrien Bardes, Michael Rabbat +3

Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual streams naturally co-occur in vid…

cs.CV2025

DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models

Revant Teotia, Candace Ross, Karen Ullrich +4

Recent advances in text-to-image (T2I) models have achieved impressive quality and consistency. However, this has come at the cost of representation diversity. While automatic eval…

cs.CV2025

On Improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models

Tariq Berrada Ifriqi, Pietro Astolfi, Melissa Hall +8

Large-scale training of latent diffusion models (LDMs) has enabled unprecedented quality in image generation. However, the key components of the best performing LDM training recipe…

cs.LG2025

Lossless Compression of Vector IDs for Approximate Nearest Neighbor Search

Daniel Severo, Giuseppe Ottaviano, Matthew Muckley +2

Approximate nearest neighbor search for vectors relies on indexes that are most often accessed from RAM. Therefore, storage is the factor limiting the size of the database that can…

cs.LG2025

Qinco2: Vector Compression and Search with Improved Implicit Neural Codebooks

Théophane Vallaeys, Matthew Muckley, Jakob Verbeek +1

Vector quantization is a fundamental technique for compression and large-scale nearest neighbor search. For high-accuracy operating points, multi-codebook quantization associates d…