activity
20242026
collaborators

9 papers

cs.IR2026

Subtraction Gets You More: Gap-Aware Retrieval for Multimodal Multi-Hop QA

Sunah O, Jay-Yoon Lee

In multimodal multi-hop question answering, we focus on the initial retrieval stage via two distinct tasks: (1) evidence set completion, retrieving missing evidence given context,…

cs.CV2026

Efficient Universal Perception Encoder

Chenchen Zhu, Saksham Suri, Cijo Jose +8

Running AI models on smart edge devices can unlock versatile user experiences, but presents challenges due to limited compute and the need to handle multiple tasks simultaneously.…

cs.AI2025

Disentangling the Factors of Convergence between Brains and Computer Vision Models

Joséphine Raugel, Marc Szafraniec, Huy V. Vo +5

Many AI models trained on natural images develop representations that resemble those of the human brain. However, the factors that drive this brain-model similarity remain poorly u…

cs.CV2025

DINOv3

Oriane Siméoni, Huy V. Vo, Maximilian Seitzer +23

Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. B…

cs.CV2025

Back to the Features: DINO as a Foundation for Video World Models

Federico Baldassarre, Marc Szafraniec, Basile Terver +6

We present DINO-world, a powerful generalist video world model trained to predict future frames in the latent space of DINOv2. By leveraging a pre-trained image encoder and trainin…

cs.CV2025

Cluster and Predict Latent Patches for Improved Masked Image Modeling

Timothée Darcet, Federico Baldassarre, Maxime Oquab +2

Masked Image Modeling (MIM) offers a promising approach to self-supervised representation learning, however existing MIM models still lag behind the state-of-the-art. In this paper…