activity
20242026
collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

NewtPhys: Do Foundation Models Understand Newtonian Physics?

Sebastian Cavada, Soumava Paul, Tuan-Hung Vu +2

Previous work has evaluated physics reasoning in foundation models using synthetic or semi-synthetic scenes and visual question-answering tasks. However, these benchmarks emphasize…

cs.CV2026

Domain Adaptation with a Single Vision-Language Embedding

Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc +2

Domain adaptation has been extensively investigated in computer vision but still requires access to target data at the training time, which might be difficult to obtain in real-wor…

cs.CV2026

Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning

Shashanka Venkataramanan, Valentinos Pariza, Mohammadreza Salehi +5

We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance…

cs.CV2026

3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds

Ryousuke Yamada, Kohsuke Ide, Yoshihiro Fukuhara +4

Despite recent progress in 3D self-supervised learning, collecting large-scale 3D scene scans remains expensive and labor-intensive. In this work, we investigate whether 3D represe…

cs.CV2026

Driving on Registers

Ellington Kirby, Alexandre Boulch, Yihong Xu +11

We present DrivoR, a simple and efficient transformer-based architecture for end-to-end autonomous driving. Our approach builds on pretrained Vision Transformers (ViTs) and introdu…

cs.CV2026

CLIP's Visual Embedding Projector is a Few-shot Cornucopia

Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc +2

We introduce ProLIP, a simple and architecture-agnostic method for adapting contrastively pretrained vision-language models, such as CLIP, to few-shot classification. ProLIP fine-t…