13 papers
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
Gilles Puy, Nermin Samet, Alexandre Boulch +3
Plain Transformers have become the de-facto architecture for processing text, audio, image, and video, offering a unified backbone for multimodal learning. However, state-of-the-ar…
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
Shashanka Venkataramanan, Valentinos Pariza, Mohammadreza Salehi +5
We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance…
Coevolving Representations in Joint Image-Feature Diffusion
Theodoros Kouzelis, Spyros Gidaris, Nikos Komodakis
Joint image-feature generative modeling has recently emerged as an effective strategy for improving diffusion training by coupling low-level VAE latents with high-level semantic fe…
Boosting Visual Instruction Tuning with Self-Supervised Guidance
Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc +2
Multimodal large language models (MLLMs) perform well on many vision-language tasks but often struggle with vision-centric problems that require fine-grained visual reasoning. Rece…
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis
Accurate future video prediction requires both high visual fidelity and consistent scene semantics, particularly in complex dynamic environments such as autonomous driving. We pres…
Driving on Registers
Ellington Kirby, Alexandre Boulch, Yihong Xu +11
We present DrivoR, a simple and efficient transformer-based architecture for end-to-end autonomous driving. Our approach builds on pretrained Vision Transformers (ViTs) and introdu…