collaborators

13 papers

cs.CV2026

Vanilla ViT for Automotive Point Cloud Semantic Segmentation

Gilles Puy, Nermin Samet, Alexandre Boulch +3

Plain Transformers have become the de-facto architecture for processing text, audio, image, and video, offering a unified backbone for multimodal learning. However, state-of-the-ar…

cs.CV2026

Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning

Shashanka Venkataramanan, Valentinos Pariza, Mohammadreza Salehi +5

We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many cases surpasses the performance…

cs.CV2026

Coevolving Representations in Joint Image-Feature Diffusion

Theodoros Kouzelis, Spyros Gidaris, Nikos Komodakis

Joint image-feature generative modeling has recently emerged as an effective strategy for improving diffusion training by coupling low-level VAE latents with high-level semantic fe…

cs.CV2026

Boosting Visual Instruction Tuning with Self-Supervised Guidance

Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc +2

Multimodal large language models (MLLMs) perform well on many vision-language tasks but often struggle with vision-centric problems that require fine-grained visual reasoning. Rece…

cs.CV2026

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction

Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis

Accurate future video prediction requires both high visual fidelity and consistent scene semantics, particularly in complex dynamic environments such as autonomous driving. We pres…

cs.CV2026

Driving on Registers

Ellington Kirby, Alexandre Boulch, Yihong Xu +11

We present DrivoR, a simple and efficient transformer-based architecture for end-to-end autonomous driving. Our approach builds on pretrained Vision Transformers (ViTs) and introdu…