activity
20242026
collaborators

5 papers

cs.CV2026

Revisiting [CLS] and Patch Token Interaction in Vision Transformers

Alexis Marouani, Oriane Siméoni, Hervé Jégou +2

Vision Transformers have emerged as powerful, scalable and versatile representation learners. To capture both global and local features, a learnable [CLS] class token is typically…

cs.CV2025

DINOv3

Oriane Siméoni, Huy V. Vo, Maximilian Seitzer +23

Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. B…

cs.CV2025

Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models

Monika Wysoczańska, Antonin Vobecky, Amaia Cardiel +4

Recent CLIP-like Vision-Language Models (VLMs), pre-trained on large amounts of image-text pairs to align both modalities with a simple contrastive objective, have paved the way to…

cs.CV2025

LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression Comprehension

Amaia Cardiel, Eloi Zablocki, Elias Ramzi +2

Vision Language Models (VLMs) have demonstrated remarkable capabilities in various open-vocabulary tasks, yet their zero-shot performance lags behind task-specific fine-tuned model…

cs.CV2024

DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment

Cijo Jose, Théo Moutakanni, Dahyun Kang +11

Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models…