activity
20232026
most citedDINOv3

31 citations · 43 across the 7 of their papers we have counts for

collaborators
Showing 2024Show all

6 papers · 1 filter

cs.CV20244 cited

DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment

Cijo Jose, Théo Moutakanni, Dahyun Kang +11

Self-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks. However, unlike vision-language models…

cs.CV2024

LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression Comprehension

Amaia Cardiel, Eloi Zablocki, Elias Ramzi +2

Vision Language Models (VLMs) have demonstrated remarkable capabilities in various open-vocabulary tasks, yet their zero-shot performance lags behind task-specific fine-tuned model…

cs.CV2024

MILAN: Milli-Annotations for Lidar Semantic Segmentation

Nermin Samet, Gilles Puy, Oriane Siméoni +1

Annotating lidar point clouds for autonomous driving is a notoriously expensive and time-consuming task. In this work, we show that the quality of recent self-supervised lidar scan…

cs.CV2024

Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models

Monika Wysoczańska, Antonin Vobecky, Amaia Cardiel +4

Recent CLIP-like Vision-Language Models (VLMs), pre-trained on large amounts of image-text pairs to align both modalities with a simple contrastive objective, have paved the way to…

cs.CV2024

Valeo4Cast: A Modular Approach to End-to-End Forecasting

Yihong Xu, Éloi Zablocki, Alexandre Boulch +11

Motion forecasting is crucial in autonomous driving systems to anticipate the future trajectories of surrounding agents such as pedestrians, vehicles, and traffic signals. In end-t…

cs.CV20248 cited

POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images

Antonin Vobecky, Oriane Siméoni, David Hurych +4

We describe an approach to predict open-vocabulary 3D semantic voxel occupancy map from input 2D images with the objective of enabling 3D grounding, segmentation and retrieval of f…