6 papers
Jolia: Concept-Level Vision-Language Alignment for 3D CT Contrastive Learning
Julien Khlaut, Charles Corbière, Baptiste Callard +9
Vision-language contrastive pretraining has become the dominant recipe for 3D medical foundation models, leveraging the large volumes of paired scans and reports produced in clinic…
Curia-2: Scaling Self-Supervised Learning for Radiology Foundation Models
Antoine Saporta, Baptiste Callard, Corentin Dancette +5
The rapid growth of medical imaging has fueled the development of Foundation Models (FMs) to reduce the growing, unsustainable workload on radiologists. While recent FMs have shown…
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
Léo Butsanets, Charles Corbière, Julien Khlaut +2
In this work, we introduce RadImageNet-VQA, a large-scale dataset designed to advance radiologic visual question answering (VQA) on CT and MRI exams. Existing medical VQA datasets…
Boosting Omnidirectional Stereo Matching with a Pre-trained Depth Foundation Model
Jannik Endres, Oliver Hahn, Charles Corbière +3
Omnidirectional depth perception is essential for mobile robotics applications that require scene understanding across a full 360° field of view. Camera-based setups offer a cost-…
Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
Charles Corbière, Simon Roburin, Syrielle Montariol +2
While chain-of-thought (CoT) prompting improves reasoning in large language models, its effectiveness in vision-language models (VLMs) remains limited due to over-reliance on textu…
Helvipad: A Real-World Dataset for Omnidirectional Stereo Depth Estimation
Mehdi Zayene, Jannik Endres, Albias Havolli +4
Despite progress in stereo depth estimation, omnidirectional imaging remains underexplored, mainly due to the lack of appropriate data. We introduce Helvipad, a real-world dataset…