8 papers
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
Sofian Chaybouti, Sanath Narayan, Yasser Dahou +6
Vision foundation models trained via multi-teacher distillation offer a promising path toward unified visual representations, yet the learning dynamics and data efficiency of such…
Falcon Perception
Aviraj Bevli, Sofian Chaybouti, Yasser Dahou +6
Perception-centric systems are typically implemented with a modular encoder-decoder pipeline: a vision backbone for feature extraction and a separate decoder (or late-fusion module…
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
Walid Bousselham, Sofian Chaybouti, Christian Rupprecht +2
Vision-language foundation models such as CLIP have achieved tremendous results in global vision-language alignment, but still show some limitations in creating representations for…
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
Brigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh +5
Vision-Language Models (VLMs) have achieved remarkable progress across tasks such as visual question answering and image captioning. Yet, the extent to which these models perform v…
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
Sofian Chaybouti, Walid Bousselham, Moritz Wolter +1
Video-Question-Answering (VideoQA) comprises the capturing of complex visual relation changes over time, remaining a challenge even for advanced Video Language Models (VLM), i.a.,…
MIX : a Multi-task Learning Approach to Solve Open-Domain Question Answering
Sofian Chaybouti, Achraf Saghe, Aymen Shabou
This paper introduces MIX, a multi-task deep learning approach to solve open-ended question-answering. First, we design our system as a multi-stage pipeline of 3 building blocks: a…