4 papers
Falcon Perception-HD: High Density Perception via Reinforcement Learning
Sofian Chaybouti, Yasser Dahou, Ngoc Dung Huynh +2
Autoregressive perception models trained to localize visual entities under the open-vocabulary setting are mostly trained using Supervised fine-tuning (SFT) with maximum likelihood…
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
Walid Bousselham, Angie Boggust, Hendrik Strobelt +1
As Vision-Language Models (VLMs) become increasingly sophisticated and widely used, it becomes more and more crucial to understand their decision-making process. Traditional explai…
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
Sofian Chaybouti, Sanath Narayan, Yasser Dahou +6
Vision foundation models trained via multi-teacher distillation offer a promising path toward unified visual representations, yet the learning dynamics and data efficiency of such…
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
Sofian Chaybouti, Walid Bousselham, Moritz Wolter +1
Video-Question-Answering (VideoQA) comprises the capturing of complex visual relation changes over time, remaining a challenge even for advanced Video Language Models (VLM), i.a.,…