5 papers · 1 filter
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
Sofian Chaybouti, Sanath Narayan, Yasser Dahou +6
Vision foundation models trained via multi-teacher distillation offer a promising path toward unified visual representations, yet the learning dynamics and data efficiency of such…
Falcon Perception
Aviraj Bevli, Sofian Chaybouti, Yasser Dahou +6
Perception-centric systems are typically implemented with a modular encoder-decoder pipeline: a vision backbone for feature extraction and a separate decoder (or late-fusion module…
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
Brigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh +5
Vision-Language Models (VLMs) have achieved remarkable progress across tasks such as visual question answering and image captioning. Yet, the extent to which these models perform v…
Vision-Language Models Can't See the Obvious
Yasser Dahou, Ngoc Dung Huynh, Phuc H. Le-Khac +3
We present Saliency Benchmark (SalBench), a novel benchmark designed to assess the capability of Large Vision-Language Models (LVLM) in detecting visually salient features that are…
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
Wamiq Reyaz Para, Abdelrahman Eldesokey, Zhenyu Li +3
We introduce an approach for 3D head avatar generation and editing with multi-modal conditioning based on a 3D Generative Adversarial Network (GAN) and a Latent Diffusion Model (LD…