11 papers
Test-Time Hallucination Control in Large Vision-Language Models
Mehran Tamjidi, Hamidreza Dastmalchi, Ali Cheraghian +3
Object Hallucination in large vision-language models (LVLMs), where models generate non-factual content about input images, remains a critical barrier to their reliability in real-…
Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs
Ali Cheraghian, Hamidreza Dastmalchi, Hamed Barzamini +4
Recent advances in large vision-language models (LVLMs) have enabled powerful multimodal reasoning by integrating visual encoders with large language models (LLMs). However, their…
SteerSeg: Attention Steering for Reasoning Video Segmentation
Ali Cheraghian, Hamidreza Dastmalchi, Abdelwahed Khamis +3
Video reasoning segmentation requires localizing objects across video frames from natural language expressions, often involving spatial reasoning and implicit references. Recent ap…
TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
Ahmed Akl, Abdelwahed Khamis, Zhe Wang +3
Visual Question Answering (VQA) systems are notoriously brittle under distribution shifts and data scarcity. While previous solutions-such as ensemble methods and data augmentation…
Fighting Hallucinations with Counterfactuals: Diffusion-Guided Perturbations for LVLM Hallucination Suppression
Hamidreza Dastmalchi, Aijun An, Ali Cheraghian +1
While large vision-language models (LVLMs) achieve strong performance on multimodal tasks, they frequently generate hallucinations -- unfaithful outputs misaligned with the visual…
HIME: Mitigating Object Hallucinations in LVLMs via Hallucination Insensitivity Model Editing
Ahmed Akl, Abdelwahed Khamis, Ali Cheraghian +3
Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal understanding capabilities, yet they remain prone to object hallucination, where models describe non-ex…