5 citations · 8 across the 7 of their papers we have counts for
6 papers · 1 filter
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
Peng Ding, Jingyu Wu, Jun Kuang +6
Multi-modal Large Language Models (MLLMs) have demonstrated remarkable performance on various visual-language understanding and generation tasks. However, MLLMs occasionally genera…
What Do Deep Saliency Models Learn about Visual Attention?
Shi Chen, Ming Jiang, Qi Zhao
In recent years, deep saliency models have made significant progress in predicting human visual attention. However, the mechanisms behind their success remain largely unexplained d…
Attention in Reasoning: Dataset, Analysis, and Modeling
Shi Chen, Ming Jiang, Jinhui Yang +1
While attention has been an increasingly popular component in deep neural networks to both interpret and boost the performance of models, little work has examined how attention pro…
REX: Reasoning-aware and Grounded Explanation
Shi Chen, Qi Zhao
Effectiveness and interpretability are two essential properties for trustworthy AI systems. Most recent studies in visual reasoning are dedicated to improving the accuracy of predi…
AiR: Attention with Reasoning Capability
Shi Chen, Ming Jiang, Jinhui Yang +1
While attention has been an increasingly popular component in deep neural networks to both interpret and boost performance of models, little work has examined how attention progres…
Boosted Attention: Leveraging Human Attention for Image Captioning
Shi Chen, Qi Zhao
Visual attention has shown usefulness in image captioning, with the goal of enabling a caption model to selectively focus on regions of interest. Existing models typically rely on…