5 papers
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
Jiahao Nie, Gongjie Zhang, Wenbin An +4
Though Multi-modal Large Language Models (MLLMs) have recently achieved significant progress, they often struggle to understand diverse and complicated inter-object relations. Spec…
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
Wenbin An, Jiahao Nie, Yaqiang Wu +3
By integrating the perception capabilities of multimodal encoders with the generative power of Large Language Models (LLMs), Multimodal Large Language Models (MLLMs), exemplified b…
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
Wenbin An, Feng Tian, Sicong Leng +6
Despite great success across various multimodal tasks, Large Vision-Language Models (LVLMs) often encounter object hallucinations with generated textual responses being inconsisten…
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
Shilin Sun, Wenbin An, Feng Tian +5
Artificial intelligence (AI) has rapidly developed through advancements in computational power and the growth of massive datasets. However, this progress has also heightened challe…
Unleashing the Potential of Model Bias for Generalized Category Discovery
Wenbin An, Haonan Lin, Jiahao Nie +5
Generalized Category Discovery is a significant and complex task that aims to identify both known and undefined novel categories from a set of unlabeled data, leveraging another la…