10 citations · 16 across the 8 of their papers we have counts for
8 papers · 1 filter
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
Wenbin An, Jiahao Nie, Yaqiang Wu +3
By integrating the perception capabilities of multimodal encoders with the generative power of Large Language Models (LLMs), Multimodal Large Language Models (MLLMs), exemplified b…
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
Shilin Sun, Wenbin An, Feng Tian +5
Artificial intelligence (AI) has rapidly developed through advancements in computational power and the growth of massive datasets. However, this progress has also heightened challe…
Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image Editing
Haonan Lin, Mengmeng Wang, Jiahao Wang +7
Text-guided diffusion models have significantly advanced image editing, enabling high-quality and diverse modifications driven by text prompts. However, effective editing requires…
Flipped Classroom: Aligning Teacher Attention with Student in Generalized Category Discovery
Haonan Lin, Wenbin An, Jiahao Wang +6
Recent advancements have shown promise in applying traditional Semi-Supervised Learning strategies to the task of Generalized Category Discovery (GCD). Typically, this involves a t…
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
Wenbin An, Feng Tian, Jiahao Nie +7
Knowledge-based Visual Question Answering (KVQA) requires both image and world knowledge to answer questions. Current methods first retrieve knowledge from the image and external k…
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
Wenbin An, Feng Tian, Sicong Leng +6
Despite great success across various multimodal tasks, Large Vision-Language Models (LVLMs) often encounter object hallucinations with generated textual responses being inconsisten…