3 papers
cs.CV2024
Automated Multi-level Preference for MLLMs
Mengxi Zhang, Wenhao Wu, Yu Lu +8
Current multimodal Large Language Models (MLLMs) suffer from ``hallucination'', occasionally generating responses that are not grounded in the input images. To tackle this challeng…
cs.CV2024
Dense Connector for MLLMs
Huanjin Yao, Wenhao Wu, Taojiannan Yang +7
Do we fully leverage the potential of visual encoder in Multimodal Large Language Models (MLLMs)? The recent outstanding performance of MLLMs in multimodal understanding has garner…
cs.CV2023
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
Wenhao Wu, Huanjin Yao, Mengxi Zhang +3
This paper does not present a novel method. Instead, it delves into an essential, yet must-know baseline in light of the latest advancements in Generative Artificial Intelligence (…