1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.CV2024★ 1 cited
FreeVA: Offline MLLM as Training-Free Video Assistant
Wenhao Wu
This paper undertakes an empirical study to revisit the latest advancements in Multimodal Large Language Models (MLLMs): Video Assistant. This study, namely FreeVA, aims to extend…
cs.CV2024
Automated Multi-level Preference for MLLMs
Mengxi Zhang, Wenhao Wu, Yu Lu +8
Current multimodal Large Language Models (MLLMs) suffer from ``hallucination'', occasionally generating responses that are not grounded in the input images. To tackle this challeng…
cs.CV2024
Dense Connector for MLLMs
Huanjin Yao, Wenhao Wu, Taojiannan Yang +7
Do we fully leverage the potential of visual encoder in Multimodal Large Language Models (MLLMs)? The recent outstanding performance of MLLMs in multimodal understanding has garner…