10 citations · 12 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 10 cited
VL-Mamba: Exploring State Space Models for Multimodal Learning
Yanyuan Qiao, Zheng Yu, Longteng Guo +5
Multimodal large language models (MLLMs) have attracted widespread interest and have rich applications. However, the inherent attention mechanism in its Transformer structure requi…
cs.CV2024★ 2 cited
Knowledge Condensation and Reasoning for Knowledge-based VQA
Dongze Hao, Jian Jia, Longteng Guo +8
Knowledge-based visual question answering (KB-VQA) is a challenging task, which requires the model to leverage external knowledge for comprehending and answering questions grounded…