2 citations · 2 across the 3 of their papers we have counts for
5 papers · 1 filter
Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning
Wenxi Gao, Guanxi Lu, Didi Zhu +5
Unified multimodal models (UMMs) with interleaved reasoning, which generate both textual and visual steps as part of intermediate reasoning traces, have demonstrated great potentia…
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
Junxiao Xue, Quan Deng, Tingqi Hu +4
Despite recent progress in multimodal large language models (MLLMs), reliable visual question answering in aerial scenes remains challenging. In such scenes, task-critical evidence…
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
Fei Yu, Quan Deng, Shengeng Tang +2
Understanding 3D scenes in open-world settings poses fundamental challenges for vision and robotics, particularly due to the limitations of closed-vocabulary supervision and static…
Towards Comprehensive Interactive Change Understanding in Remote Sensing: A Large-scale Dataset and Dual-granularity Enhanced VLM
Junxiao Xue, Quan Deng, Xuecheng Wu +7
Remote sensing change understanding (RSCU) is essential for analyzing remote sensing images and understanding how human activities affect the environment. However, existing dataset…
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
Junxiao Xue, Quan Deng, Fei Yu +3
Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tas…