6 papers
AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning
Junxiao Xue, Quan Deng, Tingqi Hu +4
Despite recent progress in multimodal large language models (MLLMs), reliable visual question answering in aerial scenes remains challenging. In such scenes, task-critical evidence…
Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning
Wenxi Gao, Guanxi Lu, Didi Zhu +5
Unified multimodal models (UMMs) with interleaved reasoning, which generate both textual and visual steps as part of intermediate reasoning traces, have demonstrated great potentia…
Towards Comprehensive Interactive Change Understanding in Remote Sensing: A Large-scale Dataset and Dual-granularity Enhanced VLM
Junxiao Xue, Quan Deng, Xuecheng Wu +7
Remote sensing change understanding (RSCU) is essential for analyzing remote sensing images and understanding how human activities affect the environment. However, existing dataset…
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
Fei Yu, Quan Deng, Shengeng Tang +2
Understanding 3D scenes in open-world settings poses fundamental challenges for vision and robotics, particularly due to the limitations of closed-vocabulary supervision and static…
Scene Understanding Enabled Semantic Communication with Open Channel Coding
Zhe Xiang, Fei Yu, Quan Deng +2
As communication systems transition from symbol transmission to conveying meaningful information, sixth-generation (6G) networks emphasize semantic communication. This approach pri…
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
Junxiao Xue, Quan Deng, Fei Yu +3
Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tas…