6 papers
Multimodal Fusion via Self-Consistent Task-Gradient Fields
Jiayu Xiong, Jing Wang, Jun Xue +4
Multimodal learning aims to preserve as much task-related information as possible from different inputs. However, current fusion designs often distort the feedback loop to feature…
Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown
Bowen Wang, Zhouqiang Jiang, Yasuaki Susumu +3
The real value of knowledge lies not just in its accumulation, but in its potential to be harnessed effectively to conquer the unknown. Although recent multimodal large language mo…
DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models
Bowen Wang, Jiuyang Chang, Yiming Qian +6
Large language models (LLMs) have recently showcased remarkable capabilities, spanning a wide range of tasks and applications, including those in the medical domain. Models like GP…
Exploring Visual Prompting: Robustness Inheritance and Beyond
Qi Li, Liangzhi Li, Zhouqiang Jiang +2
Visual Prompting (VP), an efficient method for transfer learning, has shown its potential in vision tasks. However, previous works focus exclusively on VP from standard source mode…
Putting People in LLMs' Shoes: Generating Better Answers via Question Rewriter
Junhao Chen, Bowen Wang, Zhouqiang Jiang +1
Large Language Models (LLMs) have demonstrated significant capabilities, particularly in the domain of question answering (QA). However, their effectiveness in QA is often undermin…
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training
Zhouqiang Jiang, Bowen Wang, Junhao Chen +1
Recent approaches for visually-rich document understanding (VrDU) uses manually annotated semantic groups, where a semantic group encompasses all semantically relevant but not obvi…