activity
20242026
collaborators

6 papers

cs.CV2026

Multimodal Fusion via Self-Consistent Task-Gradient Fields

Jiayu Xiong, Jing Wang, Jun Xue +4

Multimodal learning aims to preserve as much task-related information as possible from different inputs. However, current fusion designs often distort the feedback loop to feature…

cs.AI2025

Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown

Bowen Wang, Zhouqiang Jiang, Yasuaki Susumu +3

The real value of knowledge lies not just in its accumulation, but in its potential to be harnessed effectively to conquer the unknown. Although recent multimodal large language mo…

cs.CL2025

DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models

Bowen Wang, Jiuyang Chang, Yiming Qian +6

Large language models (LLMs) have recently showcased remarkable capabilities, spanning a wide range of tasks and applications, including those in the medical domain. Models like GP…

cs.CV2025

Exploring Visual Prompting: Robustness Inheritance and Beyond

Qi Li, Liangzhi Li, Zhouqiang Jiang +2

Visual Prompting (VP), an efficient method for transfer learning, has shown its potential in vision tasks. However, previous works focus exclusively on VP from standard source mode…

cs.CL2025

Putting People in LLMs' Shoes: Generating Better Answers via Question Rewriter

Junhao Chen, Bowen Wang, Zhouqiang Jiang +1

Large Language Models (LLMs) have demonstrated significant capabilities, particularly in the domain of question answering (QA). However, their effectiveness in QA is often undermin…

cs.CV2024

ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training

Zhouqiang Jiang, Bowen Wang, Junhao Chen +1

Recent approaches for visually-rich document understanding (VrDU) uses manually annotated semantic groups, where a semantic group encompasses all semantically relevant but not obvi…