collaborators

6 papers

cs.AI2026

REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot Alignment

Kai Ye, Xianwei Mao, Sheng Zhou +6

Knowledge-intensive Visual Question Answering (KI-VQA) frequently suffers from severe knowledge conflicts caused by the inherent limitations of open-domain retrieval. However, exis…

cs.AI2026

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning

Ye Mo, Kai Ye, Xianwei Mao +9

Document understanding aims to perform question answering and information extraction over document images, where the visual content is highly information-dense and most queries rel…

cs.AI2025

Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding

Zirui Shao, Feiyu Gao, Zhaoqing Zhu +6

Multimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand. As a mult…

cs.CL2025

BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks

Tianyuan Huang, Zepeng Zhu, Hangdi Xing +6

Braille plays a vital role in education and information accessibility for visually impaired individuals. However, Braille information processing faces challenges such as data scarc…

cs.CV2025

A Simple yet Effective Layout Token in Large Language Models for Document Understanding

Zhaoqing Zhu, Chuwei Luo, Zirui Shao +4

Recent methods that integrate spatial layouts with text for document understanding in large language models (LLMs) have shown promising results. A commonly used method is to repres…

cs.CV2025

MP-GUI: Modality Perception with MLLMs for GUI Understanding

Ziwei Wang, Weizhi Chen, Leyang Yang +7

Graphical user interface (GUI) has become integral to modern society, making it crucial to be understood for human-centric systems. However, unlike natural images or documents, GUI…