collaborators

6 papers

cs.CL2026

AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking

Zining Wang, Tongkun Guan, Boming Chen +8

Multimodal large language models have demonstrated strong document reasoning capabilities by incorporating explicit thinking processes. While this capability significantly improves…

cs.CV2026

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering

Zhentao Guo, Chen Duan, Tongkun Guan +3

Despite remarkable progress in multimodal understanding, current MLLMs still exhibit limitations in video text understanding, particularly when semantics emerge through the integra…

cs.CV2026

InstructTable: Improving Table Structure Recognition Through Instructions

Boming Chen, Zining Wang, Zhentao Guo +5

Table structure recognition (TSR) holds widespread practical importance by parsing tabular images into structured representations, yet encounters significant challenges when proces…

cs.CV2026

PositionOCR: Augmenting Positional Awareness in Multi-Modal Models via Hybrid Specialist Integration

Chen Duan, Zhentao Guo, Pei Fu +3

In recent years, Multi-modal Large Language Models (MLLMs) have achieved strong performance in OCR-centric Visual Question Answering (VQA) tasks, illustrating their capability to p…

cs.CV2025

Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding

Zining Wang, Tongkun Guan, Pei Fu +7

Multi-modal Large Language Models (MLLMs) have introduced a novel dimension to document understanding, i.e., they endow large language models with visual comprehension capabilities…

cs.CV2025

Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review

Pei Fu, Tongkun Guan, Zining Wang +8

The recent emergence of Multi-modal Large Language Models (MLLMs) has introduced a new dimension to the Text-rich Image Understanding (TIU) field, with models demonstrating impress…