1 citations · 1 across the 9 of their papers we have counts for
Showing 2026 · cs.CVShow all
3 papers · 2 filters
cs.CV2026
ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering
Zhentao Guo, Chen Duan, Tongkun Guan +3
Despite remarkable progress in multimodal understanding, current MLLMs still exhibit limitations in video text understanding, particularly when semantics emerge through the integra…
cs.CV2026
InstructTable: Improving Table Structure Recognition Through Instructions
Boming Chen, Zining Wang, Zhentao Guo +5
Table structure recognition (TSR) holds widespread practical importance by parsing tabular images into structured representations, yet encounters significant challenges when proces…
cs.CV2026
PositionOCR: Augmenting Positional Awareness in Multi-Modal Models via Hybrid Specialist Integration
Chen Duan, Zhentao Guo, Pei Fu +3
In recent years, Multi-modal Large Language Models (MLLMs) have achieved strong performance in OCR-centric Visual Question Answering (VQA) tasks, illustrating their capability to p…