5 papers
See then Tell: Enhancing Key Information Extraction with Vision Grounding
Shuhang Liu, Zhenrong Zhang, Pengfei Hu +5
In the digital era, the ability to understand visually rich documents that integrate text, complex layouts, and imagery is critical. Traditional Key Information Extraction (KIE) me…
MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
Shuhang Liu, Zhenrong Zhang, Pengfei Hu +7
Visual language models (VLMs) have demonstrated strong performance across diverse multimodal reasoning tasks but still face challenges such as hallucinations, resulting in incorrec…
PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search
Pengfei Hu, Zhenrong Zhang, Qikai Chang +8
Recent work increasingly focuses on improving the reasoning capabilities of Multimodal Large Language Models (MLLMs). Among existing methods, Process Reward Models (PRMs) stand out…
DocMamba: Efficient Document Pre-training with State Space Model
Pengfei Hu, Zhenrong Zhang, Jiefeng Ma +3
In recent years, visually-rich document understanding has attracted increasing attention. Transformer-based pre-trained models have become the mainstream approach, yielding signifi…
UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition
Zhenrong Zhang, Shuhang Liu, Pengfei Hu +4
In the digital era, table structure recognition technology is a critical tool for processing and analyzing large volumes of tabular data. Previous methods primarily focus on visual…