collaborators

5 papers

cs.CV2025

See then Tell: Enhancing Key Information Extraction with Vision Grounding

Shuhang Liu, Zhenrong Zhang, Pengfei Hu +5

In the digital era, the ability to understand visually rich documents that integrate text, complex layouts, and imagery is critical. Traditional Key Information Extraction (KIE) me…

cs.MM2025

MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique

Shuhang Liu, Zhenrong Zhang, Pengfei Hu +7

Visual language models (VLMs) have demonstrated strong performance across diverse multimodal reasoning tasks but still face challenges such as hallucinations, resulting in incorrec…

cs.MM2025

PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search

Pengfei Hu, Zhenrong Zhang, Qikai Chang +8

Recent work increasingly focuses on improving the reasoning capabilities of Multimodal Large Language Models (MLLMs). Among existing methods, Process Reward Models (PRMs) stand out…

cs.CL2025

DocMamba: Efficient Document Pre-training with State Space Model

Pengfei Hu, Zhenrong Zhang, Jiefeng Ma +3

In recent years, visually-rich document understanding has attracted increasing attention. Transformer-based pre-trained models have become the mainstream approach, yielding signifi…

cs.CV2024

UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition

Zhenrong Zhang, Shuhang Liu, Pengfei Hu +4

In the digital era, table structure recognition technology is a critical tool for processing and analyzing large volumes of tabular data. Previous methods primarily focus on visual…