activity
20242026
most citedUniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing 2025Show all

5 papers · 1 filter

cs.AI2025

THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning

Qikai Chang, Zhenrong Zhang, Pengfei Hu +6

Large Language Models (LLMs) have made remarkable progress in mathematical reasoning, but still continue to struggle with high-precision tasks like numerical computation and formal…

cs.CL2025

Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration

Yicheng Pan, Zhenrong Zhang, Pengfei Hu +6

Recent advances in Multimodal Large Language Models (MLLMs) have achieved remarkable progress in general domains and demonstrated promise in multimodal mathematical reasoning. Howe…

cs.MM2025

MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique

Shuhang Liu, Zhenrong Zhang, Pengfei Hu +7

Visual language models (VLMs) have demonstrated strong performance across diverse multimodal reasoning tasks but still face challenges such as hallucinations, resulting in incorrec…

cs.MM2025

PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search

Pengfei Hu, Zhenrong Zhang, Qikai Chang +8

Recent work increasingly focuses on improving the reasoning capabilities of Multimodal Large Language Models (MLLMs). Among existing methods, Process Reward Models (PRMs) stand out…

cs.CV2025

Skeleton and Font Generation Network for Zero-shot Chinese Character Generation

Mobai Xue, Jun Du, Zhenrong Zhang +5

Automatic font generation remains a challenging research issue, primarily due to the vast number of Chinese characters, each with unique and intricate structures. Our investigation…