4 papers
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
Wenjie Wang, Wei Wu, Ying Liu +10
Medical document OCR is challenging due to complex layouts, domain-specific terminology, and noisy annotations, while requiring strict field-level exact matching. Existing OCR syst…
Team PA-VCG's Solution for Competition on Understanding Chinese College Entrance Exam Papers in ICDAR'25
Wei Wu, Wenjie Wang, Yang Tan +6
This report presents Team PA-VGG's solution for the ICDAR'25 Competition on Understanding Chinese College Entrance Exam Papers. In addition to leveraging high-resolution image proc…
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
Shuai Li, Jian Xu, Xiao-Hui Li +2
Recent advances in Multi-modal Large Language Models (MLLMs) have shown significant progress in open-world Visual Question Answering (VQA). However, integrating visual information…
SuperNeRF-GAN: A Universal 3D-Consistent Super-Resolution Framework for Efficient and Enhanced 3D-Aware Image Synthesis
Peng Zheng, Linzhi Huang, Yizhou Yu +3
Neural volume rendering techniques, such as NeRF, have revolutionized 3D-aware image synthesis by enabling the generation of images of a single scene or object from various camera…