5 papers
AlphaOracle: Oracle bone script decipherment via human-workflow-inspired deep learning
Yuliang Liu, Haisu Guan, Pengjie Wang +11
Approximately 3,000 of the 4,500 oracle bone script (OBS) characters remain undeciphered due to fragmentary inscriptions and sparse evidence. Current AI approaches fail to replicat…
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI
Yuliang Liu, Zhang Li, Ziyang Zhang +11
Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-graine…
Multimodal OCR: Parse Anything from Documents
Handong Zheng, Yumeng Li, Kaile Zhang +22
We present Multimodal OCR (MOCR), a document parsing paradigm that jointly parses text and graphics into unified textual representations. Unlike conventional OCR systems that focus…
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
Cheng Cui, Ting Sun, Suyin Liang +15
In this report, we propose PaddleOCR-VL, a SOTA and resource-efficient model tailored for document parsing. Its core component is PaddleOCR-VL-0.9B, a compact yet powerful vision-l…
VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
Hao Yan, Xingchen Liu, Hao Wang +11
Recent strides in multimodal large language models (MLLMs) have significantly advanced their performance in many reasoning tasks. However, Abstract Visual Reasoning (AVR) remains a…