8 papers
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
Gengluo Li, Pengyuan Lyu, Chengquan Zhang +7
Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Traditional cascaded pipelines depend…
Revealing the Learning Dynamics of Long-Context Continual Pre-training
Yupu Liang, Shuang Chen, Guanwei Zhang +2
Existing studies on Long-Context Continual Pre-training (LCCP) mainly focus on small-scale models and limited data regimes (tens of billions of tokens). We argue that directly migr…
MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
Gengluo Li, Chengquan Zhang, Yupu Liang +9
End-to-end text-image machine translation (TIMT), which directly translates textual content in images across languages, is crucial for real-world multilingual scene understanding.…
ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
Yaping Zhang, Yupu Liang, Zhiyang Zhang +5
Document Image Machine Translation (DIMT) seeks to translate text embedded in document images from one language to another by jointly modeling both textual content and page layout,…
HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery
Yaping Zhang, Qixuan Zhang, Xingquan Zhang +8
The rapid advancement of large language models (LLMs) and multimodal foundation models has sparked growing interest in their potential for scientific research. However, scientific…
Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency
Yupu Liang, Yaping Zhang, Zhiyang Zhang +5
Multimodal Large Language Models (MLLMs) have shown strong performance in document image tasks, especially Optical Character Recognition (OCR). However, they struggle with Document…