3 papers
cs.CV2026
Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding
Haoyu Cao, Kun Yin, Yunfei Wu +16
This paper presents Youtu-Parsing, an efficient and versatile document parsing model designed for high-performance content extraction. The architecture employs a native Vision Tran…
cs.CV2025
DREAM: Document Reconstruction via End-to-end Autoregressive Model
Xin Li, Mingming Gong, Yunfei Wu +7
Document reconstruction constitutes a significant facet of document analysis and recognition, a field that has been progressively accruing interest within the scholarly community.…
cs.CV2024
Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language Models
Xin Li, Yunfei Wu, Xinghua Jiang +6
Recently, the advent of Large Visual-Language Models (LVLMs) has received increasing attention across various domains, particularly in the field of visual document understanding (V…