6 papers
Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis
Zhipeng Xu, Zulong Chen, Qing Liu +6
Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-s…
CC-OCR V2: Fine-Grained Attribution of LMM Failures in Real-World Visual Document Understanding
Zhipeng Xu, Junhao Ji, Yuqi Xiong +13
Recent Large Multimodal Models (LMMs) have achieved remarkable progress on OCR-centric document understanding and processing tasks. Existing benchmarks primarily evaluate LMMs acro…
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
Wenwen Yu, Zhibo Yang, Jianqiang Wan +5
Visually-situated text parsing (VsTP) has recently seen notable advancements, driven by the growing demand for automated document understanding and the emergence of large language…
Qwen3-VL Technical Report
Shuai Bai, Yuxuan Cai, Ruizhe Chen +61
We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively…
Qwen2.5-VL Technical Report
Shuai Bai, Keqin Chen, Xuejing Liu +24
We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational capabilities and innovative func…
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
Zhibo Yang, Jun Tang, Zhaohai Li +9
Large Multimodal Models (LMMs) have demonstrated impressive performance in recognizing document images with natural language instructions. However, it remains unclear to what exten…