2 papers
cs.CV2026
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
Daxiang Dong, Mingming Zheng, Dong Xu +17
We present Qianfan-OCR, a 4B-parameter end-to-end vision-language model that unifies document parsing, layout analysis, and document understanding within a single architecture. It…
cs.CV2025
Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
Mi Zheng, Guanglei Yang, Zitong Huang +3
With the emergence of transformer-based architectures and large language models (LLMs), the accuracy of road scene perception has substantially advanced. Nonetheless, current road…