5 papers
PaddleOCR 3.0 Technical Report
Cheng Cui, Ting Sun, Manhui Lin +16
This technical report introduces PaddleOCR 3.0, an Apache-licensed open-source toolkit for OCR and document parsing. To address the growing demand for document understanding in the…
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
Feng Ni, Kui Huang, Yao Lu +4
With the rapid advancement of digitalization, various document images are being applied more extensively in production and daily life, and there is an increasingly urgent need for…
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
Kui Huang, Xinrong Chen, Wenyu Lv +3
This report introduces PP-DocBee2, an advanced version of the PP-DocBee, designed to enhance multimodal document understanding. Built on a large multimodal model architecture, PP-D…
Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information
Xu Chu, Xinrong Chen, Guanyu Wang +5
Inference time scaling drives extended reasoning to enhance the performance of Vision-Language Models (VLMs), thus forming powerful Vision-Language Reasoning Models (VLRMs). Howeve…
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer
Wenyu Lv, Yian Zhao, Qinyao Chang +3
In this report, we present RT-DETRv2, an improved Real-Time DEtection TRansformer (RT-DETR). RT-DETRv2 builds upon the previous state-of-the-art real-time detector, RT-DETR, and op…