Publications (28)
RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild
Cheng Cui, Tingquan Gao, Xueqing Wang +11
Accurate document layout analysis remains a critical bottleneck for document parsing systems, due to the intricate coupling among heterogeneous document layout elements, geometric…
WebCanvas: Benchmarking Web Agents in Online Environments
Yichen Pan, Dehan Kong, Sida Zhou +8
For web agents to be practically useful, they must adapt to the continuously evolving web environment characterized by frequent updates to user interfaces and content. However, mos…
PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices
Guanghua Yu, Qinyao Chang, Wenyu Lv +12
The better accuracy and efficiency trade-off has been a challenging problem in object detection. In this work, we are dedicated to studying key optimizations and neural network arc…
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
Changda Zhou, Ziyue Gao, Xueqing Wang +4
While Vision-Language Models (VLMs) achieve near-perfect scores on digital document benchmarks like OmniDocBench, their performance in the unpredictable physical world remains larg…
PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System
Yuning Du, Chenxia Li, Ruoyu Guo +9
Optical Character Recognition (OCR) systems have been widely used in various of application scenarios. Designing an OCR system is still a challenging task. In previous work, we pro…
PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training
Zelun Zhang, Hongen Liu, Suyin Liang +12
We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. Although PaddleOCR-VL-1.5 establishes a strong 0.9B baseline, its remaining e…
Beyond Self-Supervision: A Simple Yet Effective Network Distillation Alternative to Improve Backbones
Cheng Cui, Ruoyu Guo, Yuning Du +10
Recently, research efforts have been concentrated on revealing how pre-trained model makes a difference in neural network performance. Self-supervision and semi-supervised learning…
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
Cheng Cui, Ting Sun, Suyin Liang +12
We introduce PaddleOCR-VL-1.5, an upgraded model achieving a new state-of-the-art (SOTA) accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real-worl…
PaddleOCR 3.0 Technical Report
Cheng Cui, Ting Sun, Manhui Lin +16
This technical report introduces PaddleOCR 3.0, an Apache-licensed open-source toolkit for OCR and document parsing. To address the growing demand for document understanding in the…
Huawei Cloud Model-as-a-Service on the CloudMatrix384 SuperPod
Ao Xiao, Bangzheng He, Baoquan Zhang +125
Scaled-out MoE LLMs and scaled-up SuperPods create new systems challenges for production Model-as-a-Service (MaaS), requiring disaggregation, low-latency communication, and decentr…
PP-FormulaNet: Bridging Accuracy and Efficiency in Advanced Formula Recognition
Hongen Liu, Cheng Cui, Yuning Du +2
Formula recognition is an important task in document intelligence. It involves converting mathematical expressions from document images into structured symbolic formats that comput…
2nd Place and 2nd Place Solution to Kaggle Landmark Recognition andRetrieval Competition 2019
Kaibing Chen, Cheng Cui, Yuning Du +2
We present a retrieval based system for landmark retrieval and recognition challenge.There are five parts in retrieval competition system, including feature extraction and matching…
HS-ResNet: Hierarchical-Split Block on Convolutional Neural Network
Pengcheng Yuan, Shufei Lin, Cheng Cui +5
This paper addresses representational block named Hierarchical-Split Block, which can be taken as a plug-and-play block to upgrade existing convolutional neural networks, improves…
Semi-Supervised Recognition under a Noisy and Fine-grained Dataset
Cheng Cui, Zhi Ye, Yangxi Li +7
Simi-Supervised Recognition Challenge-FGVC7 is a challenging fine-grained recognition competition. One of the difficulties of this competition is how to use unlabeled data. We adop…
2nd Place Solution in Google AI Open Images Object Detection Track 2019
Ruoyu Guo, Cheng Cui, Yuning Du +6
We present an object detection framework based on PaddlePaddle. We put all the strategies together (multi-scale training, FPN, Cascade, Dcnv2, Non-local, libra loss) based on ResNe…
Cast: Automated Resilience Testing for Production Cloud Service Systems
Zhuangbin Chen, Zhiling Deng, Kaiming Zhang +4
The distributed nature of microservice architecture introduces significant resilience challenges. Traditional testing methods, limited by extensive manual effort and oversimplified…
GLAD: Grounded Layered Autonomous Driving for Complex Service Tasks
Yan Ding, Cheng Cui, Xiaohan Zhang +1
Given the current point-to-point navigation capabilities of autonomous vehicles, researchers are looking into complex service requests that require the vehicles to visit multiple p…
PP-ShiTu: A Practical Lightweight Image Recognition System
Shengyu Wei, Ruoyu Guo, Cheng Cui +10
In recent years, image recognition applications have developed rapidly. A large number of studies and techniques have emerged in different fields, such as face recognition, pedestr…
PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks
Yubo Zhang, Xueqing Wang, Manhui Lin +13
Vision-Language Models (VLMs) have achieved impressive results on general vision-language tasks, yet they suffer from hallucination, imprecise localization, and prohibitive computa…
2nd Place Solution to Google Landmark Retrieval 2020
Min Yang, Cheng Cui, Xuetong Xue +2
This paper presents the 2nd place solution to the Google Landmark Retrieval Competition 2020. We propose a training method of global feature model for landmark retrieval without po…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
Cheng Cui, Ting Sun, Suyin Liang +15
In this report, we propose PaddleOCR-VL, a SOTA and resource-efficient model tailored for document parsing. Its core component is PaddleOCR-VL-0.9B, a compact yet powerful vision-l…
PP-YOLOE: An evolved version of YOLO
Shangliang Xu, Xinxin Wang, Wenyu Lv +8
In this report, we present PP-YOLOE, an industrial state-of-the-art object detector with high performance and friendly deployment. We optimize on the basis of the previous PP-YOLOv…
PP-LCNet: A Lightweight CPU Convolutional Neural Network
Cheng Cui, Tingquan Gao, Shengyu Wei +10
We propose a lightweight CPU network based on the MKLDNN acceleration strategy, named PP-LCNet, which improves the performance of lightweight models on multiple tasks. This paper l…
HPD-Parsing: Hierarchical Parallel Document Parsing
Shu Wei, Jingjing Wu, Lingshu Zhang +10
Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers…
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
Cheng Cui, Ting Sun, Suyin Liang +15
Document parsing is a fine-grained task where image resolution significantly impacts performance. While advanced research leveraging vision-language models benefits from high-resol…
PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks
Cheng Cui, Yubo Zhang, Ting Sun +11
The advent of "OCR 2.0" and large-scale vision-language models (VLMs) has set new benchmarks in text recognition. However, these unified architectures often come with significant c…
PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction
Ting Sun, Cheng Cui, Yuning Du +1
Document layout analysis is a critical preprocessing step in document intelligence, enabling the detection and localization of structural elements such as titles, text blocks, tabl…