papers

Publications (39)

cs.CV2022

PP-ShiTu: A Practical Lightweight Image Recognition System

Shengyu Wei, Ruoyu Guo, Cheng Cui +10

In recent years, image recognition applications have developed rapidly. A large number of studies and techniques have emerged in different fields, such as face recognition, pedestr…

cs.CL2026

ERNIE 5.0 Technical Report

Haifeng Wang, Hua Wu, Tian Wu +432

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…

cs.CV2025

PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model

Cheng Cui, Ting Sun, Suyin Liang +15

In this report, we propose PaddleOCR-VL, a SOTA and resource-efficient model tailored for document parsing. Its core component is PaddleOCR-VL-0.9B, a compact yet powerful vision-l…

quant-ph2008

Intrinsic Quantum Noise in Faraday Rotation Measurements of a Single Electron Spin

Yanjun Ma, Jeremy Levy

Faraday rotation is one way to realize quantum non-demolition measurement of electron spin in quantum dots. To describe Faraday rotation, semiclassical models are typically used, b…

cs.LG2025

SDWPF: A Dataset for Spatial Dynamic Wind Power Forecasting Challenge at KDD Cup 2022

Jingbo Zhou, Xinjiang Lu, Yixiong Xiao +4

The variability of wind power supply can present substantial challenges to incorporating wind power into a grid system. Thus, Wind Power Forecasting (WPF) has been widely recognize…

quant-ph2025

Efficient Implementation of Arbitrary Two-Qubit Gates via Unified Control

Zhen Chen, Weiyang Liu, Yanjun Ma +16

The native gate set is fundamental to the performance of quantum devices, as it governs the accuracy of basic quantum operations and dictates the complexity of implementing quantum…

cs.LG2025

GraphNet: A Large-Scale Computational Graph Dataset for Tensor Compiler Research

Xinqi Li, Yiqun Liu, Shan Jiang +6

We introduce GraphNet, a dataset of 2.7K real-world deep learning computational graphs with rich metadata, spanning six major task categories across multiple deep learning framewor…

cs.CV2021

PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System

Yuning Du, Chenxia Li, Ruoyu Guo +9

Optical Character Recognition (OCR) systems have been widely used in various of application scenarios. Designing an OCR system is still a challenging task. In previous work, we pro…

cs.LG2026

AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping

Guoxia Wang, Shuai Li, Congliang Chen +5

Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous research has attempted to identify the root cause of loss spikes by investigating…

cs.CV2021

Beyond Self-Supervision: A Simple Yet Effective Network Distillation Alternative to Improve Backbones

Cheng Cui, Ruoyu Guo, Yuning Du +10

Recently, research efforts have been concentrated on revealing how pre-trained model makes a difference in neural network performance. Self-supervision and semi-supervised learning…

cs.CV2026

PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing

Cheng Cui, Ting Sun, Suyin Liang +12

We introduce PaddleOCR-VL-1.5, an upgraded model achieving a new state-of-the-art (SOTA) accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real-worl…

cs.DC2022

Boosting Distributed Training Performance of the Unpadded BERT Model

Jinle Zeng, Min Li, Zhihua Wu +4

Pre-training models are an important tool in Natural Language Processing (NLP), while the BERT model is a classic pre-training model whose structure has been widely adopted by foll…

cs.DC2022

HelixFold: An Efficient Implementation of AlphaFold2 using PaddlePaddle

Guoxia Wang, Xiaomin Fang, Zhihua Wu +6

Accurate protein structure prediction can significantly accelerate the development of life science. The accuracy of AlphaFold2, a frontier end-to-end structure prediction system, i…

cond-mat.mtrl-sci2021

Engineering of Ferroic Orders in Thin Films by Anionic Substitution

A. C. Garcia-Castro, Yanjun Ma, Zachary Romestan +3

Multiferroics are a unique class of materials where magnetic and ferroelectric orders coexist. The research on multiferroics contributes significantly to the fundamental understand…

cond-mat.mtrl-sci2019

Realization of Epitaxial Thin Films of the Topological Crystalline Insulator SrSnO

Yanjun Ma, Anthony Edgeton, Hanjong Paik +8

Topological materials are derived from the interplay between symmetry and topology. Advances in topological band theories have led to the prediction that the antiperovskite oxide S…

eess.AS2022

PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit

Hui Zhang, Tian Yuan, Junkun Chen +10

PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command…

cs.CV2022

PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System

Chenxia Li, Weiwei Liu, Ruoyu Guo +9

Optical character recognition (OCR) technology has been widely used in various scenes, as shown in Figure 1. Designing a practical OCR system is still a meaningful but challenging…

cond-mat.str-el2015

Ultrafast observation of electron hybridization and in-gap states formation in Kondo insulator SmB6

Sanjay Adhikari, Yanjun Ma, Zachary Fisk +3

SmB6 is a promising candidate for topological Kondo insulator. In this letter, we report ultrafast carrier dynamics of SmB6. Two characteristic temperatures: T1=100 K and T2= 20 K…

cs.DC2023

HeterPS: Distributed Deep Learning With Reinforcement Learning Based Scheduling in Heterogeneous Environments

Ji Liu, Zhihua Wu, Dianhai Yu +6

Deep neural networks (DNNs) exploit many layers and a large number of parameters to achieve excellent performance. The training process of DNN models generally handles large-scale…

cs.CV2026

PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks

Yubo Zhang, Xueqing Wang, Manhui Lin +13

Vision-Language Models (VLMs) have achieved impressive results on general vision-language tasks, yet they suffer from hallucination, imprecise localization, and prohibitive computa…

cs.CV2021

PP-LCNet: A Lightweight CPU Convolutional Neural Network

Cheng Cui, Tingquan Gao, Shengyu Wei +10

We propose a lightweight CPU network based on the MKLDNN acceleration strategy, named PP-LCNet, which improves the performance of lightweight models on multiple tasks. This paper l…

cs.DC2021

End-to-end Adaptive Distributed Training on PaddlePaddle

Yulong Ao, Zhihua Wu, Dianhai Yu +10

Distributed training has become a pervasive and effective approach for training a large neural network (NN) model with processing massive data. However, it is very challenging to s…

cs.CV2026

Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing

Cheng Cui, Ting Sun, Suyin Liang +15

Document parsing is a fine-grained task where image resolution significantly impacts performance. While advanced research leveraging vision-language models benefits from high-resol…

cs.CV2021

PP-YOLOv2: A Practical Object Detector

Xin Huang, Xinxin Wang, Wenyu Lv +10

Being effective and efficient is essential to an object detector for practical use. To meet these two concerns, we comprehensively evaluate a collection of existing refinements to…

cs.IT2012

On the Achievability of Interference Alignment for Three-Cell Constant Cellular Interfering Networks

Yanjun Ma, Jiandong Li, Rui Chen +1

For a three-cell constant cellular interfering network, a new property of alignment is identified, i.e., interference alignment (IA) solution obtained in an user-cooperation scenar…

cond-mat.mtrl-sci2023

Room-temperature ferromagnetism in epitaxial bilayer FeSb/SrTiO3(001) terminated with a Kagome lattice

Huimin Zhang, Qinxi Liu, Liangzi Deng +10

Two-dimensional (2D) magnets exhibit unique physical properties for potential applications in spintronics. To date, most 2D ferromagnets are obtained by mechanical exfoliation of b…

cs.DC2022

Efficient AlphaFold2 Training using Parallel Evoformer and Branch Parallelism

Guoxia Wang, Zhihua Wu, Xiaomin Fang +4

The accuracy of AlphaFold2, a frontier end-to-end structure prediction system, is already close to that of the experimental determination techniques. Due to the complex model archi…

cs.LG2022

Nebula-I: A General Framework for Collaboratively Training Deep Learning Models on Low-Bandwidth Cloud Clusters

Yang Xiang, Zhihua Wu, Weibao Gong +15

The ever-growing model size and scale of compute have attracted increasing interests in training deep learning models over multiple nodes. However, when it comes to training on clo…

cs.IT2010

Group Based Interference Alignment

Yanjun Ma, Jiandong Li, Qin Liu +1

In the -user single-input single-output (SISO) frequency-selective fading interference channel, it is shown that the maximal achievable multiplexing gain is almost surely

cond-mat.mes-hall2010

Rewritable nanoscale oxide photodetector

Patrick Irvin, Yanjun Ma, Daniela F. Bogorin +5

Nanophotonic devices seek to generate, guide, and/or detect light using structures whose nanoscale dimensions are closely tied to their functionality. Semiconducting nanowires, gro…

cs.CV2021

PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices

Guanghua Yu, Qinyao Chang, Wenyu Lv +12

The better accuracy and efficiency trade-off has been a challenging problem in object detection. In this work, we are dedicated to studying key optimizations and neural network arc…

cs.IT2011

Distributed Interference Alignment with Low Overhead

Yanjun Ma, Jiandong Li, Rui Chen

Based on closed-form interference alignment (IA) solutions, a low overhead distributed interference alignment (LOIA) scheme is proposed in this paper for the -user SISO interfer…

cs.CV2026

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training

Zelun Zhang, Hongen Liu, Suyin Liang +12

We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. Although PaddleOCR-VL-1.5 establishes a strong 0.9B baseline, its remaining e…

quant-ph2026

Mitigating Measurement Crosstalk via Pulse Shaping

Yang Gao, Feiyu Li, Yang Liu +10

Quantum error correction protocols require rapid and repeated qubit measurements. While multiplexed readout in superconducting quantum systems improves efficiency, fast probe pulse…

cs.LG2025

CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs

Zhaojing Zhou, Xunchao Li, Minghao Li +8

The rapid scaling of Large Language Models (LLMs) elevates inference costs and compounds substantial deployment barriers. While quantization to 8 or 4 bits mitigates this, sub-3-bi…

cs.CV2025

PaddleOCR 3.0 Technical Report

Cheng Cui, Ting Sun, Manhui Lin +16

This technical report introduces PaddleOCR 3.0, an Apache-licensed open-source toolkit for OCR and document parsing. To address the growing demand for document understanding in the…

cs.CV2022

PP-LiteSeg: A Superior Real-Time Semantic Segmentation Model

Juncai Peng, Yi Liu, Shiyu Tang +13

Real-world applications have high demands for semantic segmentation methods. Although semantic segmentation has made remarkable leap-forwards with deep learning, the performance of…

cs.CV2021

PAFNet: An Efficient Anchor-Free Object Detector Guidance

Ying Xin, Guanzhong Wang, Mingyuan Mao +5

Object detection is a basic but challenging task in computer vision, which plays a key role in a variety of industrial applications. However, object detectors based on deep learnin…

cs.CL2021

ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation

Shuohuan Wang, Yu Sun, Yang Xiang +26

Pre-trained language models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks. GPT-3 has shown that scaling up pre-trained language models c…