Publications (39)
PP-ShiTu: A Practical Lightweight Image Recognition System
Shengyu Wei, Ruoyu Guo, Cheng Cui +10
In recent years, image recognition applications have developed rapidly. A large number of studies and techniques have emerged in different fields, such as face recognition, pedestr…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
Cheng Cui, Ting Sun, Suyin Liang +15
In this report, we propose PaddleOCR-VL, a SOTA and resource-efficient model tailored for document parsing. Its core component is PaddleOCR-VL-0.9B, a compact yet powerful vision-l…
Intrinsic Quantum Noise in Faraday Rotation Measurements of a Single Electron Spin
Yanjun Ma, Jeremy Levy
Faraday rotation is one way to realize quantum non-demolition measurement of electron spin in quantum dots. To describe Faraday rotation, semiclassical models are typically used, b…
SDWPF: A Dataset for Spatial Dynamic Wind Power Forecasting Challenge at KDD Cup 2022
Jingbo Zhou, Xinjiang Lu, Yixiong Xiao +4
The variability of wind power supply can present substantial challenges to incorporating wind power into a grid system. Thus, Wind Power Forecasting (WPF) has been widely recognize…
Efficient Implementation of Arbitrary Two-Qubit Gates via Unified Control
Zhen Chen, Weiyang Liu, Yanjun Ma +16
The native gate set is fundamental to the performance of quantum devices, as it governs the accuracy of basic quantum operations and dictates the complexity of implementing quantum…
GraphNet: A Large-Scale Computational Graph Dataset for Tensor Compiler Research
Xinqi Li, Yiqun Liu, Shan Jiang +6
We introduce GraphNet, a dataset of 2.7K real-world deep learning computational graphs with rich metadata, spanning six major task categories across multiple deep learning framewor…
PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System
Yuning Du, Chenxia Li, Ruoyu Guo +9
Optical Character Recognition (OCR) systems have been widely used in various of application scenarios. Designing an OCR system is still a challenging task. In previous work, we pro…
AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping
Guoxia Wang, Shuai Li, Congliang Chen +5
Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous research has attempted to identify the root cause of loss spikes by investigating…
Beyond Self-Supervision: A Simple Yet Effective Network Distillation Alternative to Improve Backbones
Cheng Cui, Ruoyu Guo, Yuning Du +10
Recently, research efforts have been concentrated on revealing how pre-trained model makes a difference in neural network performance. Self-supervision and semi-supervised learning…
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
Cheng Cui, Ting Sun, Suyin Liang +12
We introduce PaddleOCR-VL-1.5, an upgraded model achieving a new state-of-the-art (SOTA) accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real-worl…
Boosting Distributed Training Performance of the Unpadded BERT Model
Jinle Zeng, Min Li, Zhihua Wu +4
Pre-training models are an important tool in Natural Language Processing (NLP), while the BERT model is a classic pre-training model whose structure has been widely adopted by foll…
HelixFold: An Efficient Implementation of AlphaFold2 using PaddlePaddle
Guoxia Wang, Xiaomin Fang, Zhihua Wu +6
Accurate protein structure prediction can significantly accelerate the development of life science. The accuracy of AlphaFold2, a frontier end-to-end structure prediction system, i…
Engineering of Ferroic Orders in Thin Films by Anionic Substitution
A. C. Garcia-Castro, Yanjun Ma, Zachary Romestan +3
Multiferroics are a unique class of materials where magnetic and ferroelectric orders coexist. The research on multiferroics contributes significantly to the fundamental understand…
Realization of Epitaxial Thin Films of the Topological Crystalline Insulator SrSnO
Yanjun Ma, Anthony Edgeton, Hanjong Paik +8
Topological materials are derived from the interplay between symmetry and topology. Advances in topological band theories have led to the prediction that the antiperovskite oxide S…
PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit
Hui Zhang, Tian Yuan, Junkun Chen +10
PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command…
PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System
Chenxia Li, Weiwei Liu, Ruoyu Guo +9
Optical character recognition (OCR) technology has been widely used in various scenes, as shown in Figure 1. Designing a practical OCR system is still a meaningful but challenging…
Ultrafast observation of electron hybridization and in-gap states formation in Kondo insulator SmB6
Sanjay Adhikari, Yanjun Ma, Zachary Fisk +3
SmB6 is a promising candidate for topological Kondo insulator. In this letter, we report ultrafast carrier dynamics of SmB6. Two characteristic temperatures: T1=100 K and T2= 20 K…
HeterPS: Distributed Deep Learning With Reinforcement Learning Based Scheduling in Heterogeneous Environments
Ji Liu, Zhihua Wu, Dianhai Yu +6
Deep neural networks (DNNs) exploit many layers and a large number of parameters to achieve excellent performance. The training process of DNN models generally handles large-scale…
PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks
Yubo Zhang, Xueqing Wang, Manhui Lin +13
Vision-Language Models (VLMs) have achieved impressive results on general vision-language tasks, yet they suffer from hallucination, imprecise localization, and prohibitive computa…
PP-LCNet: A Lightweight CPU Convolutional Neural Network
Cheng Cui, Tingquan Gao, Shengyu Wei +10
We propose a lightweight CPU network based on the MKLDNN acceleration strategy, named PP-LCNet, which improves the performance of lightweight models on multiple tasks. This paper l…
End-to-end Adaptive Distributed Training on PaddlePaddle
Yulong Ao, Zhihua Wu, Dianhai Yu +10
Distributed training has become a pervasive and effective approach for training a large neural network (NN) model with processing massive data. However, it is very challenging to s…
Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing
Cheng Cui, Ting Sun, Suyin Liang +15
Document parsing is a fine-grained task where image resolution significantly impacts performance. While advanced research leveraging vision-language models benefits from high-resol…
PP-YOLOv2: A Practical Object Detector
Xin Huang, Xinxin Wang, Wenyu Lv +10
Being effective and efficient is essential to an object detector for practical use. To meet these two concerns, we comprehensively evaluate a collection of existing refinements to…
On the Achievability of Interference Alignment for Three-Cell Constant Cellular Interfering Networks
Yanjun Ma, Jiandong Li, Rui Chen +1
For a three-cell constant cellular interfering network, a new property of alignment is identified, i.e., interference alignment (IA) solution obtained in an user-cooperation scenar…
Room-temperature ferromagnetism in epitaxial bilayer FeSb/SrTiO3(001) terminated with a Kagome lattice
Huimin Zhang, Qinxi Liu, Liangzi Deng +10
Two-dimensional (2D) magnets exhibit unique physical properties for potential applications in spintronics. To date, most 2D ferromagnets are obtained by mechanical exfoliation of b…
Efficient AlphaFold2 Training using Parallel Evoformer and Branch Parallelism
Guoxia Wang, Zhihua Wu, Xiaomin Fang +4
The accuracy of AlphaFold2, a frontier end-to-end structure prediction system, is already close to that of the experimental determination techniques. Due to the complex model archi…
Nebula-I: A General Framework for Collaboratively Training Deep Learning Models on Low-Bandwidth Cloud Clusters
Yang Xiang, Zhihua Wu, Weibao Gong +15
The ever-growing model size and scale of compute have attracted increasing interests in training deep learning models over multiple nodes. However, when it comes to training on clo…
Group Based Interference Alignment
Yanjun Ma, Jiandong Li, Qin Liu +1
In the -user single-input single-output (SISO) frequency-selective fading interference channel, it is shown that the maximal achievable multiplexing gain is almost surely …
Rewritable nanoscale oxide photodetector
Patrick Irvin, Yanjun Ma, Daniela F. Bogorin +5
Nanophotonic devices seek to generate, guide, and/or detect light using structures whose nanoscale dimensions are closely tied to their functionality. Semiconducting nanowires, gro…
PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices
Guanghua Yu, Qinyao Chang, Wenyu Lv +12
The better accuracy and efficiency trade-off has been a challenging problem in object detection. In this work, we are dedicated to studying key optimizations and neural network arc…
Distributed Interference Alignment with Low Overhead
Yanjun Ma, Jiandong Li, Rui Chen
Based on closed-form interference alignment (IA) solutions, a low overhead distributed interference alignment (LOIA) scheme is proposed in this paper for the -user SISO interfer…
PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training
Zelun Zhang, Hongen Liu, Suyin Liang +12
We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. Although PaddleOCR-VL-1.5 establishes a strong 0.9B baseline, its remaining e…
Mitigating Measurement Crosstalk via Pulse Shaping
Yang Gao, Feiyu Li, Yang Liu +10
Quantum error correction protocols require rapid and repeated qubit measurements. While multiplexed readout in superconducting quantum systems improves efficiency, fast probe pulse…
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs
Zhaojing Zhou, Xunchao Li, Minghao Li +8
The rapid scaling of Large Language Models (LLMs) elevates inference costs and compounds substantial deployment barriers. While quantization to 8 or 4 bits mitigates this, sub-3-bi…
PaddleOCR 3.0 Technical Report
Cheng Cui, Ting Sun, Manhui Lin +16
This technical report introduces PaddleOCR 3.0, an Apache-licensed open-source toolkit for OCR and document parsing. To address the growing demand for document understanding in the…
PP-LiteSeg: A Superior Real-Time Semantic Segmentation Model
Juncai Peng, Yi Liu, Shiyu Tang +13
Real-world applications have high demands for semantic segmentation methods. Although semantic segmentation has made remarkable leap-forwards with deep learning, the performance of…
PAFNet: An Efficient Anchor-Free Object Detector Guidance
Ying Xin, Guanzhong Wang, Mingyuan Mao +5
Object detection is a basic but challenging task in computer vision, which plays a key role in a variety of industrial applications. However, object detectors based on deep learnin…
ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
Shuohuan Wang, Yu Sun, Yang Xiang +26
Pre-trained language models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks. GPT-3 has shown that scaling up pre-trained language models c…