48 citations · 125 across the 8 of their papers we have counts for
7 papers
Label Information Enhanced Fraud Detection against Low Homophily in Graphs
Yuchen Wang, Jinghui Zhang, Zhengjie Huang +9
Node classification is a substantial problem in graph-based fraud detection. Many existing works adopt Graph Neural Networks (GNNs) to enhance fraud detectors. While promising, cur…
TA-MoE: Topology-Aware Large Scale Mixture-of-Expert Training
Chang Chen, Min Li, Zhihua Wu +2
Sparsely gated Mixture-of-Expert (MoE) has demonstrated its effectiveness in scaling up deep neural networks to an extreme scale. Despite that numerous efforts have been made to im…
Boosting Distributed Training Performance of the Unpadded BERT Model
Jinle Zeng, Min Li, Zhihua Wu +4
Pre-training models are an important tool in Natural Language Processing (NLP), while the BERT model is a classic pre-training model whose structure has been widely adopted by foll…
Large-scale Knowledge Distillation with Elastic Heterogeneous Computing Resources
Ji Liu, Daxiang Dong, Xi Wang +5
Although more layers and more parameters generally improve the accuracy of the models, such big models generally have high computational complexity and require big memory, which ex…
HelixFold: An Efficient Implementation of AlphaFold2 using PaddlePaddle
Guoxia Wang, Xiaomin Fang, Zhihua Wu +6
Accurate protein structure prediction can significantly accelerate the development of life science. The accuracy of AlphaFold2, a frontier end-to-end structure prediction system, i…
ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
Shuohuan Wang, Yu Sun, Yang Xiang +26
Pre-trained language models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks. GPT-3 has shown that scaling up pre-trained language models c…