activity
20212024
most citedERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation

37 citations · 73 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2024

Efficient LLM Inference with Kcache

Qiaozhi He, Zhihua Wu

Large Language Models(LLMs) have had a profound impact on AI applications, particularly in the domains of long-text comprehension and generation. KV Cache technology is one of the…

cs.DC20224 cited

Boosting Distributed Training Performance of the Unpadded BERT Model

Jinle Zeng, Min Li, Zhihua Wu +4

Pre-training models are an important tool in Natural Language Processing (NLP), while the BERT model is a classic pre-training model whose structure has been widely adopted by foll…

cs.LG20222 cited

Nebula-I: A General Framework for Collaboratively Training Deep Learning Models on Low-Bandwidth Cloud Clusters

Yang Xiang, Zhihua Wu, Weibao Gong +15

The ever-growing model size and scale of compute have attracted increasing interests in training deep learning models over multiple nodes. However, when it comes to training on clo…

cs.CV202130 cited

ERNIE-ViLG: Unified Generative Pre-training for Bidirectional Vision-Language Generation

Han Zhang, Weichong Yin, Yewei Fang +7

Conventional methods for the image-text generation tasks mainly tackle the naturally bidirectional generation tasks separately, focusing on designing task-specific frameworks to im…

cs.CL202137 cited

ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation

Shuohuan Wang, Yu Sun, Yang Xiang +26

Pre-trained language models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks. GPT-3 has shown that scaling up pre-trained language models c…

cs.DC2021

End-to-end Adaptive Distributed Training on PaddlePaddle

Yulong Ao, Zhihua Wu, Dianhai Yu +10

Distributed training has become a pervasive and effective approach for training a large neural network (NN) model with processing massive data. However, it is very challenging to s…