37 citations · 73 across the 6 of their papers we have counts for
6 papers
Efficient LLM Inference with Kcache
Qiaozhi He, Zhihua Wu
Large Language Models(LLMs) have had a profound impact on AI applications, particularly in the domains of long-text comprehension and generation. KV Cache technology is one of the…
Boosting Distributed Training Performance of the Unpadded BERT Model
Jinle Zeng, Min Li, Zhihua Wu +4
Pre-training models are an important tool in Natural Language Processing (NLP), while the BERT model is a classic pre-training model whose structure has been widely adopted by foll…
Nebula-I: A General Framework for Collaboratively Training Deep Learning Models on Low-Bandwidth Cloud Clusters
Yang Xiang, Zhihua Wu, Weibao Gong +15
The ever-growing model size and scale of compute have attracted increasing interests in training deep learning models over multiple nodes. However, when it comes to training on clo…
ERNIE-ViLG: Unified Generative Pre-training for Bidirectional Vision-Language Generation
Han Zhang, Weichong Yin, Yewei Fang +7
Conventional methods for the image-text generation tasks mainly tackle the naturally bidirectional generation tasks separately, focusing on designing task-specific frameworks to im…
ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
Shuohuan Wang, Yu Sun, Yang Xiang +26
Pre-trained language models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks. GPT-3 has shown that scaling up pre-trained language models c…
End-to-end Adaptive Distributed Training on PaddlePaddle
Yulong Ao, Zhihua Wu, Dianhai Yu +10
Distributed training has become a pervasive and effective approach for training a large neural network (NN) model with processing massive data. However, it is very challenging to s…