activity
20192022
most citedEVA: An Open-Domain Chinese Dialogue System with Large-Scale Generative Pre-Training

29 citations · 86 across the 9 of their papers we have counts for

collaborators

11 papers

cs.CL2022

Learning Instructions with Unlabeled Data for Zero-Shot Cross-Task Generalization

Yuxian Gu, Pei Ke, Xiaoyan Zhu +1

Training language models to learn from human instructions for zero-shot cross-task generalization has attracted much attention in NLP communities. Recently, instruction tuning (IT)…

cs.CL2022

Rethinking and Refining the Distinct Metric

Siyang Liu, Sahand Sabour, Yinhe Zheng +3

Distinct- score\cite{Li2016} is a widely used automatic metric for evaluating diversity in language generation tasks. However, we observed that the original approach for calcula…

cs.CL202129 cited

EVA: An Open-Domain Chinese Dialogue System with Large-Scale Generative Pre-Training

Hao Zhou, Pei Ke, Zheng Zhang +11

Although pre-trained language models have remarkably enhanced the generation ability of dialogue systems, open-domain Chinese dialogue systems are still limited by the dialogue dat…

cs.CL202115 cited

CPM-2: Large-scale Cost-effective Pre-trained Language Models

Zhengyan Zhang, Yuxian Gu, Xu Han +16

In recent years, the size of pre-trained language models (PLMs) has grown by leaps and bounds. However, efficiency issues of these large-scale PLMs limit their utilization in real-…

cs.CL2021

JointGT: Graph-Text Joint Representation Learning for Text Generation from Knowledge Graphs

Pei Ke, Haozhe Ji, Yu Ran +5

Existing pre-trained models for knowledge-graph-to-text (KG-to-text) generation simply fine-tune text-to-text pre-trained models such as BART or T5 on KG-to-text datasets, which la…

cs.CL202022 cited

CPM: A Large-scale Generative Chinese Pre-trained Language Model

Zhengyan Zhang, Xu Han, Hao Zhou +22

Pre-trained Language Models (PLMs) have proven to be beneficial for various downstream NLP tasks. Recently, GPT-3, with 175 billion parameters and 570GB training data, drew a lot o…