PLATO-XL: Exploring the Large-scale Pre-training of Dialogue Generation
arXiv:2109.09519
Abstract
To explore the limit of dialogue generation pre-training, we present the models of PLATO-XL with up to 11 billion parameters, trained on both Chinese and English social media conversations. To train such large models, we adopt the architecture of unified transformer with high computation and parameter efficiency. In addition, we carry out multi-party aware pre-training to better distinguish the characteristic information in social media conversations. With such designs, PLATO-XL successfully achieves superior performances as compared to other approaches in both Chinese and English chitchat. We further explore the capacity of PLATO-XL on other conversational tasks, such as knowledge grounded dialogue and task-oriented conversation. The experimental results indicate that PLATO-XL obtains state-of-the-art results across multiple conversational tasks, verifying its potential as a foundation model of conversational AI.
Findings of AACL-IJCNLP 2022. First four authors contributed equally to this work
References in corpus (13)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Scaling Laws for Neural Language Models
- The Next Decade in AI: Four Steps Towards Robust Artificial Intelligence
- Towards a Human-like Open-Domain Chatbot
- ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- PanGu-: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation
- EVA: An Open-Domain Chinese Dialogue System with Large-Scale Generative Pre-Training
- MultiWOZ 2.2 : A Dialogue Dataset with Additional Annotation Corrections and State Tracking Baselines
- Addressee and Response Selection in Multi-Party Conversations with Speaker Interaction RNNs
- CPM: A Large-scale Generative Chinese Pre-trained Language Model
- ProphetNet-X: Large-Scale Pre-training Models for English, Chinese, Multi-lingual, Dialog, and Code Generation
- Know More about Each Other: Evolving Dialogue Strategy via Compound Assessment