activity
20152022
most citedTowards a Human-like Open-Domain Chatbot

263 citations · 604 across the 10 of their papers we have counts for

collaborators
Showing cs.CLShow all

18 papers · 1 filter

cs.CL2022★ 9 cited

MTet: Multi-domain Translation for English and Vietnamese

Chinh Ngo, Trieu H. Trinh, Long Phan +5

We introduce MTet, the largest publicly available parallel corpus for English-Vietnamese translation. MTet consists of 4.2M high-quality training sentence pairs and a multi-domain…

cs.CL2022

Enriching Biomedical Knowledge for Low-resource Language Through Large-Scale Translation

Long Phan, Tai Dang, Hieu Tran +4

Biomedical data and benchmarks are highly valuable yet very limited in low-resource languages other than English such as Vietnamese. In this paper, we make use of a state-of-the-ar…

cs.CL2021

Beyond Distillation: Task-level Mixture-of-Experts for Efficient Inference

Sneha Kudugunta, Yanping Huang, Ankur Bapna +4

Sparse Mixture-of-Experts (MoE) has been a successful approach for scaling multilingual translation models to billions of parameters without a proportional increase in training com…

cs.CL2021

STraTA: Self-Training with Task Augmentation for Better Few-shot Learning

Tu Vu, Minh-Thang Luong, Quoc V. Le +2

Despite their recent successes in tackling many NLP tasks, large-scale pre-trained language models do not perform as well in few-shot settings where only a handful of training exam…

cs.CL2020

Pre-Training Transformers as Energy-Based Cloze Models

Kevin Clark, Minh-Thang Luong, Quoc V. Le +1

We introduce Electric, an energy-based cloze model for representation learning over text. Like BERT, it is a conditional generative model of tokens given their contexts. However, E…

cs.CL2020★ 263 cited

Towards a Human-like Open-Domain Chatbot

Daniel Adiwardana, Minh-Thang Luong, David R. So +8

We present Meena, a multi-turn open-domain chatbot trained end-to-end on data mined and filtered from public domain social media conversations. This 2.6B parameter neural network i…