2 papers
cs.CL2022
Improving Large-scale Paraphrase Acquisition and Generation
Yao Dou, Chao Jiang, Wei Xu
This paper addresses the quality issues in existing Twitter-based paraphrase datasets, and discusses the necessity of using two separate definitions of paraphrase for identificatio…
cs.CL2021
MultiTalk: A Highly-Branching Dialog Testbed for Diverse Conversations
Yao Dou, Maxwell Forbes, Ari Holtzman +1
We study conversational dialog in which there are many possible responses to a given history. We present the MultiTalk Dataset, a corpus of over 320,000 sentences of written conver…