Sequence-to-Sequence Data Augmentation for Dialogue Language Understanding
arXiv:1807.01554
Abstract
In this paper, we study the problem of data augmentation for language understanding in task-oriented dialogue system. In contrast to previous work which augments an utterance without considering its relation with other utterances, we propose a sequence-to-sequence generation based data augmentation framework that leverages one utterance's same semantic alternatives in the training data. A novel diversity rank is incorporated into the utterance representation to make the model produce diverse utterances and these diversely augmented utterances help to improve the language understanding module. Experimental results on the Airline Travel Information System dataset and a newly created semantic frame annotation on Stanford Multi-turn, Multidomain Dialogue Dataset show that our framework achieves significant improvements of 6.38 and 10.04 F-scores respectively when only a training set of hundreds utterances is represented. Case studies also confirm that our method generates diverse utterances.
Accepted By COLING2018
References in corpus (3)
Cited by in corpus (6)
- An Empirical Survey of Data Augmentation for Time Series Classification with Neural Networks
- End-to-End Spoken Language Understanding: Performance analyses of a voice command task in a low resource setting
- Paraphrase Augmented Task-Oriented Dialog Generation
- Cross-Modal Generative Augmentation for Visual Question Answering
- Learning to Bridge Metric Spaces: Few-shot Joint Learning of Intent Detection and Slot Filling
- Linguistic Knowledge in Data Augmentation for Natural Language Processing: An Example on Chinese Question Matching