Speech Intention Understanding in a Head-final Language: A Disambiguation Utilizing Intonation-dependency
arXiv:1811.04231 · doi:10.1145/3529648
Abstract
For a large portion of real-life utterances, the intention cannot be solely decided by either their semantic or syntactic characteristics. Although not all the sociolinguistic and pragmatic information can be digitized, at least phonetic features are indispensable in understanding the spoken language. Especially in head-final languages such as Korean, sentence-final prosody has great importance in identifying the speaker's intention. This paper suggests a system which identifies the inherent intention of a spoken utterance given its transcript, in some cases using auxiliary acoustic features. The main point here is a separate distinction for cases where discrimination of intention requires an acoustic cue. Thus, the proposed classification system decides whether the given utterance is a fragment, statement, question, command, or a rhetorical question/command, utilizing the intonation-dependency coming from the head-finality. Based on an intuitive understanding of the Korean language that is engaged in the data annotation, we construct a network which identifies the intention of a speech, and validate its utility with the test sentences. The system, if combined with up-to-date speech recognizers, is expected to be flexibly inserted into various language understanding modules.
14 pages, 2 figures, 7 tables; Identical to the previous revision. The latest version of this manuscript is recently accepted at ACM TALLIP, with the modified title, authors, and contents (see the DOI below). Please refer to THIS version only when relevant to the analysis with speech data, and refer to the journal version to cite the protocol and dataset
References in corpus (7)
- Adam: A Method for Stochastic Optimization
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Scaling Laws for Neural Language Models
- A Structured Self-attentive Sentence Embedding
- DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
- Transformer-based Korean Pretrained Language Models: A Survey on Three Years of Progress
- Giving Space to Your Message: Assistive Word Segmentation for the Electronic Typing of Digital Minorities