Component-Enhanced Chinese Character Embeddings
arXiv:1508.06669
Abstract
Distributed word representations are very useful for capturing semantic information and have been successfully applied in a variety of NLP tasks, especially on English. In this work, we innovatively develop two component-enhanced Chinese character embedding models and their bigram extensions. Distinguished from English word embeddings, our models explore the compositions of Chinese characters, which often serve as semantic indictors inherently. The evaluations on both word similarity and text classification demonstrate the effectiveness of our models.
6 pages, 2 figures, conference, EMNLP 2015
References in corpus (1)
Cited by in corpus (5)
- Glyce: Glyph-vectors for Chinese Character Representations
- Learning Chinese Word Representations From Glyphs Of Characters
- Phonetic-enriched Text Representation for Chinese Sentiment Analysis with Reinforcement Learning
- Improving Word Vector with Prior Knowledge in Semantic Dictionary
- Enhancing Chinese Intent Classification by Dynamically Integrating Character Features into Word Embeddings with Ensemble Techniques