Synergistic Union of Word2Vec and Lexicon for Domain Specific Semantic Similarity
arXiv:1706.01967 · doi:10.1109/ICIINFS.2017.8300343
Abstract
Semantic similarity measures are an important part in Natural Language Processing tasks. However Semantic similarity measures built for general use do not perform well within specific domains. Therefore in this study we introduce a domain specific semantic similarity measure that was created by the synergistic union of word2vec, a word embedding method that is used for semantic similarity calculation and lexicon based (lexical) semantic similarity methods. We prove that this proposed methodology out performs word embedding methods trained on generic corpus and methods trained on domain specific corpus but do not use lexical semantic similarity methods to augment the results. Further, we prove that text lemmatization can improve the performance of word embedding methods.
6 Pages, 3 figures
References in corpus (2)
Cited by in corpus (5)
- SigmaLaw-ABSA: Dataset for Aspect-Based Sentiment Analysis in Legal Opinion Texts
- Identifying Relationships Among Sentences in Court Case Transcripts Using Discourse Relations
- Rule-Based Approach for Party-Based Sentiment Analysis in Legal Opinion Texts
- SHADE: Semantic Hypernym Annotator for Domain-specific Entities -- DnD Domain Use Case
- Fine Tuning Named Entity Extraction Models for the Fantasy Domain