Wikipedia-based Semantic Interpretation for Natural Language Processing
arXiv:1401.5697 · doi:10.1613/jair.2669
Abstract
Adequate representation of natural language semantics requires access to vast amounts of common sense and domain-specific world knowledge. Prior work in the field was based on purely statistical techniques that did not make use of background knowledge, on limited lexicographic knowledge bases such as WordNet, or on huge manual efforts such as the CYC project. Here we propose a novel method, called Explicit Semantic Analysis (ESA), for fine-grained semantic interpretation of unrestricted natural language texts. Our method represents meaning in a high-dimensional space of concepts derived from Wikipedia, the largest encyclopedia in existence. We explicitly represent the meaning of any text in terms of Wikipedia-based concepts. We evaluate the effectiveness of our method on text categorization and on computing the degree of semantic relatedness between fragments of natural language text. Using ESA results in significant improvements over the previous state of the art in both tasks. Importantly, due to the use of natural concepts, the ESA model is easy to explain to human users.
References in corpus (5)
- Thumbs up? Sentiment Classification using Machine Learning Techniques
- Semantic Similarity in a Taxonomy: An Information-Based Measure and its Application to Problems of Ambiguity in Natural Language
- Thumbs Up or Thumbs Down? Semantic Orientation Applied to Unsupervised Classification of Reviews
- Unsupervised Learning of Semantic Orientation from a Hundred-Billion-Word Corpus
- Measuring Semantic Similarity by Latent Relational Analysis
Cited by in corpus (20)
- Analogical Inference for Multi-Relational Embeddings
- The Tower of Babel Meets Web 2.0: User-Generated Content and its Applications in a Multilingual Context
- Text Relatedness Based on a Word Thesaurus
- Semantic Similarity from Natural Language and Ontology Analysis
- Semantic homophily in online communication: evidence from Twitter
- LinkNBed: Multi-Graph Representation Learning with Entity Linkage
- Fast and accurate annotation of short texts with Wikipedia pages
- Machine Learning with World Knowledge: The Position and Survey
- The CQC Algorithm: Cycling in Graphs to Semantically Enrich and Enhance a Bilingual Dictionary
- Cross-lingual Dataless Classification for Languages with Small Wikipedia Presence
- Recursive Feature Generation for Knowledge-based Learning
- Automated Word Puzzle Generation via Topic Dictionaries
- WikiContradiction: Detecting Self-Contradiction Articles on Wikipedia
- Entities of Interest
- Knowledge-Based Learning through Feature Generation
- Statistical Analysis of Multi-Relational Network Recovery
- Eliminating Search Intent Bias in Learning to Rank
- World Knowledge as Indirect Supervision for Document Clustering
- Modeling problems of identity in Little Red Riding Hood
- Semantic Sort: A Supervised Approach to Personalized Semantic Relatedness