Document Embedding with Paragraph Vectors
arXiv:1507.07998
Abstract
Paragraph Vectors has been recently proposed as an unsupervised method for learning distributed representations for pieces of texts. In their work, the authors showed that the method can learn an embedding of movie review texts which can be leveraged for sentiment analysis. That proof of concept, while encouraging, was rather narrow. Here we consider tasks other than sentiment analysis, provide a more thorough comparison of Paragraph Vectors to other document modelling algorithms such as Latent Dirichlet Allocation, and evaluate performance of the method as we vary the dimensionality of the learned representation. We benchmarked the models on two document similarity data sets, one from Wikipedia, one from arXiv. We observe that the Paragraph Vector method performs significantly better than other methods, and propose a simple improvement to enhance embedding quality. Somewhat surprisingly, we also show that much like word embeddings, vector operations on Paragraph Vectors can perform useful semantic results.
8 pages
Cited by in corpus (52)
- TensorFlow: A system for large-scale machine learning
- Evaluation of Text Generation: A Survey
- Contextual LSTM (CLSTM) models for Large scale NLP tasks
- Efficient Vector Representation for Documents through Corruption
- HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework
- Target specific mining of COVID-19 scholarly articles using one-class approach
- A Comparison of Machine Learning Algorithms for the Surveillance of Autism Spectrum Disorder
- From Free Text to Clusters of Content in Health Records: An Unsupervised Graph Partitioning Approach
- On the Use of ArXiv as a Dataset
- Cognitive Database: A Step towards Endowing Relational Databases with Artificial Intelligence Capabilities
- Visual Text Correction
- Sherlock: A Deep Learning Approach to Semantic Data Type Detection
- Speeding up Word Mover's Distance and its variants via properties of distances between embeddings
- Enabling Cognitive Intelligence Queries in Relational Databases using Low-dimensional Word Embeddings
- FlagIt: A System for Minimally Supervised Human Trafficking Indicator Mining
- Context based Text-generation using LSTM networks
- Document Embedding for Scientific Articles: Efficacy of Word Embeddings vs TFIDF
- Semantic Regularities in Document Representations
- Context Aware Document Embedding
- PRIVEE: A Visual Analytic Workflow for Proactive Privacy Risk Inspection of Open Data
- Improving Active Learning in Systematic Reviews
- Scientific Statement Classification over arXiv.org
- Adversarial Privacy Preserving Graph Embedding against Inference Attack
- Combining Code Embedding with Static Analysis for Function-Call Completion
- Getting Started with Neural Models for Semantic Matching in Web Search
- A Scalable Chatbot Platform Leveraging Online Community Posts: A Proof-of-Concept Study
- Utilizing Embeddings for Ad-hoc Retrieval by Document-to-document Similarity
- Towards Music Captioning: Generating Music Playlist Descriptions
- Hierarchical Document Encoder for Parallel Corpus Mining
- Structure-Invariant Testing for Machine Translation
- Textual Data for Time Series Forecasting
- SimDoc: Topic Sequence Alignment based Document Similarity Framework
- SLAM-Inspired Simultaneous Contextualization and Interpreting for Incremental Conversation Sentences
- Explaining Away Syntactic Structure in Semantic Document Representations
- Triple Scoring Using Paragraph Vector - The Gailan Triple Scorer at WSDM Cup 2017
- KeyVec: Key-semantics Preserving Document Representations
- Encouraging Paragraph Embeddings to Remember Sentence Identity Improves Classification
- Generative Interest Estimation for Document Recommendations
- Iterative Relevance Feedback for Answer Passage Retrieval with Passage-level Semantic Match
- Fast Intent Classification for Spoken Language Understanding
- DrugDBEmbed : Semantic Queries on Relational Database using Supervised Column Encodings
- A Robotic Dating Coaching System Leveraging Online Communities Posts
- XtraLibD: Detecting Irrelevant Third-Party libraries in Java and Python Applications
- Multi-view user representation learning for user matching without personal information
- Using Paragraph Vectors to improve our existing code review assisting tool-CRUSO
- A Study into patient similarity through representation learning from medical records
- Clustering via Content-Augmented Stochastic Blockmodels
- Leveraging Distributional Semantics for Multi-Label Learning
- Geosocial Location Classification: Associating Type to Places Based on Geotagged Social-Media Posts
- NMT-based Cross-lingual Document Embeddings
- A New Framework for Machine Intelligence: Concepts and Prototype
- Recurrent Neural Network Language Model Adaptation Derived Document Vector