Deep Sentence Embedding Using Long Short-Term Memory Networks: Analysis and Application to Information Retrieval
arXiv:1502.06922 · doi:10.1109/TASLP.2016.2520371
Abstract
This paper develops a model that addresses sentence embedding, a hot topic in current natural language processing research, using recurrent neural networks with Long Short-Term Memory (LSTM) cells. Due to its ability to capture long term memory, the LSTM-RNN accumulates increasingly richer information as it goes through the sentence, and when it reaches the last word, the hidden layer of the network provides a semantic representation of the whole sentence. In this paper, the LSTM-RNN is trained in a weakly supervised manner on user click-through data logged by a commercial web search engine. Visualization and analysis are performed to understand how the embedding process works. The model is found to automatically attenuate the unimportant words and detects the salient keywords in the sentence. Furthermore, these detected keywords are found to automatically activate different cells of the LSTM-RNN, where words belonging to a similar topic activate the same cell. As a semantic representation of the sentence, the embedding vector can be used in many different applications. These automatic keyword detection and topic allocation abilities enabled by the LSTM-RNN allow the network to perform document retrieval, a difficult language processing task, where the similarity between the query and documents can be measured by the distance between their corresponding sentence embedding vectors computed by the LSTM-RNN. On a web search task, the LSTM-RNN embedding is shown to significantly outperform several existing state of the art methods. We emphasize that the proposed model generates sentence embedding vectors that are specially useful for web document retrieval tasks. A comparison with a well known general sentence embedding method, the Paragraph Vector, is performed. The results show that the proposed method in this paper significantly outperforms it for web document retrieval task.
To appear in IEEE/ACM Transactions on Audio, Speech, and Language Processing
Cited by in corpus (27)
- A Survey of Multi-View Representation Learning
- Multimodal Intelligence: Representation Learning, Information Fusion, and Applications
- Information Retrieval: Recent Advances and Beyond
- Distributed Compressive Sensing: A Deep Learning Approach
- Learning to Rank Question Answer Pairs with Holographic Dual LSTM Architecture
- DeepIST: Deep Image-based Spatio-Temporal Network for Travel Time Estimation
- Attentive Deep Neural Networks for Legal Document Retrieval
- Neural ranking models for document retrieval
- Leveraging Structural Context Models and Ranking Score Fusion for Human Interaction Prediction
- Learning a Product Relevance Model from Click-Through Data in E-Commerce
- Match-Ignition: Plugging PageRank into Transformer for Long-form Text Matching
- A Multi-task Learning Framework for Product Ranking with BERT
- Query Rewriting via Cycle-Consistent Translation for E-Commerce Search
- Learning to Ask: Conversational Product Search via Representation Learning
- The Power of Selecting Key Blocks with Local Pre-ranking for Long Document Information Retrieval
- The Winning Solution to the IEEE CIG 2017 Game Data Mining Competition
- Learning Contextualized Document Representations for Healthcare Answer Retrieval
- MS MARCO Web Search: a Large-scale Information-rich Web Dataset with Millions of Real Click Labels
- Multi-Task Learning for Email Search Ranking with Auxiliary Query Clustering
- Deep Bag-of-Words Model: An Efficient and Interpretable Relevance Architecture for Chinese E-Commerce
- A Novel Scholar Embedding Model for Interdisciplinary Collaboration
- A Survey on Awesome Korean NLP Datasets
- Q-fid: Quantum Circuit Fidelity Improvement with LSTM Networks
- Explainable LLM-driven Multi-dimensional Distillation for E-Commerce Relevance Learning
- On the Complexity of Opinions and Online Discussions
- Water Quality Prediction on a Sigfox-compliant IoT Device: The Road Ahead of WaterS
- Neural Network Architecture for Credibility Assessment of Textual Claims