A Dual Embedding Space Model for Document Ranking
arXiv:1602.01137
Abstract
A fundamental goal of search engines is to identify, given a query, documents that have relevant text. This is intrinsically difficult because the query and the document may use different vocabulary, or the document may contain query words without being relevant. We investigate neural word embeddings as a source of evidence in document ranking. We train a word2vec embedding model on a large unlabelled query corpus, but in contrast to how the model is commonly used, we retain both the input and the output projections, allowing us to leverage both the embedding spaces to derive richer distributional relationships. During ranking we map the query words into the input space and the document words into the output space, and compute a query-document relevance score by aggregating the cosine similarities across all the query-document word pairs. We postulate that the proposed Dual Embedding Space Model (DESM) captures evidence on whether a document is about a query term in addition to what is modelled by traditional term-frequency based approaches. Our experiments show that the DESM can re-rank top documents returned by a commercial Web search engine, like Bing, better than a term-matching based signal like TF-IDF. However, when ranking a larger set of candidate documents, we find the embeddings-based approach is prone to false positives, retrieving documents that are only loosely related to the query. We demonstrate that this problem can be solved effectively by ranking based on a linear mixture of the DESM and the word counting features.
This paper is an extended evaluation and analysis of the model proposed in a poster to appear in WWW'16, April 11 - 15, 2016, Montreal, Canada
References in corpus (5)
Cited by in corpus (27)
- Information Retrieval: Recent Advances and Beyond
- Ad Hoc Table Retrieval using Semantic Similarity
- Semantic Models for the First-stage Retrieval: A Comprehensive Review
- Pretrained Transformers for Text Ranking: BERT and Beyond
- Neural Models for Information Retrieval
- Neural Information Retrieval: A Literature Review
- Semantic Specialisation of Distributional Word Vector Spaces using Monolingual and Cross-Lingual Constraints
- Conformer-Kernel with Query Term Independence for Document Retrieval
- Efficient Neural Ranking using Forward Indexes
- Semantic Product Search
- Learning to Match Using Local and Distributed Representations of Text for Web Search
- Co-PACRR: A Context-Aware Neural IR Model for Ad-hoc Retrieval
- Improving Low-Resource Cross-lingual Document Retrieval by Reranking with Deep Bilingual Representations
- Toward a Deep Neural Approach for Knowledge-Based IR
- One word at a time: adversarial attacks on retrieval models
- Toward Incorporation of Relevant Documents in word2vec
- Deep Learning Relevance: Creating Relevant Information (as Opposed to Retrieving it)
- EXS: Explainable Search Using Local Model Agnostic Interpretability
- Utilizing Embeddings for Ad-hoc Retrieval by Document-to-document Similarity
- Improving Document Representations by Generating Pseudo Query Embeddings for Dense Retrieval
- Query Clustering using Segment Specific Context Embeddings
- An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking
- Conformer-Kernel with Query Term Independence at TREC 2020 Deep Learning Track
- Duet at TREC 2019 Deep Learning Track
- Intrinsic analysis for dual word embedding space models
- RelEmb: A relevance-based application embedding for Mobile App retrieval and categorization
- A Line in the Sand: Recommendation or Ad-hoc Retrieval?