Foundations of Vector Retrieval
arXiv:2401.09350 · doi:10.1007/978-3-031-55182-6
Abstract
Vectors are universal mathematical objects that can represent text, images, speech, or a mix of these data modalities. That happens regardless of whether data is represented by hand-crafted features or learnt embeddings. Collect a large enough quantity of such vectors and the question of retrieval becomes urgently relevant: Finding vectors that are more similar to a query vector. This monograph is concerned with the question above and covers fundamental concepts along with advanced data structures and algorithms for vector retrieval. In doing so, it recaps this fascinating topic and lowers barriers of entry into this rich area of research.
References in corpus (12)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Asymmetric LSH (ALSH) for Sublinear Time Maximum Inner Product Search (MIPS)
- An Efficiency Study for SPLADE Models
- EFANNA : An Extremely Fast Approximate Nearest Neighbor Search Algorithm Based on kNN Graph
- On Symmetric and Asymmetric LSHs for Inner Product Search
- Improved Asymmetric Locality Sensitive Hashing (ALSH) for Maximum Inner Product Search (MIPS)
- Efficient and Effective Tree-based and Neural Learning to Rank
- From Distillation to Hard Negative Sampling: Making Sparse Neural IR Models More Effective
- FreshDiskANN: A Fast and Accurate Graph-Based ANN Index for Streaming Similarity Search
- Worst-case Performance of Popular Approximate Nearest Neighbor Search Implementations: Guarantees and Limitations
- Automating Nearest Neighbor Search Configuration with Constrained Optimization
- A Bandit Approach to Maximum Inner Product Search