6 papers
Neural Prioritisation for Web Crawling
Francesca Pezzuti, Sean MacAvaney, Nicola Tonellotto
Given the vast scale of the Web, crawling prioritisation techniques based on link graph traversal, popularity, link analysis, and textual content are frequently applied to surface…
Document Quality Scoring for Web Crawling
Francesca Pezzuti, Ariane Mueller, Sean MacAvaney +1
The internet contains large amounts of low-quality content, yet users expect web search engines to deliver high-quality, relevant results. The abundant presence of low-quality page…
MURR: Model Updating with Regularized Replay for Searching a Document Stream
Eugene Yang, Nicola Tonellotto, Dawn Lawrie +4
The Internet produces a continuous stream of new documents and user-generated queries. These naturally change over time based on events in the world and the evolution of language.…
Efficient Constant-Space Multi-Vector Retrieval
Sean MacAvaney, Antonio Mallia, Nicola Tonellotto
Multi-vector retrieval methods, exemplified by the ColBERT architecture, have shown substantial promise for retrieval by providing strong trade-offs in terms of retrieval latency a…
Faster Learned Sparse Retrieval with Block-Max Pruning
Antonio Mallia, Torten Suel, Nicola Tonellotto
Learned sparse retrieval systems aim to combine the effectiveness of contextualized language models with the scalability of conventional data structures such as inverted indexes. N…
Two-Step SPLADE: Simple, Efficient and Effective Approximation of SPLADE
Carlos Lassance, Hervé Dejean, Stéphane Clinchant +1
Learned sparse models such as SPLADE have successfully shown how to incorporate the benefits of state-of-the-art neural information retrieval models into the classical inverted ind…