collaborators

6 papers

cs.IR2025

Neural Prioritisation for Web Crawling

Francesca Pezzuti, Sean MacAvaney, Nicola Tonellotto

Given the vast scale of the Web, crawling prioritisation techniques based on link graph traversal, popularity, link analysis, and textual content are frequently applied to surface…

cs.IR2025

Document Quality Scoring for Web Crawling

Francesca Pezzuti, Ariane Mueller, Sean MacAvaney +1

The internet contains large amounts of low-quality content, yet users expect web search engines to deliver high-quality, relevant results. The abundant presence of low-quality page…

cs.IR2025

MURR: Model Updating with Regularized Replay for Searching a Document Stream

Eugene Yang, Nicola Tonellotto, Dawn Lawrie +4

The Internet produces a continuous stream of new documents and user-generated queries. These naturally change over time based on events in the world and the evolution of language.…

cs.IR2025

Efficient Constant-Space Multi-Vector Retrieval

Sean MacAvaney, Antonio Mallia, Nicola Tonellotto

Multi-vector retrieval methods, exemplified by the ColBERT architecture, have shown substantial promise for retrieval by providing strong trade-offs in terms of retrieval latency a…

cs.IR2024

Faster Learned Sparse Retrieval with Block-Max Pruning

Antonio Mallia, Torten Suel, Nicola Tonellotto

Learned sparse retrieval systems aim to combine the effectiveness of contextualized language models with the scalability of conventional data structures such as inverted indexes. N…

cs.IR2024

Two-Step SPLADE: Simple, Efficient and Effective Approximation of SPLADE

Carlos Lassance, Hervé Dejean, Stéphane Clinchant +1

Learned sparse models such as SPLADE have successfully shown how to incorporate the benefits of state-of-the-art neural information retrieval models into the classical inverted ind…