7 papers
MICE: Minimal Interaction Cross-Encoders for efficient Re-ranking
Mathias Vast, Victor Morand, Basile van Cooten +3
Cross-encoders deliver state-of-the-art ranking effectiveness in information retrieval, but have a high inference cost. This prevents them from being used as first-stage rankers, b…
From Tokens to Concepts: Leveraging SAE for SPLADE
Yuxuan Zong, Mathias Vast, Basile Van Cooten +2
Learned Sparse IR models, such as SPLADE, offer an excellent efficiency-effectiveness tradeoff. However, they rely on the underlying backbone vocabulary, which might hinder perform…
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
Pierre-Antoine Lequeu, Léo Labat, Laurène Cave +3
LLMs are ubiquitous in modern NLP, and while their applicability extends to texts produced for democratic activities such as online deliberations or large-scale citizen consultatio…
ToMMeR -- Efficient Entity Mention Detection from Large Language Models
Victor Morand, Nadi Tomeh, Josiane Mothe +1
Identifying which text spans refer to entities - mention detection - is both foundational for information extraction and a known performance bottleneck. We introduce ToMMeR, a ligh…
Reproducing and Comparing Distillation Techniques for Cross-Encoders
Victor Morand, Mathias Vast, Basile Van Cooten +3
Recent advances in Information Retrieval have established transformer-based cross-encoders as a keystone in IR. Recent studies have focused on knowledge distillation and showed tha…
On the Representations of Entities in Auto-regressive Large Language Models
Victor Morand, Josiane Mothe, Benjamin Piwowarski
Named entities are fundamental building blocks of knowledge in text, grounding factual information and structuring relationships within language. Despite their importance, it remai…