Showing cs.IRShow all
3 papers · 1 filter
cs.IR2026
When Is Complex Chunking Worth It? A Multi-Objective Evaluation of Chunking Methods at Scale
Laura Caspari, Kanishka Ghosh Dastidar, Michael Dinzinger +2
Dense retrieval is commonly evaluated on benchmarks that represent each document with a single embedding, even though real-world retrieval systems often index long documents that r…
cs.IR2026
WebFAQ 2.0: A Multilingual QA Dataset with Mined Hard Negatives for Dense Retrieval
Michael Dinzinger, Laura Caspari, Ali Salman +3
We introduce WebFAQ 2.0, a new version of the WebFAQ dataset, containing 198 million FAQ-based natural question-answer pairs across 108 languages. Compared to the previous version,…
cs.IR2025
CoRECT: A Framework for Evaluating Embedding Compression Techniques at Scale
L. Caspari, M. Dinzinger, K. Ghosh Dastidar +3
Dense retrieval systems have proven to be effective across various benchmarks, but require substantial memory to store large search indices. Recent advances in embedding compressio…