4 papers
WebFAQ 2.0: A Multilingual QA Dataset with Mined Hard Negatives for Dense Retrieval
Michael Dinzinger, Laura Caspari, Ali Salman +3
We introduce WebFAQ 2.0, a new version of the WebFAQ dataset, containing 198 million FAQ-based natural question-answer pairs across 108 languages. Compared to the previous version,…
CoRECT: A Framework for Evaluating Embedding Compression Techniques at Scale
L. Caspari, M. Dinzinger, K. Ghosh Dastidar +3
Dense retrieval systems have proven to be effective across various benchmarks, but require substantial memory to store large search indices. Recent advances in embedding compressio…
Revisiting the Relation Between Robustness and Universality
M. Klabunde, L. Caspari, F. Lemmerich
The modified universality hypothesis proposed by Jones et al. (2022) suggests that adversarially robust models trained for a given task are highly similar. We revisit the hypothesi…
WebFAQ: A Multilingual Collection of Natural Q&A Datasets for Dense Retrieval
Michael Dinzinger, Laura Caspari, Kanishka Ghosh Dastidar +2
We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from FAQ-style schema.org annotations. In total, the data collection consists of 96 m…