30 citations · 107 across the 9 of their papers we have counts for
11 papers · 1 filter
Overview of the TREC 2022 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz +4
This is the fourth year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human annotated training labels…
Overview of the TREC 2023 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz +5
This is the fifth year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human-annotated training labels a…
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur +5
The application of large language models to provide relevance assessments presents exciting opportunities to advance information retrieval, natural language processing, and beyond,…
LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?
Rikiya Takehi, Ellen M. Voorhees, Tetsuya Sakai +1
Test collections are information-retrieval tools that allow researchers to quickly and easily evaluate ranking algorithms. While test collections have become an integral part of IR…
Corrected Evaluation Results of the NTCIR WWW-2, WWW-3, and WWW-4 English Subtasks
Tetsuya Sakai, Sijie Tao, Maria Maistro +7
Unfortunately, the official English (sub)task results reported in the NTCIR-14 WWW-2, NTCIR-15 WWW-3, and NTCIR-16 WWW-4 overview papers are incorrect due to noise in the official…
Can Old TREC Collections Reliably Evaluate Modern Neural Retrieval Models?
Ellen M. Voorhees, Ian Soboroff, Jimmy Lin
Neural retrieval models are generally regarded as fundamentally different from the retrieval techniques used in the late 1990's when the TREC ad hoc test collections were construct…