8 citations · 11 across the 3 of their papers we have counts for
7 papers · 1 filter
Retrieving Supporting Evidence for LLMs Generated Answers
Siqing Huo, Negar Arabzadeh, Charles L. A. Clarke
Current large language models (LLMs) can exhibit near-human levels of performance on many natural language tasks, including open-domain question answering. Unfortunately, they also…
Perspectives on Large Language Models for Relevance Judgment
Guglielmo Faggioli, Laura Dietz, Charles Clarke +8
When asked, large language models (LLMs) like ChatGPT claim that they can assist with relevance judgments but it is not clear whether automated judgments can reliably be used in ev…
Human Preferences as Dueling Bandits
Xinyi Yan, Chengxi Luo, Charles L. A. Clarke +3
The dramatic improvements in core information retrieval tasks engendered by neural rankers create a need for novel evaluation methods. If every ranker returns highly relevant items…
Predicting Efficiency/Effectiveness Trade-offs for Dense vs. Sparse Retrieval Strategy Selection
Negar Arabzadeh, Xinyi Yan, Charles L. A. Clarke
Over the last few years, contextualized pre-trained transformer models such as BERT have provided substantial improvements on information retrieval tasks. Recent approaches based o…
Assessing top- preferences
Charles L. A. Clarke, Alexandra Vtyurina, Mark D. Smucker
Assessors make preference judgments faster and more consistently than graded judgments. Preference judgments can also recognize distinctions between items that appear equivalent un…
The Effects of Latency Penalties in Evaluating Push Notification Systems
Luchen Tan, Jimmy Lin, Adam Roegiest +1
We examine the effects of different latency penalties in the evaluation of push notification systems, as operationalized in the TREC 2015 Microblog track evaluation. The purpose of…