3 citations · 7 across the 5 of their papers we have counts for
5 papers
LLMJudge: LLMs for Relevance Judgments
Hossein A. Rahmani, Emine Yilmaz, Nick Craswell +6
The LLMJudge challenge is organized as part of the LLM4Eval workshop at SIGIR 2024. Test collections are essential for evaluating information retrieval (IR) systems. The evaluation…
Report on the 1st Workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) at SIGIR 2024
Hossein A. Rahmani, Clemencia Siro, Mohammad Aliannejadi +6
The first edition of the workshop on Large Language Model for Evaluation in Information Retrieval (LLM4Eval 2024) took place in July 2024, co-located with the ACM SIGIR Conference…
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
Clemencia Siro, Mohammad Aliannejadi, Maarten de Rijke
In ad-hoc retrieval, evaluation relies heavily on user actions, including implicit feedback. In a conversational setting such signals are usually unavailable due to the nature of t…
Context Does Matter: Implications for Crowdsourced Evaluation Labels in Task-Oriented Dialogue Systems
Clemencia Siro, Mohammad Aliannejadi, Maarten de Rijke
Crowdsourced labels play a crucial role in evaluating task-oriented dialogue systems (TDSs). Obtaining high-quality and consistent ground-truth labels from annotators presents chal…
Asking Multimodal Clarifying Questions in Mixed-Initiative Conversational Search
Yifei Yuan, Clemencia Siro, Mohammad Aliannejadi +2
In mixed-initiative conversational search systems, clarifying questions are used to help users who struggle to express their intentions in a single query. These questions aim to un…