27 citations · 44 across the 10 of their papers we have counts for
7 papers · 1 filter
SWAN: A Generic Framework for Auditing Textual Conversational Systems
Tetsuya Sakai
We present a simple and generic framework for auditing a given textual conversational system, given some samples of its conversation sessions as its input. The framework computes a…
Relevance Assessments for Web Search Evaluation: Should We Randomise or Prioritise the Pooled Documents? (CORRECTED VERSION)
Tetsuya Sakai, Sijie Tao, Zhaohao Zeng
In the context of depth- pooling for constructing web search test collections, we compare two approaches to ordering pooled documents for relevance assessors: the prioritisation…
Corrected Evaluation Results of the NTCIR WWW-2, WWW-3, and WWW-4 English Subtasks
Tetsuya Sakai, Sijie Tao, Maria Maistro +7
Unfortunately, the official English (sub)task results reported in the NTCIR-14 WWW-2, NTCIR-15 WWW-3, and NTCIR-16 WWW-4 overview papers are incorrect due to noise in the official…
On Variants of Root Normalised Order-aware Divergence and a Divergence based on Kendall's Tau
Tetsuya Sakai
This paper reports on a follow-up study of the work reported in Sakai, which explored suitable evaluation measures for ordinal quantification tasks. More specifically, the present…
A Versatile Framework for Evaluating Ranked Lists in terms of Group Fairness and Relevance
Tetsuya Sakai, Jin Young Kim, Inho Kang
We present a simple and versatile framework for evaluating ranked lists in terms of group fairness and relevance, where the groups (i.e., possible attribute values) can be either n…
How to Measure the Reproducibility of System-oriented IR Experiments
Timo Breuer, Nicola Ferro, Norbert Fuhr +4
Replicability and reproducibility of experimental results are primary concerns in all the areas of science and IR is not an exception. Besides the problem of moving the field towar…