27 citations · 35 across the 9 of their papers we have counts for
11 papers
Relevance Assessments for Web Search Evaluation: Should We Randomise or Prioritise the Pooled Documents? (CORRECTED VERSION)
Tetsuya Sakai, Sijie Tao, Zhaohao Zeng
In the context of depth- pooling for constructing web search test collections, we compare two approaches to ordering pooled documents for relevance assessors: the prioritisation…
Corrected Evaluation Results of the NTCIR WWW-2, WWW-3, and WWW-4 English Subtasks
Tetsuya Sakai, Sijie Tao, Maria Maistro +7
Unfortunately, the official English (sub)task results reported in the NTCIR-14 WWW-2, NTCIR-15 WWW-3, and NTCIR-16 WWW-4 overview papers are incorrect due to noise in the official…
On Variants of Root Normalised Order-aware Divergence and a Divergence based on Kendall's Tau
Tetsuya Sakai
This paper reports on a follow-up study of the work reported in Sakai, which explored suitable evaluation measures for ordinal quantification tasks. More specifically, the present…
A Versatile Framework for Evaluating Ranked Lists in terms of Group Fairness and Relevance
Tetsuya Sakai, Jin Young Kim, Inho Kang
We present a simple and versatile framework for evaluating ranked lists in terms of group fairness and relevance, where the groups (i.e., possible attribute values) can be either n…
AxIoU: An Axiomatically Justified Measure for Video Moment Retrieval
Riku Togashi, Mayu Otani, Yuta Nakashima +3
Evaluation measures have a crucial impact on the direction of research. Therefore, it is of utmost importance to develop appropriate and reliable evaluation measures for new applic…
DCH-2: A Parallel Customer-Helpdesk Dialogue Corpus with Distributions of Annotators' Labels
Zhaohao Zeng, Tetsuya Sakai
We introduce a data set called DCH-2, which contains 4,390 real customer-helpdesk dialogues in Chinese and their English translations. DCH-2 also contains dialogue-level annotation…