13 papers
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
Xanh Ho, Jiahao Huang, Florian Boudin +1
Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Recent studies have proposed using l…
Refining and Reusing Annotation Guidelines for LLM Annotation
Kon Woo Kim, Jin-Dong Kim, Akiko Aizawa
While Large Language Models (LLMs) demonstrate remarkable performance on zero-shot annotation tasks, they often struggle with the specialized conventions of gold-standard benchmark…
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
Xanh Ho, Yun-Ang Wu, Sunisth Kumar +4
We present SciClaimEval, a new scientific dataset for the claim verification task. Unlike existing resources, SciClaimEval features authentic claims, including refuted ones, direct…
FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems
Juan Junqueras, Florian Boudin, May-Myo Zin +5
Hate speech (HS) is a critical issue in online discourse, and one promising strategy to counter it is through the use of counter-narratives (CNs). Datasets linking HS with CNs are…
Revisiting Bi-Encoder Neural Search: An Encoding--Searching Separation Perspective
Hung-Nghiep Tran, Akiko Aizawa, Atsuhiro Takasu
This paper reviews, analyzes, and proposes a new perspective on the bi-encoder architecture for neural search. While the bi-encoder architecture is widely used due to its simplicit…
Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts
Xanh Ho, Yun-Ang Wu, Sunisth Kumar +3
With the growing number of submitted scientific papers, there is an increasing demand for systems that can assist reviewers in evaluating research claims. Experimental results are…