collaborators

13 papers

cs.CL2026

Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses

Xanh Ho, Jiahao Huang, Florian Boudin +1

Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Recent studies have proposed using l…

cs.CL2026

Refining and Reusing Annotation Guidelines for LLM Annotation

Kon Woo Kim, Jin-Dong Kim, Akiko Aizawa

While Large Language Models (LLMs) demonstrate remarkable performance on zero-shot annotation tasks, they often struggle with the specialized conventions of gold-standard benchmark…

cs.CL2026

SciClaimEval: Cross-modal Claim Verification in Scientific Papers

Xanh Ho, Yun-Ang Wu, Sunisth Kumar +4

We present SciClaimEval, a new scientific dataset for the claim verification task. Unlike existing resources, SciClaimEval features authentic claims, including refuted ones, direct…

cs.CL2026

FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems

Juan Junqueras, Florian Boudin, May-Myo Zin +5

Hate speech (HS) is a critical issue in online discourse, and one promising strategy to counter it is through the use of counter-narratives (CNs). Datasets linking HS with CNs are…

cs.LG2025

Revisiting Bi-Encoder Neural Search: An Encoding--Searching Separation Perspective

Hung-Nghiep Tran, Akiko Aizawa, Atsuhiro Takasu

This paper reviews, analyzes, and proposes a new perspective on the bi-encoder architecture for neural search. While the bi-encoder architecture is widely used due to its simplicit…

cs.CL2025

Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts

Xanh Ho, Yun-Ang Wu, Sunisth Kumar +3

With the growing number of submitted scientific papers, there is an increasing demand for systems that can assist reviewers in evaluating research claims. Experimental results are…