15 papers
Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
Sunisth Kumar, Xanh Ho, Tim Schopf +3
Multimodal LLMs are increasingly used to assist scientific peer review, where a core requirement is verifying whether claims in a paper are supported by its evidence. Prior work ha…
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
Xanh Ho, Jiahao Huang, Florian Boudin +1
Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Recent studies have proposed using l…
EarlySciRev: A Dataset of Early-Stage Scientific Revisions Extracted from LaTeX Writing Traces
Léane Jourdan, Julien Aubert-Béduchaud, Yannis Chupin +2
Scientific writing is an iterative process that generates rich revision traces, yet publicly available resources typically expose only final or near-final versions of papers. This…
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
Xanh Ho, Yun-Ang Wu, Sunisth Kumar +4
We present SciClaimEval, a new scientific dataset for the claim verification task. Unlike existing resources, SciClaimEval features authentic claims, including refuted ones, direct…
Identifying Reliable Evaluation Metrics for Scientific Text Revision
Léane Jourdan, Florian Boudin, Richard Dufour +1
Evaluating text revision in scientific writing remains a challenge, as traditional metrics such as ROUGE and BERTScore primarily focus on similarity rather than capturing meaningfu…
FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems
Juan Junqueras, Florian Boudin, May-Myo Zin +5
Hate speech (HS) is a critical issue in online discourse, and one promising strategy to counter it is through the use of counter-narratives (CNs). Datasets linking HS with CNs are…