2 citations · 2 across the 2 of their papers we have counts for
6 papers
Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
Sunisth Kumar, Xanh Ho, Tim Schopf +3
Multimodal LLMs are increasingly used to assist scientific peer review, where a core requirement is verifying whether claims in a paper are supported by its evidence. Prior work ha…
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
Xanh Ho, Jiahao Huang, Florian Boudin +1
Extractive QA tasks are commonly evaluated using Exact Match (EM) and F1-score, but these metrics often fail to reflect true model performance. Recent studies have proposed using l…
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
Xanh Ho, Yun-Ang Wu, Sunisth Kumar +4
We present SciClaimEval, a new scientific dataset for the claim verification task. Unlike existing resources, SciClaimEval features authentic claims, including refuted ones, direct…
Format Matters: The Robustness of Multimodal LLMs in Reviewing Evidence from Tables and Charts
Xanh Ho, Yun-Ang Wu, Sunisth Kumar +3
With the growing number of submitted scientific papers, there is an increasing demand for systems that can assist reviewers in evaluating research claims. Experimental results are…
Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers
Xanh Ho, Sunisth Kumar, Yun-Ang Wu +3
Scientific claim verification against tables typically requires predicting whether a claim is supported or refuted given a table. However, we argue that predicting the final label…
MoreHopQA: More Than Multi-hop Reasoning
Julian Schnitzler, Xanh Ho, Jiahao Huang +3
Most existing multi-hop datasets are extractive answer datasets, where the answers to the questions can be extracted directly from the provided context. This often leads models to…