13 citations · 47 across the 15 of their papers we have counts for
7 papers · 1 filter
Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
Sunisth Kumar, Xanh Ho, Tim Schopf +3
Multimodal LLMs are increasingly used to assist scientific peer review, where a core requirement is verifying whether claims in a paper are supported by its evidence. Prior work ha…
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
Xanh Ho, Yun-Ang Wu, Sunisth Kumar +4
We present SciClaimEval, a new scientific dataset for the claim verification task. Unlike existing resources, SciClaimEval features authentic claims, including refuted ones, direct…
Overview of PAN 2026: Voight-Kampff Generative AI Detection, Text Watermarking, Multi-Author Writing Style Analysis, Generative Plagiarism Detection, and Reasoning Trajectory Detection
Janek Bevendorff, Maik Fröbe, André Greiner-Petter +9
The goal of the PAN workshop is to advance computational stylometry and text forensics via objective and reproducible evaluation. In 2026, we run the following five tasks: (1) Voig…
Overview of the Plagiarism Detection Task at PAN 2025
André Greiner-Petter, Maik Fröbe, Jan Philip Wahle +4
The generative plagiarism detection task at PAN 2025 aims at identifying automatically generated textual plagiarism in scientific articles and aligning them with their respective s…
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection
Tomas Horych, Christoph Mandl, Terry Ruas +4
High annotation costs from hiring or crowdsourcing complicate the creation of large, high-quality datasets needed for training reliable text classifiers. Recent research suggests u…
Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange
Ankit Satpute, Noah Giessing, Andre Greiner-Petter +4
Large Language Models (LLMs) have demonstrated exceptional capabilities in various natural language tasks, often achieving performances that surpass those of humans. Despite these…