4 papers
How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation
Aida Usmanova, Zangir Iklassov, Markus Leippold +1
Automated fact-checking (AFC) systems retrieve evidence and predict claim veracity, yet evaluations omit simple baselines, systems are developed for a single benchmark and cannot b…
SymStep: Symbolic Step Verification for Logical Reasoning
Aida Usmanova, Rui Gao, Dilshod Azizov +2
Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps. We introduce SymStep: an L…
Reporting and Analysing the Environmental Impact of Language Models on the Example of Commonsense Question Answering with External Knowledge
Aida Usmanova, Junbo Huang, Debayan Banerjee +1
Human-produced emissions are growing at an alarming rate, causing already observable changes in the climate and environment in general. Each year global carbon dioxide emissions hi…
Scholarly Question Answering using Large Language Models in the NFDI4DataScience Gateway
Hamed Babaei Giglou, Tilahun Abedissa Taffa, Rana Abdullah +4
This paper introduces a scholarly Question Answering (QA) system on top of the NFDI4DataScience Gateway, employing a Retrieval Augmented Generation-based (RAG) approach. The NFDI4D…