3 papers
cs.CL2025
ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations
Yindong Wang, Martin Preiß, Margarita Bugueño +4
The mechanisms underlying scientific confabulation in Large Language Models (LLMs) remain poorly understood. We introduce ReFACT (Reddit False And Correct Texts), a benchmark of 1,…
cs.CL2025
Rethinking Graph-Based Document Classification: Learning Data-Driven Structures Beyond Heuristic Approaches
Margarita Bugueño, Gerard de Melo
In document classification, graph-based models effectively capture document structure, overcoming sequence length limitations and enhancing contextual understanding. However, most…
cs.CL2024
GraphLSS: Integrating Lexical, Structural, and Semantic Features for Long Document Extractive Summarization
Margarita Bugueño, Hazem Abou Hamdan, Gerard de Melo
Heterogeneous graph neural networks have recently gained attention for long document summarization, modeling the extraction as a node classification task. Although effective, these…