activity
20242026
collaborators

5 papers

cs.CL2026

Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation

Adam Dejl, James Barry, Alessandra Pascale +1

Despite demonstrating remarkable performance across a wide range of tasks, large language models (LLMs) have also been found to frequently produce outputs that are incomplete or se…

cs.CL2026

FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models

Javier Carnerero-Cano, Massimiliano Pronesti, Radu Marinescu +6

Large language models (LLMs) are widely used in knowledge-intensive applications but often generate factually incorrect responses. A promising approach to rectify these flaws is co…

cs.CL2025

FactReasoner: A Probabilistic Approach to Long-Form Factuality Assessment for Large Language Models

Radu Marinescu, Debarun Bhattacharjya, Junkyu Lee +5

Large language models (LLMs) have achieved remarkable success in generative tasks, yet they often fall short in ensuring the factual accuracy of their outputs, thus limiting their…

cs.CL2025

Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies

Massimiliano Pronesti, Joao Bettencourt-Silva, Paul Flanagan +4

Extracting scientific evidence from biomedical studies for clinical research questions (e.g., Does stem cell transplantation improve quality of life in patients with medically refr…

cs.CL2024

WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia

Yufang Hou, Alessandra Pascale, Javier Carnerero-Cano +5

Retrieval-augmented generation (RAG) has emerged as a promising solution to mitigate the limitations of large language models (LLMs), such as hallucinations and outdated informatio…