3 papers
cs.CL2026
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
Adam Dejl, James Barry, Alessandra Pascale +1
Despite demonstrating remarkable performance across a wide range of tasks, large language models (LLMs) have also been found to frequently produce outputs that are incomplete or se…
cs.CL2026
SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures
Nedjma Ousidhoum, Junho Myung, Carla Perez-Almendros +27
We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an extended version of our manual…
cs.CL2026
FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models
Javier Carnerero-Cano, Massimiliano Pronesti, Radu Marinescu +6
Large language models (LLMs) are widely used in knowledge-intensive applications but often generate factually incorrect responses. A promising approach to rectify these flaws is co…