14 papers
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
Francesco Maria Molfese, Luca Moroni, Ciro Porcaro +2
While Small Language Models (SLMs) have demonstrated promising performance on an increasingly wide array of commonsense reasoning benchmarks, current evaluation practices rely almo…
EMERGE: A Benchmark for Updating Knowledge Graphs with Emerging Textual Knowledge
Klim Zaporojets, Daniel Daza, Edoardo Barba +3
Knowledge Graphs (KGs) are structured knowledge repositories containing entities and relations between them. In this paper, we study the problem of automatically updating KGs over…
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
Tommaso Bonomo, Luca Gioffré, Roberto Navigli
Question Answering (QA) on narrative text poses a unique challenge to current systems, requiring a deep understanding of long, complex documents. However, the reliability of Narrat…
Do Large Language Models Understand Word Senses?
Domenico Meconi, Simone Stirpe, Federico Martelli +2
Understanding the meaning of words in context is a fundamental capability for Large Language Models (LLMs). Despite extensive evaluation efforts, the extent to which LLMs show evid…
Estimating Machine Translation Difficulty
Lorenzo Proietti, Stefano Perrella, Vilém Zouhar +2
Machine translation quality has steadily improved over the years, achieving near-perfect translations in recent benchmarks. These high-quality outputs make it difficult to distingu…
BOOKCOREF: Coreference Resolution at Book Scale
Giuliano Martinelli, Tommaso Bonomo, Pere-LluÃs Huguet Cabot +1
Coreference Resolution systems are typically evaluated on benchmarks containing small- to medium-scale documents. When it comes to evaluating long texts, however, existing benchmar…