2 papers
cs.CL2026
Interpretable Coreference Resolution Evaluation Using Explicit Semantics
Bruno Gatti, Giuliano Martinelli, Roberto Navigli
Coreference resolution is typically evaluated using aggregate statistical metrics such as CoNLL-F1, which measure structural overlap between predicted and gold clusters. While wide…
cs.CL2026
Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards
Raffaele Pisano, Roberto Navigli
Process Reward Models (PRMs) have emerged as a powerful tool for providing step-level feedback when evaluating the reasoning of Large Language Models (LLMs), which frequently produ…