2 papers
cs.CL2025
The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
Maria Paz Oliva, Adriana Correia, Ivan Vankov +1
Evaluating Natural Language Generation (NLG) is crucial for the practical adoption of AI, but has been a longstanding research challenge. While human evaluation is considered the d…
cs.CL2025
ConSens: Assessing context grounding in open-book question answering
Ivan Vankov, Matyo Ivanov, Adriana Correia +1
Large Language Models (LLMs) have demonstrated considerable success in open-book question answering (QA), where the task requires generating answers grounded in a provided external…