2 papers
cs.HC2026
Human- vs. AI-generated tests: dimensionality and information accuracy in latent trait evaluation
Mario Angelelli, Morena Oliva, Serena Arima +1
Artificial Intelligence (AI) and large language models (LLMs) are increasingly used in social and psychological research. Among potential applications, LLMs can be used to generate…
cs.CL2025
The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
Maria Paz Oliva, Adriana Correia, Ivan Vankov +1
Evaluating Natural Language Generation (NLG) is crucial for the practical adoption of AI, but has been a longstanding research challenge. While human evaluation is considered the d…