Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
When Models Decide and When They Bind: A Two-Stage Computation for Multiple-Choice Question-Answering
Hugh Mee Wong, Rick Nouwen, Albert Gatt
Multiple-choice question answering (MCQA) is easy to evaluate but adds a meta-task: models must both solve the problem and output the symbol that *represents* the answer, conflatin…
cs.CL2025
Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
Dongxu Lu, Johan Jeuring, Albert Gatt
Evaluating large language models (LLMs) in long-form, knowledge-grounded role-play dialogues remains challenging. This study compares LLM-generated and human-authored responses in…
cs.CL2025
Do LLMs exhibit the same commonsense capabilities across languages?
Ivan Martínez-Murillo, Elena Lloret, Paloma Moreda +1
This paper explores the multilingual commonsense generation abilities of Large Language Models (LLMs). To facilitate this investigation, we introduce MULTICOM, a novel benchmark th…