3 papers
cs.CL2026
Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval
João Maria Janeiro, Mathurin Videau, Andrea Caciolai +3
Multiple-choice (MCQA) benchmarks are the standard for evaluating pretrained large language models, but their reliance on log-likelihood scoring makes them unreliable. Specifically…
cs.AI2025
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models
Mouadh Yagoubi, Yasser Dahou, Billel Mokeddem +12
Existing benchmarks have proven effective for assessing the performance of fully trained large language models. However, we find striking differences in the early training stages o…
cs.CL2025
SCOPE: A Self-supervised Framework for Improving Faithfulness in Conditional Text Generation
Song Duong, Florian Le Bronnec, Alexandre Allauzen +4
Large Language Models (LLMs), when used for conditional text generation, often produce hallucinations, i.e., information that is unfaithful or not grounded in the input context. Th…