2 papers
cs.CL2025
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
Alexander R. Fabbri, Diego Mares, Jorge Flores +5
Although recent Large Language Models (LLMs) have shown rapid improvement on reasoning benchmarks in English, the evaluation of such LLMs' multilingual reasoning capability across…
cs.CL2025
Can Artificial Intelligence Write Like Borges? An Evaluation Protocol for Spanish Microfiction
Gerardo Aleman Manzanarez, Nora de la Cruz Arana, Jorge Garcia Flores +3
Automated story writing has been a subject of study for over 60 years. Large language models can generate narratively consistent and linguistically coherent short fiction texts. De…