2 papers
cs.AI2025
Do Repetitions Matter? Strengthening Reliability in LLM Evaluations
Miguel Angel Alvarado Gonzalez, Michelle Bruno Hernandez, Miguel Angel Peñaloza Perez +3
LLM leaderboards often rely on single stochastic runs, but how many repetitions are required for reliable conclusions remains unclear. We re-evaluate eight state-of-the-art models…
cs.CL2025
AI4Math: A Native Spanish Benchmark for University-Level Mathematical Reasoning in Large Language Models
Miguel Angel Peñaloza Perez, Bruno Lopez Orozco, Jesus Tadeo Cruz Soto +3
Existing mathematical reasoning benchmarks are predominantly English only or translation-based, which can introduce semantic drift and mask languagespecific reasoning errors. To ad…