3 papers
cs.CL2025
Socratic-Zero : Bootstrapping Reasoning via Data-Free Agent Co-evolution
Shaobo Wang, Zhengbo Jiao, Zifan Zhang +6
Recent breakthroughs in large language models (LLMs) on reasoning tasks rely heavily on massive, high-quality datasets-typically human-annotated and thus difficult to scale. While…
cs.CL2025
SKYLENAGE Technical Report: Mathematical Reasoning and Contest-Innovation Benchmarks for Multi-Level Math Evaluation
Hu Wei, Ze Xu, Boyu Yang +15
Large language models (LLMs) now perform strongly on many public math suites, yet frontier separation within mathematics increasingly suffers from ceiling effects. We present two c…
cs.CL2025
Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs
Huichi Zhou, Zehao Xu, Munan Zhao +3
In this paper, we introduce the Multilingual Moral Reasoning Benchmark (MMRB) to evaluate the moral reasoning abilities of large language models (LLMs) across five typologically di…