8 citations · 8 across the 2 of their papers we have counts for
2 papers
cs.CL2025
RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines
Pengfei Yu, Dongming Shen, Silin Meng +8
We present RPGBench, the first benchmark designed to evaluate large language models (LLMs) as text-based role-playing game (RPG) engines. RPGBench comprises two core tasks: Game Cr…
cs.CL2025★ 8 cited
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Xinyu Guan, Li Lyna Zhang, Yifei Liu +5
We present rStar-Math to demonstrate that small language models (SLMs) can rival or even surpass the math reasoning capability of OpenAI o1, without distillation from superior mode…