2 papers
cs.CL2026
mAceReason-Math: A Dataset of High-Quality Multilingual Math Problems Ready For RLVR
Konstantin Dobler, Simon Lehnerer, Federico Scozzafava +2
Reinforcement Learning with Verifiable Rewards (RLVR) has been successfully applied to significantly boost the capabilities of pretrained large language models, especially in the m…
cs.CL2026
Multilingual Reasoning Gym: Multilingual Scaling of Procedural Reasoning Environments
Konstantin Dobler, Simon Lehnerer, Federico Scozzafava +2
We present the Multilingual Reasoning Gym, an extension of Reasoning Gym (Stojanovski et al., 2025), that procedurally generates verifiable reasoning problems across 14 languages.…