2 citations · 2 across the 3 of their papers we have counts for
2 papers
cs.AI2024★ 2 cited
Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
Zhiheng Xi, Wenxiang Chen, Boyang Hong +18
In this paper, we propose R: Learning Reasoning through Reverse Curriculum Reinforcement Learning (RL), a novel method that employs only outcome supervision to achieve the bene…
hep-ex2023
Studies of the decay
BESIII Collaboration, M. Ablikim, M. N. Achasov +608
The decay is studied based on 7.33 fb of collision data collected with the BESIII detector at center-of-mass energies in the range from 4.12…