20 citations · 20 across the 4 of their papers we have counts for
4 papers
DARO: Difficulty-Aware Reweighting Policy Optimization
Jingyu Zhou, Lu Ma, Hao Liang +3
Recent advances in large language models (LLMs) have shown that reasoning ability can be significantly enhanced through Reinforcement Learning with Verifiable Rewards (RLVR). Group…
LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
Zhen Hao Wong, Jingwen Deng, Runming He +7
Large language models (LLMs) excel at many supervised tasks but often struggle with structured reasoning in unfamiliar settings. This discrepancy suggests that standard fine-tuning…
Measurements of branching fractions of , and
BESIII Collaboration, M. Ablikim, M. N. Achasov +687
Utilizing of collision data taken with the BESIII detector at the center-of-mass energy of 3.773 GeV, we report the measurements of absolute branching f…
Strong and weak tests in sequential decays of polarized hyperons
BESIII Collaboration, M. Ablikim, M. N. Achasov +650
The processes and subsequent decays are studied using the world's largest and data samples collected with the BESIII detector. The…