2 citations · 3 across the 8 of their papers we have counts for
6 papers · 1 filter
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
Haonan He, Haodi Lei, Yun Luo +13
On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-co…
LatentMem: Customizing Latent Memory for Multi-Agent Systems
Muxin Fu, Xiangyuan Xue, Yafu Li +5
Large language model (LLM)-powered multi-agent systems (MAS) demonstrate remarkable collective intelligence, wherein multi-agent memory serves as a pivotal mechanism for continual…
FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
Runquan Gui, Yafu Li, Xiaoye Qu +3
Reinforcement Learning with Verifiable Rewards (RLVR) has markedly improved the performance of Large Language Models (LLMs) on tasks requiring multi-step reasoning. However, most R…
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
Zhilin Wang, Yafu Li, Shunkai Zhang +4
Whether Reinforcement Learning with Verifiable Rewards (RLVR) endows Large Language Models (LLMs) with new capabilities or merely elicits latent traces remains a central debate. In…
A Survey of Reinforcement Learning for Large Reasoning Models
Kaiyan Zhang, Yuxin Zuo, Bingxiang He +36
In this paper, we survey recent advances in Reinforcement Learning (RL) for reasoning with Large Language Models (LLMs). RL has achieved remarkable success in advancing the frontie…
Towards an AI Musician: Synthesizing Sheet Music Problems for Musical Reasoning
Zhilin Wang, Zhe Yang, Yun Luo +8
Enhancing the ability of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) to interpret sheet music is a crucial step toward building AI musicians. However,…