63 citations · 115 across the 16 of their papers we have counts for
20 papers · 1 filter
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance
Zhuowen Han, Jinwei Xiao, Zhengxi Lu +9
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO)…
Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models
Dan Shi, Zhuowen Han, Simon Ostermann +3
Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models (LLMs) beyond the training domain, while supervised fine-tuning (S…
Thinking Seeds: Leveraging Historical Diversity for Position-Aware RL in LLMs
Lei Yang, Wei Bi, Chenxi Sun +2
On-policy reinforcement learning (RL) for language model post-training suffers from a fundamental tension: as training progresses, policy entropy collapses and sampling diversity d…
Revisiting Entropy in Reinforcement Learning for Large Reasoning Models
Renren Jin, Pengzhi Gao, Yuqi Ren +6
Reinforcement learning with verifiable rewards (RLVR) has emerged as a prominent paradigm for enhancing the reasoning capabilities of large language models (LLMs). However, the ent…
From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan
Lei Yang, Leiyu Pan, Bojian Xiong +14
Large language models (LLMs) have achieved remarkable success across a wide range of natural language processing tasks, yet their performance remains heavily biased toward high-res…
TaP: A Taxonomy-Guided Framework for Automated and Scalable Preference Data Generation
Renren Jin, Tianhao Shen, Xinwei Wu +9
Conducting supervised and preference fine-tuning of large language models (LLMs) requires high-quality datasets to improve their ability to follow instructions and align with human…