1 citations · 1 across the 3 of their papers we have counts for
5 papers
Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic Sampling
Siyuan Gan, Yuhan Li, Xiran Wang +6
Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) is a prominent variant of Group Relative Policy Optimization (GRPO). DAPO introduces several improvements over GRPO.…
When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation
Siyuan Gan, Yuhan Li, Xiran Wang +5
On-policy distillation (OPD) has recently emerged as a popular post-training paradigm for large language models (LLMs), providing an efficient way to transfer the knowledge and cap…
Efficient Reinforcement Learning with Semantic and Token Entropy for LLM Reasoning
Hongye Cao, Zhixin Bai, Ziyue Peng +5
Reinforcement learning with verifiable rewards (RLVR) has demonstrated superior performance in enhancing the reasoning capability of large language models (LLMs). However, this acc…
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning
Siyuan Gan, Jiaheng Liu, Boyan Wang +8
Large reasoning models (LRMs) have attracted much attention due to their exceptional performance. However, their performance mainly stems from thinking, a long Chain of Thought (Co…
SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
Hongye Cao, Sijia Jing, Yanming Wang +14
With the rapid advancement of Large Language Models (LLMs), the safety of LLMs has been a critical concern requiring precise assessment. Current benchmarks primarily concentrate on…