1 citations · 2 across the 2 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
Qiang Liu, Wuganjing Song, Zhenzhou Lin +4
The reasoning capabilities of Large Language Models (LLMs) are typically developed through the single-turn reinforcement learning, whereas real-world applications often involve mul…
cs.CL2025
Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning
Chen Li, Nazhou Liu, Kai Yang
Since DeepSeek-R1 popularized, Group Relative Policy Optimization (GRPO) has become the core part of training Reasoning LLMs. However, we find some deficiency that influences RL st…