1 paper
Miaobo Hu, Shuhao Hu, Bokun Wang +5
Reinforcement learning improves LLM reasoning, but PPO/GRPO typically use fixed clipping and decoding temperature, which makes training brittle and tuning-heavy. We propose Adaptiv…