Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization
Shiye Lei, Zhihao Cheng, Dacheng Tao
Reinforcement learning (RL) post-training has increasingly demonstrated strong ability to elicit reasoning behaviors in large language models (LLMs). For training efficiency, rollo…
cs.AI2025
Revisiting LLM Reasoning via Information Bottleneck
Shiye Lei, Zhihao Cheng, Kai Jia +1
Large language models (LLMs) have recently demonstrated remarkable progress in reasoning capabilities through reinforcement learning with verifiable rewards (RLVR). By leveraging s…