Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance
Siyao Song, Cong Ma, Zhihao Cheng +5
Large language models (LLMs) have recently advanced in reasoning when optimized with reinforcement learning (RL) under verifiable rewards. Existing methods primarily rely on outcom…
cs.AI2025
Revisiting LLM Reasoning via Information Bottleneck
Shiye Lei, Zhihao Cheng, Kai Jia +1
Large language models (LLMs) have recently demonstrated remarkable progress in reasoning capabilities through reinforcement learning with verifiable rewards (RLVR). By leveraging s…