Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
Zhiqin Yang, Jingwen Fu, Yuhan Liu +16
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code,…
cs.AI2026
Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning
Qiyuan Zhu, Dezhi Li, Pengyu Cheng +8
Large Reasoning Models (LRMs) excel on complex tasks through long chain-of-thought (CoT) reasoning, but their lengthy intermediate steps cause severe overthinking that inflates inf…