2 papers
cs.LG2026
Mitigating Overthinking in Large Reasoning Models via Difficulty-aware Reinforcement Learning
Qian Wan, Ziao Xu, Luona Wei +2
Large Reasoning Models (LRMs) achieve explicit chain-of-thought expansion by imitating deep thinking behaviors of humans, demonstrating excellent performance in complex task scenar…
cs.AI2024
Gradual Vigilance and Interval Communication: Enhancing Value Alignment in Multi-Agent Debates
Rui Zou, Mengqi Wei, Jintian Feng +3
In recent years, large language models have shown exceptional performance in fulfilling diverse human needs. However, their training data can introduce harmful content, underscorin…