3 papers
cs.AI2026
Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning
Taihang Zhen, Jialiang Hong, Kai Chen +12
Large reasoning models (LRMs) often exhibit overthinking, producing verbose Chain-of-Thought (CoT) traces that increase inference cost and obscure the underlying reasoning process.…
cs.CL2026
MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models
Zecheng Tang, Baibei Ji, Ruoxi Sun +7
Existing works increasingly adopt memory-centric mechanisms to process long contexts in a segment manner, and effective memory management is one of the key capabilities that enable…
cs.LG2025
DiLoCoX: A Low-Communication Large-Scale Training Framework for Decentralized Cluster
Ji Qi, WenPeng Zhu, Li Li +6
The distributed training of foundation models, particularly large language models (LLMs), demands a high level of communication. Consequently, it is highly dependent on a centraliz…