Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Unveiling the Entropy Dynamics of Chain-of-Thought Reasoning
Ting Xu, Xu He, Yupu Lu +4
This paper investigates the entropy dynamics of Chain-of-Thought (CoT) and uncovers a consistent two-phase structure: an Uncertainty Region of exploration transitioning sharply to…
cs.CL2026
ASTER: Agentic Scaling with Tool-integrated Extended Reasoning
Xuqin Zhang, Quan He, Zhenrui Zheng +3
Reinforcement learning (RL) has emerged as a dominant paradigm for eliciting long-horizon reasoning in Large Language Models (LLMs). However, scaling Tool-Integrated Reasoning (TIR…