2 papers
cs.CL2026
From Correctness to Utility: Gain-Based Prefix Evaluation for LLM Reasoning
Yuhang Zhou, Yixin Cao, Guangnan Ye
Reasoning prefixes shape the future trajectory of LLM problem solving, yet existing process reward models usually evaluate them through local step correctness. We argue that correc…
cs.AI2026
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
Zeping Li, Hongru Wang, Yiwen Zhao +7
Tool-using agents based on Large Language Models (LLMs) excel in tasks such as mathematical reasoning and multi-hop question answering. However, in long trajectories, agents often…