#multi-step reasoning
try —
2 papers match
cs.AI2026
DeepResearch Agent System
Yong Huang, Yulu Huang, for the team Collaboration
The DeepResearch Agent System is a large language model designed for deep information retrieval and multi-step autonomous research, using a sparse activation architecture that acti…
#large language models#information retrieval#multi-step reasoning#agent systems
cs.LG2026
ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
Xingjian Wu, Xuhang Zhu, Xingchen Liu +6
The paper introduces ClawTrack, a benchmark that evaluates both the final outcomes and the step-by-step reasoning processes of LLM-based autonomous agents across multiple dimension…
#large language models#autonomous agents#benchmarking#process evaluation