1 citations · 1 across the 9 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
Pengfei He, Zhenwei Dai, Bing He +10
Large language model (LLM)-based agents increasingly rely on tool use to complete real-world tasks. While existing works evaluate the LLMs' tool use capability, they largely focus…
cs.AI2025
How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior
Zidi Xiong, Yuping Lin, Wenya Xie +5
Memory is a critical component in large language model (LLM)-based agents, enabling them to store and retrieve past executions to improve task performance over time. In this paper,…
cs.AI2025
Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
Yingqian Cui, Zhenwei Dai, Pengfei He +8
Large Language Models (LLMs) have achieved significant advances in reasoning tasks. A key approach is tree-based search with verifiers, which expand candidate reasoning paths and u…