3 papers
cs.AI2025
TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
Pengfei He, Zhenwei Dai, Bing He +10
Large language model (LLM)-based agents increasingly rely on tool use to complete real-world tasks. While existing works evaluate the LLMs' tool use capability, they largely focus…
cs.AI2025
Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
Yingqian Cui, Zhenwei Dai, Pengfei He +8
Large Language Models (LLMs) have achieved significant advances in reasoning tasks. A key approach is tree-based search with verifiers, which expand candidate reasoning paths and u…
cs.LG2025
Towards the Effect of Examples on In-Context Learning: A Theoretical Case Study
Pengfei He, Yingqian Cui, Han Xu +4
In-context learning (ICL) has emerged as a powerful capability for large language models (LLMs) to adapt to downstream tasks by leveraging a few (demonstration) examples. Despite i…