3 papers
cs.CL2026
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Jinhu Qi, Wentao Zhang, Siu Man Ng +4
Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and…
cs.AI2026
From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
Yifan Li, Shengbin Yue, Boyu Feng +6
The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks prioritize execution success, neglecting se…
cs.CL2025
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback
Chunyu Miao, Henry Peng Zou, Yangning Li +28
Large language models (LLMs) show the promise in supporting scientific research implementation, yet their ability to generate correct and executable code remains limited. Existing…