From the 1 of 12 linked papers with an AI index.
2 citations · 2 across the 6 of their papers we have counts for
12 papers
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Jinhu Qi, Wentao Zhang, Siu Man Ng +4
The paper introduces TREK, a benchmark and deterministic evaluation kit for testing large language model agents on complex travel itinerary planning, requiring joint satisfaction o…
From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
Yifan Li, Shengbin Yue, Boyu Feng +6
The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks prioritize execution success, neglecting se…
ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents
Zhihan Guo, Feiyang Xu, Yifan Li +7
A surge in academic publications calls for automated deep research (DR) systems, but accurately evaluating them is still an open problem. First, existing benchmarks often focus nar…
LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition
Yanyu Chen, Jiyue Jiang, Dianzhi Yu +8
The evolution of Large Language Model (LLM) reasoning is bottlenecked by the scarcity of high-quality process data. While self-alignment via endogenous rewards offers a solution, m…
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
Jinhu Qi, Yifan Li, Minghao Zhao +4
Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm) carry deployment-level conseque…
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
Jinhu Qi, Muzhi Li, Jiahong Liu +9
Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks autonomously, but their mul…