From the 1 of 12 linked papers with an AI index.
2 citations · 2 across the 6 of their papers we have counts for
7 papers · 1 filter
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Jinhu Qi, Wentao Zhang, Siu Man Ng +4
The paper introduces TREK, a benchmark and deterministic evaluation kit for testing large language model agents on complex travel itinerary planning, requiring joint satisfaction o…
ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents
Zhihan Guo, Feiyang Xu, Yifan Li +7
A surge in academic publications calls for automated deep research (DR) systems, but accurately evaluating them is still an open problem. First, existing benchmarks often focus nar…
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
Jinhu Qi, Yifan Li, Minghao Zhao +4
Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm) carry deployment-level conseque…
Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning
Yihong Wu, Liheng Ma, Muzhi Li +7
Large Language Models (LLMs) equipped with modern Retrieval-Augmented Generation (RAG) systems often employ multi-turn interaction pipelines to interface with search engines for co…
TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research Agents
Yanyu Chen, Jiyue Jiang, Jiahong Liu +3
The evaluation of Deep Research Agents is a critical challenge, as conventional outcome-based metrics fail to capture the nuances of their complex reasoning. Current evaluation fac…
From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Retrieval-Augmented Generation Agents Development
Muzhi Li, Jinhu Qi, Yihong Wu +9
Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction trajectories. Existing datasets provide ques…