works on

From the 1 of 12 linked papers with an AI index.

activity
20242026
most citedTowards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

2 citations · 2 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

Jinhu Qi, Wentao Zhang, Siu Man Ng +4

The paper introduces TREK, a benchmark and deterministic evaluation kit for testing large language model agents on complex travel itinerary planning, requiring joint satisfaction o…

cs.CL2026

ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents

Zhihan Guo, Feiyang Xu, Yifan Li +7

A surge in academic publications calls for automated deep research (DR) systems, but accurately evaluating them is still an open problem. First, existing benchmarks often focus nar…

cs.CL2026

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI

Jinhu Qi, Yifan Li, Minghao Zhao +4

Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm) carry deployment-level conseque…

cs.CL2026

Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning

Yihong Wu, Liheng Ma, Muzhi Li +7

Large Language Models (LLMs) equipped with modern Retrieval-Augmented Generation (RAG) systems often employ multi-turn interaction pipelines to interface with search engines for co…

cs.CL2026

TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research Agents

Yanyu Chen, Jiyue Jiang, Jiahong Liu +3

The evaluation of Deep Research Agents is a critical challenge, as conventional outcome-based metrics fail to capture the nuances of their complex reasoning. Current evaluation fac…

cs.CL2025

From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Retrieval-Augmented Generation Agents Development

Muzhi Li, Jinhu Qi, Yihong Wu +9

Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction trajectories. Existing datasets provide ques…