Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Jinhu Qi, Wentao Zhang, Siu Man Ng +4
Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and…
cs.CL2026
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
Jinhu Qi, Yifan Li, Minghao Zhao +4
Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm) carry deployment-level conseque…