Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
Weiyi Chen, Shuaixiong Wang, Ziyun Gao +5
The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations:…
cs.AI2026★ 1 cited
When AI reviews science: Can we trust the referee?
Jialiang Wang, Yuchen Liu, Hang Xu +7
The volume of scientific submissions continues to climb, outpacing the capacity of qualified human referees and stretching editorial timelines. At the same time, modern large langu…