2 papers
cs.LG2026
Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies
Luyu Qiu, Jianing Li, Hwanhee Kim +4
Transformer-based large language models (LLMs) continue to achieve state-of-the-art performance across various natural language processing tasks. However, their subpar performance…
cs.AI2026
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
Weiyi Chen, Shuaixiong Wang, Ziyun Gao +5
The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations:…