Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
Bowen Ye, Rang Li, Qibin Yang +10
Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing agent benchmarks are limited by…
cs.AI2026
CuraLight: Debate-Guided Data Curation for LLM-Centered Traffic Signal Control
Qing Guo, Xinhang Li, Junyu Chen +4
Traffic signal control (TSC) is a core component of intelligent transportation systems (ITS), aiming to reduce congestion, emissions, and travel time. Recent approaches based on re…