2 papers
cs.LG2025
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks
Pengbo Shen, Yaqing Wang, Ni Mu +8
Evaluating large language models (LLMs) in complex decision-making is essential for advancing AI's ability for strategic planning and real-time adaptation. However, existing benchm…
cs.CL2024
AutoWebGLM: A Large Language Model-based Web Navigating Agent
Hanyu Lai, Xiao Liu, Iat Long Iong +8
Large language models (LLMs) have fueled many intelligent web agents, but most existing ones perform far from satisfying in real-world web navigation tasks due to three factors: (1…