2 papers
cs.LG2026
FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches
Jiacheng Ding, Cong Guo, Jason Xu
We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events. For every one of the 104 ma…
cs.AI2026
AutoRAS: Learning Robust Agentic Systems with Primitive Representations
Yang Yue, Xuancheng Zhu, Yuyang Ma +7
The automated design of agentic systems offers a promising pathway for scaling large language models (LLMs) beyond single-agent reasoning. While prior work has advanced task perfor…