2 papers
cs.CL2026
SocietyBench: Forecasting Counterfactual Social-World Evolution
Zhenran Wang, Zhonghan Bian, Jinsong Li +1
Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A com…
cs.CL2026
WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament
Zhenran Wang, Zhonghan Bian, Jinsong Li +1
Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evalua…