3 papers
cs.CV2026
Groundbench: Multi-Resolution Polygon Grounding Exposes the Geometry Gap in Vision-Language Models
Zhonghan Bian, Zhenran Wang, Jinsong Li +1
Bounding-box scores on RefCOCO-family grounding leave little room to distinguish frontier vision-language systems, yet boxes discard object shape. We introduce GroundingBench, a ma…
cs.CL2026
SocietyBench: Forecasting Counterfactual Social-World Evolution
Zhenran Wang, Zhonghan Bian, Jinsong Li +1
Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A com…
cs.CL2026
WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament
Zhenran Wang, Zhonghan Bian, Jinsong Li +1
Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evalua…