2 papers
cs.AI2026
Interactive Benchmarks
Baoqing Yue, Zihan Zhu, Yutong Han +6
Existing reasoning evaluation paradigms suffer from different limitations: fixed benchmarks are increasingly saturated and vulnerable to contamination, while preference-based evalu…
cs.AI2025
Web World Models
Jichen Feng, Yifan Zhang, Chenggong Zhang +3
Language agents increasingly require persistent worlds in which they can act, remember, and learn. Existing approaches sit at two extremes: conventional web frameworks provide reli…