1 paper
Keyang Xuan, Peiyang Song, Pan Lu +10
AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, environments, users, and other…