1 paper · 1 filter
Shirin Shahabi, Spencer Graham, Haruna Isah
Evaluating language models and AI agents remains fundamentally challenging because static benchmarks fail to capture real-world uncertainty, distribution shift, and the gap between…