1 paper · 1 filter
Jinyeop Song, Jeff Gore, Max Kleiman-Weiner
As language model (LM) agents become increasingly capable and adopted in real-world applications, there is a growing need for scalable evaluation frameworks beyond costly, manually…