1 paper · 1 filter
Amanda Dsouza, Harit Vishwakarma, Zhengyang Qi +6
The rapid progress and widespread deployment of LLMs and LLM-powered agents has outpaced our ability to evaluate them. Hand-crafted, static benchmarks are the primary tool for asse…