1 paper · 1 filter
Victor M. dos Santos, Andre C. Castro, Samuel L. de S. Toledo +5
The rapid advancement of Large Language Models (LLMs) has outpaced the scalability of traditional evaluation benchmarks, which remain heavily dependent on labor-intensive expert cu…