14 citations · 14 across the 2 of their papers we have counts for
1 paper · 1 filter
Marco AF Pimentel, Clément Christophe, Tathagata Raha +3
As large language models (LLMs) continue to evolve, the need for robust and standardized evaluation benchmarks becomes paramount. Evaluating the performance of these models is a co…