13 citations · 16 across the 4 of their papers we have counts for
1 paper · 1 filter
Marco AF Pimentel, Clément Christophe, Tathagata Raha +3
As large language models (LLMs) continue to evolve, the need for robust and standardized evaluation benchmarks becomes paramount. Evaluating the performance of these models is a co…