15 citations · 22 across the 25 of their papers we have counts for
1 paper · 1 filter
Felipe Maia Polo, Lucas Weber, Leshem Choshen +3
The versatility of large language models (LLMs) led to the creation of diverse benchmarks that thoroughly test a variety of language models' abilities. These benchmarks consist of…