4 citations · 4 across the 4 of their papers we have counts for
1 paper · 1 filter
Ajay Patel, Kartik Hosanagar, Ramayya Krishnan +2
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answe…