1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Anjana Arunkumar, Swaroop Mishra, Bhavdeep Sachdeva +2
Recent research has shown that language models exploit `artifacts' in benchmarks to solve tasks, rather than truly learning them, leading to inflated model performance. In pursuit…