1 paper · 1 filter
Tyler Ashoff, Jordan Rodu
Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly satu…