1 paper · 1 filter
Clinton J. Wang, Dean Lee, Cristina Menghini +7
As language models master existing reasoning benchmarks, we need new challenges to evaluate their cognitive frontiers. Puzzle-solving events are rich repositories of challenging mu…