2 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.AI2025★ 1 cited
ReasoningWeekly: A General Knowledge and Verbal Reasoning Challenge for Large Language Models
Zixuan Wu, Francesca Lucchetti, Aleksander Boruch-Gruszecki +5
Existing benchmarks for frontier models often test specialized, "PhD-level" knowledge that is difficult for non-experts to grasp. In contrast, we present a benchmark with 613 probl…
cs.CL2023★ 2 cited
Solving and Generating NPR Sunday Puzzles with Large Language Models
Jingmiao Zhao, Carolyn Jane Anderson
We explore the ability of large language models to solve and generate puzzles from the NPR Sunday Puzzle game show using PUZZLEQA, a dataset comprising 15 years of on-air puzzles.…