3 papers
cs.CL2026
AGC-Bench: Measuring Artificial General Creativity
Roger Beaty, Vijeta Deshpande, Clin K. Y. Lai +9
Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from general intelligence. Both ques…
cs.AI2026
CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions
Sherin Muckatira, Jesse Geneson, Slava Gerovitch +3
Large language models have made substantial progress on mathematical reasoning, but existing benchmarks typically evaluate well-specified problems with final answers, step-by-step…
cs.CL2026
Playing with Words, Improving with Rewards: Training Language Models for Creative Association
Vijeta Deshpande, Namrata Shivagunde, Sherin Muckatira +5
Large Language Models (LLMs) are being applied to increasingly difficult problems and use cases. To navigate their vast solution spaces effectively, LLMs need to be creative. Yet t…