1 paper
Hiroshi Nonaka, K. E. Perry
Evaluating the creative capabilities of large language models (LLMs) in complex tasks often requires human assessments that are difficult to scale. We introduce a novel, scalable m…