1 paper
Dongwon Noh, Donghyeok Koh, Junghun Yuk +4
Prior benchmarks for evaluating the domain-specific knowledge of large language models (LLMs) lack the scalability to handle complex academic tasks. To address this, we introduce \…