1 paper
Seungdong Yoa, Sanghyu Yoon, Suhee Yoon +4
The evaluation of large language models (LLMs) has predominantly relied on static datasets, which offer limited scalability and fail to capture the evolving reasoning capabilities…