3 papers
cs.CL2025
A Survey on Large Language Model Benchmarks
Shiwen Ni, Guhong Chen, Shuaimin Li +11
In recent years, with the rapid development of the depth and breadth of large language models' capabilities, various corresponding evaluation benchmarks have been emerging in incre…
cs.CV2025
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
David Ma, Huaqing Yuan, Xingjian Wang +16
Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hour…
cs.CL2025
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
Jiajun Shi, Jian Yang, Jiaheng Liu +26
Recent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchm…