1 paper · 1 filter
Ruoran Xu, Wending Gao, Liyunfeng Chen +5
Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Exi…