1 paper
Ruoran Xu, Wending Gao, Liyunfeng Chen +5
Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Exi…