4 papers
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
Zeyu Wang, Jingye Xu, Xiaogang Li +7
Current multimodal benchmarks for scientific reasoning primarily evaluate local information extraction -- models recognize symbols and values and then perform textual inference. Th…
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
Ben Wang, Xiaogang Li, Ruochen Gao +6
Current multimodal models handle static image recognition well, but intuitive physical reasoning remains a weakness. Predicting how objects will move and interact from a single ima…
SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy
Peiyao Xiao, Xiaogang Li, Xinyi Gao +7
As LLMs achieved breakthroughs in general reasoning, their proficiency in specialized scientific domains reveals pronounced gaps in existing benchmarks due to data contamination, i…
CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials
Chengliang Xu, Xiaogang Li, Peiyao Xiao +3
Miller-index identification from powder XRD patterns requires capabilities untested by existing multimodal benchmarks: the model must read a narrow peak location from a rendered sc…