10 papers
ReciNet: Reciprocal Space-Aware Long-Range Modeling for Crystalline Property Prediction
Jianan Nie, Peiyao Xiao, Kaiyi Ji +1
Predicting properties of crystals from their structures is a fundamental yet challenging task in materials science. Unlike molecules, crystal structures exhibit infinite periodic a…
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
Zeyu Wang, Jingye Xu, Xiaogang Li +7
Current multimodal benchmarks for scientific reasoning primarily evaluate local information extraction -- models recognize symbols and values and then perform textual inference. Th…
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
Ben Wang, Xiaogang Li, Ruochen Gao +6
Current multimodal models handle static image recognition well, but intuitive physical reasoning remains a weakness. Predicting how objects will move and interact from a single ima…
CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials
Chengliang Xu, Xiaogang Li, Peiyao Xiao +3
Miller-index identification from powder XRD patterns requires capabilities untested by existing multimodal benchmarks: the model must read a narrow peak location from a rendered sc…
HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
Weiqi Zhai, Zhihai Wang, Jinghang Wang +35
Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions. However, community-led analyses…
DeepMTL2R: A Library for Deep Multi-task Learning to Rank
Chaosheng Dong, Peiyao Xiao, Yijia Wang +1
This paper presents DeepMTL2R, an open-source deep learning framework for Multi-task Learning to Rank (MTL2R), where multiple relevance criteria must be optimized simultaneously. D…