activity
20242026
collaborators

20 papers

cs.AI2026

VASP Agent: An Agentic Framework for Autonomous First-principles Calculations

Zeyu Xia, Jinzhe Ma, Congjie Zheng +11

Large Language Models (LLMs) are increasingly embedded in agentic frameworks for scientific discovery. First-principles materials computation imposes a demanding standard for auton…

cs.LG2026

Do LLMs Truly Generalize in the Molecular Domain? A Perturbation-Based Analysis

Jiatong Li, Weida Wang, Changmeng Zheng +4

Large Language Models (LLMs) have recently shown promise in molecular discovery, yet a gap remains between their probabilistic nature over discrete sequential tokens and the rigid…

cs.CV2026

PolyReal: A Benchmark for Real-World Polymer Science Workflows

Wanhao Liu, Weida Wang, Jiaqing Xie +12

Multimodal Large Language Models (MLLMs) excel in general domains but struggle with complex, real-world science. We posit that polymer science, an interdisciplinary field spanning…

cs.LG2026

Equivariant Evidential Deep Learning for Interatomic Potentials

Zhongyao Wang, Taoyong Cui, Jiawen Zou +5

Uncertainty quantification (UQ) is critical for assessing the reliability of machine learning interatomic potentials (MLIPs) in molecular dynamics (MD) simulations, identifying ext…

cs.AI2026

SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence

Yiheng Wang, Yixin Chen, Shuo Li +33

We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike gene…

cs.AI2025

Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows

Wanghan Xu, Yuhao Zhou, Yifan Zhou +104

Despite advances in scientific AI, a coherent framework for Scientific General Intelligence (SGI)-the ability to autonomously conceive, investigate, and reason across scientific do…