1 citations · 1 across the 5 of their papers we have counts for
10 papers
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
Yiheng Wang, Yixin Chen, Shuo Li +33
We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike gene…
Chem-R: Learning to Reason as a Chemist
Weida Wang, Benteng Chen, Di Zhang +14
Although large language models (LLMs) have significant potential to advance chemical discovery, current LLMs lack core chemical knowledge, produce unreliable reasoning trajectories…
ChemBOMAS: Accelerated BO in Chemistry with LLM-Enhanced Multi-Agent System
Dong Han, Zhehong Ai, Pengxiang Cai +16
Bayesian optimization (BO) is a powerful tool for scientific discovery in chemistry, yet its efficiency is often hampered by the sparse experimental data and vast search space. Her…
CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics
Weida Wang, Dongchen Huang, Jiatong Li +32
We introduce CMPhysBench, designed to assess the proficiency of Large Language Models (LLMs) in Condensed Matter Physics, as a novel Benchmark. CMPhysBench is composed of more than…
Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
Jiatong Li, Weida Wang, Qinggang Zhang +6
Large language models (LLMs), especially Explicit Long Chain-of-Thought (CoT) reasoning models like DeepSeek-R1 and QWQ, have demonstrated powerful reasoning capabilities, achievin…
QCBench: Evaluating Large Language Models on Domain-Specific Quantitative Chemistry
Jiaqing Xie, Weida Wang, Ben Gao +5
Quantitative chemistry is central to modern chemical research, yet the ability of large language models (LLMs) to perform its rigorous, step-by-step calculations remains underexplo…