activity
20242026
collaborators

9 papers

cs.AI2026

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification

Xiangyu Zhao, Henry Hengyuan Zhao, Yiheng Wang +7

While Process Reward Models (PRMs) have achieved remarkable success in mathematical reasoning, their application in complex scientific domains-such as biology, chemistry, and physi…

cs.AI2026

MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs

Xiangyu Zhao, Wanghan Xu, Bo Liu +7

The rapid advancement of multimodal large language models (MLLMs) offers new opportunities for complex scientific challenges, yet their application in earth science-especially at t…

cs.SE2026

Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair

Zaoyu Chen, Haoran Qin, Nuo Chen +4

Smart contracts, predominantly written in Solidity and deployed on blockchains such as Ethereum, are immutable after deployment, making functional correctness critical. However, ex…

cs.AI2026

SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence

Yiheng Wang, Yixin Chen, Shuo Li +33

We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike gene…

cs.CV2025

GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning

Bo Liu, Xiangyu Zhao, Along He +3

Medical visual question answering aims to support clinical decision-making by enabling models to answer natural language questions based on medical images. While recent advances in…

cs.CL2025

EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs

Wanghan Xu, Xiangyu Zhao, Yuhao Zhou +5

Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either prese…