works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.AI2026

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

Taolin Han, Yuchen Zhang, Jinghang Wang +22

Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce S…

cs.AI2026

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

Xuan Ren, Weiqi Zhai, Tianle Pu +3

Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasonin…

cs.CL2026

Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery

Tianyun Zhong, Wangyi Jiang, Wei Wang +15

Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured.…

cs.AI2026

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

Zihan Xu, Yanzhen Chen, Xiaocheng Zhang +4

The paper presents LongMedBench, a benchmark built from MIMIC-IV electronic health records that evaluates medical agents on long-horizon clinical decision-making across multiple vi…

cs.CL2026

ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

Xiang Zheng, Han Li, Wenjie Luo +15

Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings. We i…

cs.CL2026

MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning

Lingyan Wu, Xiang Zheng, Weiqi Zhai +5

Process-Level Reward Models (PRMs) are essential for guiding complex reasoning in large language models, yet existing PRM benchmarks cover only general domains such as mathematics,…