3 papers
cs.AI2026
Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery
Taolin Han, Yuchen Zhang, Jinghang Wang +22
Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce S…
cs.CL2026
Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
Xinke Tong, Xuanming Zhang, Tianyi Tang +10
Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and mult…
cs.CL2026
ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs
Xiang Zheng, Han Li, Wenjie Luo +15
Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings. We i…