works on

From the 1 of 29 linked papers with an AI index.

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

Xinke Tong, Xuanming Zhang, Tianyi Tang +10

Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and mult…

cs.CL2026

Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery

Tianyun Zhong, Wangyi Jiang, Wei Wang +15

Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discovery remains largely unmeasured.…

cs.CL2026

ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

Xiang Zheng, Han Li, Wenjie Luo +15

Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings. We i…

cs.CL2026

MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning

Lingyan Wu, Xiang Zheng, Weiqi Zhai +5

Process-Level Reward Models (PRMs) are essential for guiding complex reasoning in large language models, yet existing PRM benchmarks cover only general domains such as mathematics,…

cs.CL2026

HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam

Weiqi Zhai, Zhihai Wang, Jinghang Wang +35

Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions. However, community-led analyses…

cs.CL2026

Thinking by Subtraction: Confidence-Driven Contrastive Decoding for LLM Reasoning

Lexiang Tang, Weihao Gao, Bingchen Zhao +4

Recent work on test-time scaling for large language model (LLM) reasoning typically assumes that allocating more inference-time computation uniformly improves correctness. However,…