activity
20242026
most citedHuman-Level and Beyond: Benchmarking Large Language Models Against Clinical Pharmacists in Prescription Review

2 citations · 2 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Implicit Reasoning Steering via Concept Chaining

Xiao Ye, Sanika Chavan, Yuxi Huang +4

Large language models often appear to reason reliably, yet on many questions repeated sampling yields both correct and incorrect answers, revealing an underlying fragility in how f…

cs.CL2026

Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters

Xiao Ye, Jacob Dineen, Evan Zhu +3

Forecasters are evaluated by backtesting, which replays resolved questions and grades the probability the system would have assigned before the outcome was known. For LLMs, two cha…

cs.CL20252 cited

Human-Level and Beyond: Benchmarking Large Language Models Against Clinical Pharmacists in Prescription Review

Yan Yang, Mouxiao Bian, Peiling Li +10

The rapid advancement of large language models (LLMs) has accelerated their integration into clinical decision support, particularly in prescription review. To enable systematic an…

cs.CL2025

Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications

Xiao Ye, Jacob Dineen, Zhaonan Li +11

Medical Large language models achieve strong scores on standard benchmarks; however, the transfer of those results to safe and reliable performance in clinical workflows remains a…

cs.CL2025

CC-LEARN: Cohort-based Consistency Learning

Xiao Ye, Shaswat Shrivastava, Zhaonan Li +6

Large language models excel at many tasks but still struggle with consistent, robust reasoning. We introduce Cohort-based Consistency Learning (CC-Learn), a reinforcement learning…

cs.CL2025

BOW: Training Language Models to Reason Over Plausible Next Words

Ming Shen, Zhikun Xu, Jacob Dineen +2

Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next words. Recent RL-based next-word r…