2 citations · 2 across the 8 of their papers we have counts for
8 papers · 1 filter
Implicit Reasoning Steering via Concept Chaining
Xiao Ye, Sanika Chavan, Yuxi Huang +4
Large language models often appear to reason reliably, yet on many questions repeated sampling yields both correct and incorrect answers, revealing an underlying fragility in how f…
Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters
Xiao Ye, Jacob Dineen, Evan Zhu +3
Forecasters are evaluated by backtesting, which replays resolved questions and grades the probability the system would have assigned before the outcome was known. For LLMs, two cha…
Human-Level and Beyond: Benchmarking Large Language Models Against Clinical Pharmacists in Prescription Review
Yan Yang, Mouxiao Bian, Peiling Li +10
The rapid advancement of large language models (LLMs) has accelerated their integration into clinical decision support, particularly in prescription review. To enable systematic an…
Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
Xiao Ye, Jacob Dineen, Zhaonan Li +11
Medical Large language models achieve strong scores on standard benchmarks; however, the transfer of those results to safe and reliable performance in clinical workflows remains a…
CC-LEARN: Cohort-based Consistency Learning
Xiao Ye, Shaswat Shrivastava, Zhaonan Li +6
Large language models excel at many tasks but still struggle with consistent, robust reasoning. We introduce Cohort-based Consistency Learning (CC-Learn), a reinforcement learning…
BOW: Training Language Models to Reason Over Plausible Next Words
Ming Shen, Zhikun Xu, Jacob Dineen +2
Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next words. Recent RL-based next-word r…