11 papers · 1 filter
Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters
Xiao Ye, Jacob Dineen, Evan Zhu +3
Forecasters are evaluated by backtesting, which replays resolved questions and grades the probability the system would have assigned before the outcome was known. For LLMs, two cha…
Robust Asynchronous Planning via Auto-Formalization
Jiayi Zhang, Jianing Yin, Ben Zhou +1
LLMs can plan by either generating action sequences directly as a Planner or translating tasks into domain specific language for an external solver as a Formalizer. While most real…
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
Jacob Dineen, Aswin RRV, Zhikun Xu +1
Co-evolutionary self-play, where one language model generates problems and another solves them, promises curriculum learning without human supervision. The promise breaks down earl…
Reliable Use of Lemmas via Eligibility Reasoning and SectionAware Reinforcement Learning
Zhikun Xu, Xiaodong Yu, Ben Zhou +6
Recent large language models (LLMs) perform strongly on mathematical benchmarks yet often misapply lemmas, importing conclusions without validating assumptions. We formalize lemma$…
Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes
Matthew W. Kenaston, Umair Ayub, Mihir Parmar +14
Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology…
Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
Xiao Ye, Jacob Dineen, Zhaonan Li +11
Medical Large language models achieve strong scores on standard benchmarks; however, the transfer of those results to safe and reliable performance in clinical workflows remains a…