8 papers
Reliable Use of Lemmas via Eligibility Reasoning and SectionAware Reinforcement Learning
Zhikun Xu, Xiaodong Yu, Ben Zhou +6
Recent large language models (LLMs) perform strongly on mathematical benchmarks yet often misapply lemmas, importing conclusions without validating assumptions. We formalize lemma$…
Unbiased Visual Reasoning with Controlled Visual Inputs
Zhaonan Li, Shijie Lu, Fei Wang +11
End-to-end Vision-language Models (VLMs) often answer visual questions by exploiting spurious correlations instead of causal visual evidence, and can become more shortcut-prone whe…
Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes
Matthew W. Kenaston, Umair Ayub, Mihir Parmar +14
Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology…
Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
Xiao Ye, Jacob Dineen, Zhaonan Li +11
Medical Large language models achieve strong scores on standard benchmarks; however, the transfer of those results to safe and reliable performance in clinical workflows remains a…
RELATE-Sim: Leveraging Turning Point Theory and LLM Agents to Predict and Understand Long-Term Relationship Dynamics through Interactive Narrative Simulations
Matthew Yue, Zhikun Xu, Vivek Gupta +3
Most dating technologies optimize for getting together, not staying together. We present RELATE-Sim, a theory-grounded simulator that models how couples behave at consequential tur…
ThinkTuning: Instilling Cognitive Reflections without Distillation
Aswin RRV, Jacob Dineen, Divij Handa +4
Recent advances in test-time scaling have led to the emergence of thinking LLMs that exhibit self-reflective behaviors and multi-step reasoning. While RL drives this self-improveme…