collaborators

7 papers

cs.CV2025

Unbiased Visual Reasoning with Controlled Visual Inputs

Zhaonan Li, Shijie Lu, Fei Wang +11

End-to-end Vision-language Models (VLMs) often answer visual questions by exploiting spurious correlations instead of causal visual evidence, and can become more shortcut-prone whe…

cs.CL2025

Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications

Xiao Ye, Jacob Dineen, Zhaonan Li +11

Medical Large language models achieve strong scores on standard benchmarks; however, the transfer of those results to safe and reliable performance in clinical workflows remains a…

cs.CL2025

ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation

Qin Liu, Jacob Dineen, Yuxi Huang +4

Benchmarks are central to measuring the capabilities of large language models and guiding model development, yet widespread data leakage from pretraining corpora undermines their v…

cs.AI2025

ThinkTuning: Instilling Cognitive Reflections without Distillation

Aswin RRV, Jacob Dineen, Divij Handa +4

Recent advances in test-time scaling have led to the emergence of thinking LLMs that exhibit self-reflective behaviors and multi-step reasoning. While RL drives this self-improveme…

cs.CL2025

CC-LEARN: Cohort-based Consistency Learning

Xiao Ye, Shaswat Shrivastava, Zhaonan Li +6

Large language models excel at many tasks but still struggle with consistent, robust reasoning. We introduce Cohort-based Consistency Learning (CC-Learn), a reinforcement learning…

cs.CL2025

QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA

Jacob Dineen, Aswin RRV, Qin Liu +8

Alignment of large language models (LLMs) with principles like helpfulness, honesty, and harmlessness typically relies on scalar rewards that obscure which objectives drive the tra…