3 papers
cs.CL2026
FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect
Hazel H. Kim, Andrew M. Bean, Guilherme Affonso Ferreira de Camargo +10
We introduce FramingQA, a benchmark that measures the model sensitivity to question framing across law, medicine, finance, and robotic simulations. Large language models (LLMs) oft…
cs.CL2026
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
Hongjian Zhou, Xinyu Zou, Jinge Wu +19
Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increa…
cs.AI2025
Clinical-R1: Empowering Large Language Models for Faithful and Comprehensive Reasoning with Clinical Objective Relative Policy Optimization
Boyang Gu, Hongjian Zhou, Bradley Max Segal +6
Recent advances in large language models (LLMs) have shown strong reasoning capabilities through large-scale pretraining and post-training reinforcement learning, demonstrated by D…