3 papers
cs.AI2026
Do Clinical Models Change Treatment Decisions?
Dongkyu Cho, Miao Zhang, Rumi Chunara
Clinical foundation models are evaluated with factual or exam-style medical QA, but treatment decisions must change when patient context changes. We introduce ClinPivot, an auditab…
cs.LG2025
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
Dongkyu Derek Cho, Huan Song, Arijit Ghosh Chowdhury +6
Fine-tuning large language models (LLMs) for downstream tasks typically exhibit a fundamental safety-capability tradeoff, where improving task performance degrades safety alignment…
stat.ML2025
Implicit Updates for Average-Reward Temporal Difference Learning
Hwanwoo Kim, Dongkyu Derek Cho, Eric Laber
Temporal difference (TD) learning is a cornerstone of reinforcement learning. In the average-reward setting, standard TD() is highly sensitive to the choice of step-size and th…