2 papers
cs.LG2026
Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care
Prabhjot Singh, Abhishek Gupta, Chris Betz +4
We reframe clinician overrides of clinical AI recommendations as implicit preference data - the same signal structure exploited by reinforcement learning from human feedback (RLHF)…
cs.RO2026
Emergent Dexterity via Diverse Resets and Large-Scale Reinforcement Learning
Patrick Yin, Tyler Westenbroek, Zhengyu Zhang +9
Reinforcement learning in massively parallel physics simulations has driven major progress in sim-to-real robot learning. However, current approaches remain brittle and task-specif…