3 papers
cs.LG2026
Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning
Ali Larian, Qian Lin, Chang Zong Wu +1
As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands reward functions that remain robust to such ch…
cs.HC2026
Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles
Drishti Goel, Agam Goyal, Veda Duddu +8
Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond information-seeking: caregivers s…
cs.AI2026
Implicit Safety Alignment from Crowd Preferences
Qian Lin, Daniel S. Brown
Reinforcement Learning from Human Feedback (RLHF) can reveal implicit objectives such as safety considerations that go beyond task completion. In this work, we focus on the common…