Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Learning When to Trust in Contextual Social Bandits
Majid Ghasemi, Mark Crowley
Robust reinforcement learning typically assumes that feedback sources are either globally trustworthy or corrupted within a fixed global budget. We identify a more subtle failure m…
cs.AI2026
Objective Decoupling in Social Reinforcement Learning: Recovering Ground Truth from Sycophantic Majorities
Majid Ghasemi, Mark Crowley
Contemporary AI alignment strategies rely on a fragile premise: that human feedback, while noisy, remains a fundamentally truthful signal. In this paper, we identify this assumptio…