2 papers
cs.CY2026
Position: Anthropomorphic Misalignment Research Needs Stronger Evidence
Vansh Gupta, Peter Nutter, Samuel Stante +5
We argue that many Anthropomorphic Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for critical safety decisions, suc…
cs.LG2026
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
Daniel Yang, Samuel Stante, Florian Redhardt +5
Reward models are central to aligning large language models (LLMs) with human preferences. Yet most approaches rely on pointwise reward estimates that overlook the epistemic uncert…