2 papers
cs.LG2026
Out-of-Distribution Generalization of Risk Aversion in Language Models
Kristina Zhang, Junior Chinomso Okoroafor, Benjamin Maltbie +3
Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned. Misaligned but risk-averse AIs would tend to prefer low-risk, low-rewa…
cs.AI2026
Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models
Benjamin Maltbie, Shivam Raval
Large language models exhibit sycophantic tendencies, but whether this behavior varies systematically with perceived user demographics is underexplored. Inspired by intersectionali…