1 paper · 1 filter
Sanskar Pandey, Ruhaan Chopra, Angkul Puniya +1
Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that conflates helpfulness with polite subm…